Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “remap”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Optimization-based, property-preserving algorithm for passive tracer transport

Here we present a new optimization-based property-preserving algorithm for passive tracer transport. The algorithm utilizes a semi-Lagrangian approach based on incremental remapping of the mass and the total tracer. However, unlike traditional semi-Lagrangian schemes, which remap the density and the tracer mixing ratio through monotone reconstruction or flux correction, we utilize an optimization-based remapping that enforces conservation and local bounds as optimization constraints. In so doing we separate accuracy considerations from preservation of physical properties to obtain a conservative, second-order accurate transport scheme that also has a notion of optimality. Moreover, we prove that the optimization-based algorithm preserves linear relationships between tracer mixing ratios. We illustrate the properties of the new algorithm using a series of standard tracer transport test problems in a plane and on a sphere.

97 MATHEMATICS AND COMPUTING↗

Conservative high-order data transfer method on generalized polygonal meshes

A conservative data transfer (remap) between two meshes is an important step of arbitrary Lagrangian-Eulerian (ALE) hydrodynamics simulations. High-order numerical methods for ALE simulations require both high-order (curvilinear) meshes and high-order remap algorithms. Here we develop a conservative and bounds-preserving method for accurate remapping of discrete fields on generalized polygonal meshes with curvilinear edges. The properties of the proposed method are studied theoretically and numerically for various (smooth and non-smooth) mesh deformations and discrete fields that represent smooth and discontinuous functions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Divergence Reduction in Monte Carlo Neutron Transport with On-GPU Asynchronous Scheduling

While Monte Carlo Neutron Transport (MCNT) is near-embarrasingly parallel, the effectively unpredictable lifetime of neutrons can lead to divergence when MCNT is evaluated on GPUs. Divergence is the phenomenon of adjacent threads in a warp executing different control flow paths; on GPUS, it reduces performance because each work group may only execute one path at a time. The process of Thread Data Remapping (TDR) resolves these discrepancies by moving data across hardware such that data in the same warp will be processed through similar paths. A common issue among prior implementations of TDR is the synchronous nature of its remapping and processing cycles, which exhaustively sort data produced by prior processing passes and exhaustively evaluate the sorted data. In another work, we defined a method of remapping data through an asynchronous scheduler which allows for work to be stored in shared memory and deferred arbitrarily until that work is a viable option for low-divergence evaluation. This article surveys a wider set of cases, with the goal of characterizing performance trends across a more comprehensive set of parameters. These parameters include cross sections of scattering/capturing/fission, use of implicit capture, source neutron counts, simulation time spans, and tuned memory allocations. Across these cases, we have recorded minimum and average execution times, as well as a heuristically tuned near-optimal memory allocation size for both synchronous and asynchronous scheduling. Across the collected data, it is shown that the asynchronous method is faster and more memory efficient in the majority of cases, and that it requires less tuning to achieve competitive performance.

Computer Science↗

CICE on a C-grid: new momentum, stress, and transport schemes for CICEv6.5

Abstract. This article presents the C-grid implementation of the CICE sea ice model, including the C-grid discretization of the momentum equation, the boundary conditions (BCs), and the modifications to the code required to use the incremental remapping transport scheme. To validate the new C-grid implementation, many numerical experiments were conducted and compared to the B-grid solutions. In idealized experiments, the standard advection method (incremental remapping with C-grid velocities interpolated to the cell corners) leads to a checkerboard pattern. A modal analysis demonstrates that this computational noise originates from the spatial averaging of C-grid velocities at corners. The checkerboard pattern can be eliminated by adjusting the departure regions to match the divergence obtained from the solution of the momentum equation. We refer to this novel approach as the edge flux adjustment (EFA) method. The C-grid discretization with edge flux adjustment allows for transport in channels that are one grid cell wide – a capability that is not possible with the B-grid discretization nor with the C-grid and standard remapping advection. Simulation results match the predicted values of a novel analytical solution for one-grid-cell-wide channels.

Lemieux, Jean-François (ORCID:0000000320845759)↗

A general interactive system for compositing digital radar and satellite data

Reynolds and Smith (1979) have considered the combined use of digital weather radar and satellite data in interactive systems for case study analysis and forecasting. Satellites view the top of clouds, whereas radar is capable of observing the detailed internal structure of clouds. The considered approach requires the use of a common coordinate system. In the present investigation, it was decided to use the satellite coordinate system as the base system in order to maintain the fullest resolution of the satellite data. The investigation is concerned with the development of a general interactive software system called RADPAK for remapping and analyzing conventional and Doppler radar data. RADPAK is implemented as a part of a minicomputer-based image processing system, called Atmospheric and Oceanographic Image Processing System. Attention is given to a general description of the RADPAK system, remapping methodology, and an example of satellite remapping.

Ghosh, K. K.↗

An interactive system for compositing digital radar and satellite data

This paper describes an approach for compositing digital radar data and GOES satellite data for meteorological analysis. The processing is performed on a user-oriented image processing system, and is designed to be used in the research mode. It has a capability to construct PPIs and three-dimensional CAPPIs using conventional as well as Doppler data, and to composite other types of data. In the remapping of radar data to satellite coordinates, two steps are necessary. First, PPI or CAPPI images are remapped onto a latitude-longitude projection. Then, the radar data are projected into satellite coordinates. The exact spherical trigonometric equations, and the approximations derived for simplifying the computations are given. The use of these approximations appears justified for most meteorological applications. The largest errors in the remapping procedure result from the satellite viewing angle parallax, which varies according to the cloud top height. The horizontal positional error due to this is of the order of the error in the assumed cloud height in mid-latitudes. Examples of PPI and CAPPI data composited with satellite data are given for Hurricane Frederic on 13 September 1979 and for a squall line on 2 May 1979 in Oklahoma.

Heymsfield, G. M.↗

Revisiting the pseudo-supercritical path method: An improved formulation for the alchemical calculation of solid–liquid coexistence

Alchemical free energy calculations via molecular dynamics have been applied to obtain thermodynamic properties related to solid–liquid equilibrium conditions, such as melting points. In recent years, the pseudo-supercritical path (PSCP) method has proved to be an important approach to melting point prediction due to its flexibility and applicability. In the present work, we propose improvements to the PSCP alchemical cycle to make it more compact and efficient through a concerted evaluation of different potential energies. The multistate Bennett acceptance ratio (MBAR) estimator was applied at all stages of the new cycle to provide greater accuracy and uniformity, which is essential concerning uncertainty calculations. In particular, for the multistate expansion stage from solid to liquid, we employed the MBAR estimator with a reduced energy function that allows affine transformations of coordinates. Free energy and mean derivative profiles were calculated at different cycle stages for argon, triazole, propenal, and the ionic liquid 1-ethyl-3-methyl-imidazolium hexafluorophosphate. Comparisons showed a better performance of the proposed method than the original PSCP cycle for systems with higher complexity, especially the ionic liquid. Furthermore, a detailed study of the expansion stage revealed that remapping the centers of mass of the molecules or ions is preferable to remapping the coordinates of each atom, yielding better overlap between adjacent states and improving the accuracy of the methodology.

25 ENERGY STORAGE↗

A seamless approach for evaluating climate models across spatial scales

In regions of the world where topography varies significantly with distance, most global climate models (GCMs) have spatial resolutions that are too coarse to accurately simulate key meteorological variables that are influenced by topography, such as clouds, precipitation, and surface temperatures. One approach to tackle this challenge is to run climate models of sufficiently high resolution in those topographically complex regions such as the North American Regionally Refined Model (NARRM) subset of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 (E3SM v2). Although high-resolution simulations are expected to provide unprecedented details of atmospheric processes, running models at such high resolutions remains computationally expensive compared to lower-resolution models such as the E3SM Low Resolution (LR). Moreover, because regionally refined and high-resolution GCMs are relatively new, there are a limited number of observational datasets and frameworks available for evaluating climate models with regionally varying spatial resolutions. As such, we developed a new framework to quantify the added value of high spatial resolution in simulating precipitation over the contiguous United States (CONUS). To determine its viability, we applied the framework to two model simulations and an observational dataset. We first remapped all the data into Hierarchical Equal-Area Iso-Latitude Pixelization (HEALPix) pixels. HEALPix offers several mathematical properties that enable seamless evaluation of climate models across different spatial resolutions including its equal-area and partitioning properties. The remapped HEALPix-based data are used to show how the spatial variability of both observed and simulated precipitation changes with resolution increases. This study provides valuable insights into the requirements for achieving accurate simulations of precipitation patterns over the CONUS. It highlights the importance of allocating sufficient computational resources to run climate models at higher temporal and spatial resolutions to capture spatial patterns effectively. Furthermore, the study demonstrates the effectiveness of the HEALPix framework in evaluating precipitation simulations across different spatial resolutions. This framework offers a viable approach for comparing observed and simulated data when dealing with datasets of varying spatial resolutions. By employing this framework, researchers can extend its usage to other climate variables, datasets, and disciplines that require comparing datasets with different spatial resolutions.

54 ENVIRONMENTAL SCIENCES↗

Impacts of spatial heterogeneity of anthropogenic aerosol emissions in a regionally refined global aerosol–climate model

Abstract. Emissions of anthropogenic aerosol and their precursors are often prescribed in global aerosol models. Most of these emissions are spatially heterogeneous at model grid scales. When remapped from low-resolution data, the spatial heterogeneity in emissions can be lost, leading to large errors in the simulation. It can also cause the conservation problem if non-conservative remapping is used. The default anthropogenic emission treatment in the Energy Exascale Earth System Model (E3SM) is subject to both problems. In this study, we introduce a revised emission treatment for the E3SM Atmosphere Model (EAM) that ensures conservation of mass fluxes and preserves the original emission heterogeneity at the model-resolved grid scale. We assess the error estimates associated with the default emission treatment and the impact of improved heterogeneity and mass conservation in both globally uniform standard-resolution (∼ 165 km) and regionally refined high-resolution (∼ 42 km) simulations. The default treatment incurs significant errors near the surface, particularly over sharp emission gradient zones. Much larger errors are observed in high-resolution simulations. It substantially underestimates the aerosol burden, surface concentration, and aerosol sources over highly polluted regions, while it overestimates these quantities over less-polluted adjacent areas. Large errors can persist at higher elevation for daily mean estimates, which can affect aerosol extinction profiles and aerosol optical depth (AOD). We find that the revised treatment significantly improves the accuracy of the aerosol emissions from surface and elevated sources near sharp spatial gradient regions, with significant improvement in the spatial heterogeneity and variability of simulated surface concentration in high-resolution simulations. In the next-generation E3SM running at convection-permitting scales where the resolved spatial heterogeneity is significantly increased, the revised emission treatment is expected to better represent the aerosol emissions as well as their lifecycle and impacts on climate.

54 ENVIRONMENTAL SCIENCES↗

Accurate modeling of parallel scientific computations

Scientific codes are usually parallelized by partitioning a grid among processors. To achieve top performance it is necessary to partition the grid so as to balance workload and minimize communication/synchronization costs. This problem is particularly acute when the grid is irregular, changes over the course of the computation, and is not known until load time. Critical mapping and remapping decisions rest on the ability to accurately predict performance, given a description of a grid and its partition. This paper discusses one approach to this problem, and illustrates its use on a one-dimensional fluids code. The models constructed are shown to be accurate, and are used to find optimal remapping schedules.

Nicol, David M.↗

Programmable remapper with single flow architecture

The invention relates to image processing systems and methods and in particular to a machine which accepts a real time video image in the form of a matrix of picture elements (pixels) and remaps such image according to a selectable one of a plurality of mapping functions to create an output matrix of pixels. Such mapping functions, or transformations, may be any one of a number of different transformations depending on the objective of the user of the system. The system remaps input images from one coordinate system to another using a set of look-up tables for the data necessary for the transform. The transforms, which are operator selectable, are precomputed and loaded into massive look-up tables. Input pixels, via the look-up tables of any particular transform selected, are mapped into output pixels with the radiance information of the input pixels being appropriately weighted. An earlier embodiment of the system included two parallel processors: a collective processor which mapped multiple input pixels into a single output pixel and an interpolative processor. The interpolative processor performed an interpolation among pixels in the input image where a given input pixel may affect the value of many output pixels. Several advantages are provided over previous embodiments in that the two distinct processors are replaced by a single processor capable of performing both types of operations (collective and interpolative) with no more complexity. Previously, there has existed no image processor or 'remapper' that can operate with sufficient speed and flexibility to permit investigating different transformation patterns in real time.

Fisher, Timothy E.↗

An implicit method for two-dimensional hydrodynamics

An implicit method for compressible multidimensional flows is presented. The method, which is strongly oriented toward astrophysical applications, enables one to simulate very subsonic flows by removing the Courant condition upon time steps. It consists of an implicit purely Lagrangian step, followed by an explicit and second-order accurate (at least in one dimension) remapping step, which is optional. When the remapping step is performed the time step is limited by the 'particle crossing time' and otherwise it is limited only by accuracy considerations. The suggested method, which results from a compromise between accuracy and efficiency, is very efficient relative to other methods. It enables the computation of many multidimensional problems in stellar evolution, such as those governed by very subsonic flows, which were not calculable with existing explicit methods.

Livne, Eli↗

Tests of a simple data merging algorithm for the GONG project

The GONG (Global Oscillation Network Group) project proposes to reduce the impact of diurnal variations on helioseismic measurements by making long-term observations of solar images from six sites placed around the globe. The sun will be observed nearly constantly for three years, resulting in the acquisition of l+ terabyte of image data. To use the solar network to maximum advantage, the images from the sites must be combined into a single time series to determine mode frequencies, amplitudes, and line widths. Initial versions of combined, i.e., merged, time series were made using a simple weighted average of data from different sites taken simultaneously. In order to accurately assess the impact of the data merge on the helioseismic measurements, a set of artificial solar disk images was made using a standard solar model and containing a well known set of oscillation modes and frequencies. This undegraded data set and data products computed from it were used to judge the relative merits of various data merging schemes. The artificial solar disk images were subjected to various instrumental and atmospheric degradations, dependent on site and time, in order to create a set of images simulating those likely to be taken at the site. The degraded artificial solar disk images for the six observing sites were combined in various ways to form merged time series of images and mode coefficients. Various forms of a weighted average were used, including an equally-weighted average, an average with weights dependent upon air mass and averages with weights dependent on various quality assurance parameters. Both the undegraded solar disk image time series and several time series made up of various combinations of the degraded solar disk images from the six sites were subjected to standard helioseismic measurement processing. This processing consisted of coordinate remapping, detrending, spherical harmonic transformation, computation of power series for the oscillation mode coefficients, and mode frequency identification. Visual and statistical evaluation of the merged data sets themselves and differences between the merged and undegraded data set shows good agreement between the two data sets. Some slight differences in image scale and registration appear between the undegraded data set and the various merged data sets. In the set of power series made from the mode coefficients of the merged data sets, some power leakage is observed into the background and into slightly lower l-value modes, especially at higher l-value mode frequencies. The results of the comparison of time series and mode oscillation frequencies of the undegraded data with those of the data merged using weighted averages indicate that, at least for p-mode solar oscillations, a weighted average of either the detrended remapped images or the mode coefficients gives good determination of the mode frequencies and adequate-to-good determination of the amplitudes and widths of the mode frequency lines. This conclusion is most advantageous to the analysis of the massive amounts of data to be received by the GONG network.

Williams, W. E.↗

An adaptive technique to maximize lossless image data compression of satellite images

Data compression will pay an increasingly important role in the storage and transmission of image data within NASA science programs as the Earth Observing System comes into operation. It is important that the science data be preserved at the fidelity the instrument and the satellite communication systems were designed to produce. Lossless compression must therefore be applied, at least, to archive the processed instrument data. In this paper, we present an analysis of the performance of lossless compression techniques and develop an adaptive approach which applied image remapping, feature-based image segmentation to determine regions of similar entropy and high-order arithmetic coding to obtain significant improvements over the use of conventional compression techniques alone. Image remapping is used to transform the original image into a lower entropy state. Several techniques were tested on satellite images including differential pulse code modulation, bi-linear interpolation, and block-based linear predictive coding. The results of these experiments are discussed and trade-offs between computation requirements and entropy reductions are used to identify the optimum approach for a variety of satellite images. Further entropy reduction can be achieved by segmenting the image based on local entropy properties then applying a coding technique which maximizes compression for the region. Experimental results are presented showing the effect of different coding techniques for regions of different entropy. A rule-base is developed through which the technique giving the best compression is selected. The paper concludes that maximum compression can be achieved cost effectively and at acceptable performance rates with a combination of techniques which are selected based on image contextual information.

Stewart, Robert J.↗

Scalability study of parallel spatial direct numerical simulation code on IBM SP1 parallel supercomputer

The implementation and the performance of a parallel spatial direct numerical simulation (PSDNS) code are reported for the IBM SP1 supercomputer. The spatially evolving disturbances that are associated with laminar-to-turbulent in three-dimensional boundary-layer flows are computed with the PS-DNS code. By remapping the distributed data structure during the course of the calculation, optimized serial library routines can be utilized that substantially increase the computational performance. Although the remapping incurs a high communication penalty, the parallel efficiency of the code remains above 40% for all performed calculations. By using appropriate compile options and optimized library routines, the serial code achieves 52-56 Mflops on a single node of the SP1 (45% of theoretical peak performance). The actual performance of the PSDNS code on the SP1 is evaluated with a 'real world' simulation that consists of 1.7 million grid points. One time step of this simulation is calculated on eight nodes of the SP1 in the same time as required by a Cray Y/MP for the same simulation. The scalability information provides estimated computational costs that match the actual costs relative to changes in the number of grid points.

Hanebutte, Ulf R.↗

Load Balancing Sequences of Unstructured Adaptive Grids

Mesh adaption is a powerful tool for efficient unstructured grid computations but causes load imbalance on multiprocessor systems. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. This paper makes several important additions to our previous work. First, a new remapping cost model is presented and empirically validated on an SP2. Next, our load balancing strategy is applied to sequences of dynamically adapted unstructured grids. Results indicate that our framework is effective on many processors for both steady and unsteady problems with several levels of adaption. Additionally, we demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required for a fine initial mesh. Finally, we show that the data remapping overhead can be significantly reduced by applying our heuristic processor reassignment algorithm.

Biswas, Rupak↗

PLUM: Parallel Load Balancing for Unstructured Adaptive Meshes

Dynamic mesh adaption on unstructured grids is a powerful tool for computing large-scale problems that require grid modifications to efficiently resolve solution features. Unfortunately, an efficient parallel implementation is difficult to achieve, primarily due to the load imbalance created by the dynamically-changing nonuniform grid. To address this problem, we have developed PLUM, an automatic portable framework for performing adaptive large-scale numerical computations in a message-passing environment. First, we present an efficient parallel implementation of a tetrahedral mesh adaption scheme. Extremely promising parallel performance is achieved for various refinement and coarsening strategies on a realistic-sized domain. Next we describe PLUM, a novel method for dynamically balancing the processor workloads in adaptive grid computations. This research includes interfacing the parallel mesh adaption procedure based on actual flow solutions to a data remapping module, and incorporating an efficient parallel mesh repartitioner. A significant runtime improvement is achieved by observing that data movement for a refinement step should be performed after the edge-marking phase but before the actual subdivision. We also present optimal and heuristic remapping cost metrics that can accurately predict the total overhead for data redistribution. Several experiments are performed to verify the effectiveness of PLUM on sequences of dynamically adapted unstructured grids. Portability is demonstrated by presenting results on the two vastly different architectures of the SP2 and the Origin2OOO. Additionally, we evaluate the performance of five state-of-the-art partitioning algorithms that can be used within PLUM. It is shown that for certain classes of unsteady adaption, globally repartitioning the computational mesh produces higher quality results than diffusive repartitioning schemes. We also demonstrate that a coarse starting mesh produces high quality load balancing, at a fraction of the cost required a fine initial mesh. Results indicate that our parallel load balancing strategy will remain viable on large numbers of processors.

Oliker, Leonid↗