Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51

High Performance Fortran for Aerospace Applications

This paper focuses on the use of High Performance Fortran (HPF) for important classes of algorithms employed in aerospace applications. HPF is a set of Fortran extensions designed to provide users with a high-level interface for programming data parallel scientific applications, while delegating to the compiler/runtime system the task of generating explicitly parallel message-passing programs. We begin by providing a short overview of the HPF language. This is followed by a detailed discussion of the efficient use of HPF for applications involving multiple structured grids such as multiblock and adaptive mesh refinement (AMR) codes as well as unstructured grid codes. We focus on the data structures and computational structures used in these codes and on the high-level strategies that can be expressed in HPF to optimally exploit the parallelism in these algorithms.

Mehrotra, Piyush↗

Parallelized modelling and solution scheme for hierarchically scaled simulations

This two-part paper presents the results of a benchmarked analytical-numerical investigation into the operational characteristics of a unified parallel processing strategy for implicit fluid mechanics formulations. This hierarchical poly tree (HPT) strategy is based on multilevel substructural decomposition. The Tree morphology is chosen to minimize memory, communications and computational effort. The methodology is general enough to apply to existing finite difference (FD), finite element (FEM), finite volume (FV) or spectral element (SE) based computer programs without an extensive rewrite of code. In addition to finding large reductions in memory, communications, and computational effort associated with a parallel computing environment, substantial reductions are generated in the sequential mode of application. Such improvements grow with increasing problem size. Along with a theoretical development of general 2-D and 3-D HPT, several techniques for expanding the problem size that the current generation of computers are capable of solving, are presented and discussed. Among these techniques are several interpolative reduction methods. It was found that by combining several of these techniques that a relatively small interpolative reduction resulted in substantial performance gains. Several other unique features/benefits are discussed in this paper. Along with Part 1's theoretical development, Part 2 presents a numerical approach to the HPT along with four prototype CFD applications. These demonstrate the potential of the HPT strategy.

Padovan, Joe↗

GPU Acceleration of Large-Scale Full-Frequency GW Calculations

Many-body perturbation theory is a powerful method to simulate electronic excitations in molecules and materials starting from the output of density functional theory calculations. By implementing the theory efficiently so as to run at scale on the latest leadership high-performance computing systems it is possible to extend the scope of GW calculations. Here, we present a GPU acceleration study of the full-frequency GW method as implemented in the WEST code. Excellent performance is achieved through the use of (i) optimized GPU libraries, e.g., cuFFT and cuBLAS, (ii) a hierarchical parallelization strategy that minimizes CPU-CPU, CPU-GPU, and GPU-GPU data transfer operations, (iii) nonblocking MPI communications that overlap with GPU computations, and (iv) mixed precision in selected portions of the code. A series of performance benchmarks has been carried out on leadership high-performance computing systems, showing a substantial speedup of the GPU-accelerated version of WEST with respect to its CPU version. Good strong and weak scaling is demonstrated using up to 25 920 GPUs. Finally, we showcase the capability of the GPU version of WEST for large-scale, full-frequency GW calculations of realistic systems, e.g., a nanostructure, an interface, and a defect, comprising up to 10 368 valence electrons.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The OMPS Limb Profiler Instrument: Two-Dimensional Retrieval Algorithm

The upcoming Ozone Mapper and Profiler Suite (OMPS), which will be launched on the NPOESS Preparatory Project (NPP) platform in early 2011, will continue monitoring the global distribution of the Earth's middle atmosphere ozone and aerosol. OMPS is composed of three instruments, namely the Total Column Mapper (heritage: TOMS, OMI), the Nadir Profiler (heritage: SBUV) and the Limb Profiler (heritage: SOLSE/LORE, OSIRIS, SCIAMACHY, SAGE III). The ultimate goal of the mission is to better understand and quantify the rate of stratospheric ozone recovery. The focus of the paper will be on the Limb Profiler (LP) instrument. The LP instrument will measure the Earth's limb radiance (which is due to the scattering of solar photons by air molecules, aerosol and Earth surface) in the ultra-violet (UV), visible and near infrared, from 285 to 1000 nm. The LP simultaneously images the whole vertical extent of the Earth's limb through three vertical slits, each covering a vertical tangent height range of 100 km and each horizontally spaced by 250 km in the cross-track direction. Measurements are made every 19 seconds along the orbit track, which corresponds to a distance of about 150km. Several data analysis tools are presently being constructed and tested to retrieve ozone and aerosol vertical distribution from limb radiance measurements. The primary NASA algorithm is based on earlier algorithms developed for the SOLSE/LORE and SAGE III limb scatter missions. All the existing retrieval algorithms rely on a spherical symmetry assumption for the atmosphere structure. While this assumption is reasonable in most of the stratosphere, it is no longer valid in regions of prime scientific interest, such as polar vortex and UTLS regions. The paper will describe a two-dimensional retrieval algorithm whereby the ozone distribution is simultaneously retrieved vertically and horizontally for a whole orbit. The retrieval code relies on (1) a forward 2D Radiative Transfer code (to model limb radiances within a non-uniform atmosphere and evaluate 2D analytical partial derivatives) and (2) an optimal estimator inversion routine. The algorithm uses the typically sparse nature of the kernel matrices as well as fast matrix inversion techniques to allow for fast inversion of limb data with efficient memory management (as was done for MIPAS data processing). While the method has so far only been developed in the context of Single Scatter, the paper will show how the CPU intensive Multiple Scatter modeling can be implemented using parallel CPU processing. Initial results will be presented in terms of retrieved ozone profiles and code performance.

Rault, Didier F.↗

A Versatile Simulation Framework for Elastodynamic Modeling of Structural Health Monitoring

Structural health monitoring (SHM) has the capacity to reduce failure by detecting damage during service life, by periodic, automated monitoring. Guided Wave (GW) Ultrasound is a common SHM approach for aerospace structures. Modelling the physics of GW SHM systems provides a route for understanding system dependencies, capabilities and limitations as damage evolves during service life. Such a toolset can strengthen the understanding of the connection between GW SHM results and the true material state. The most useful modelling tools are those that provide versatile solutions with respect to the simulated component geometry and computational grid connectivity. This work details a versatile application programming interface (API) for the elastodynamic finite integration technique for modelling GW SHM of metals. The custom code implementation, EFIT-CompCell, allows for the modelling of diverse geometries by automatically balancing the message passing interface parallelization layout. The user provides the basic parameters of the simulation and the software automatically performs an initial balancing based on anticipated computational loads, and establishes the CPU communication patterns for any geometry. This work describes the programming philosophy and code structure used to create EFIT-CompCell and compares its performance and capacity to simulation tools that are more specialized for specific architectures. Results are presented for a simulation of GW SHM of an aluminum fuselage section being tested by the FAA. The simulation consists of 733M voxels which took approximately 70 hours to complete 25000 time steps using 40 Intel Xeon E5-4650v2 Ivy Bridge processor cores.

Gregory, Elizabeth D.↗

Coupling Carbon Oxidation and Surface Recession in Direct-Simulation Monte Carlo Code, SPARTA

Ablative thermal protection system (TPS) materials for spacecraft are composites that are often made out of carbon-based reinforcement and a polymeric matrix. They endure high-temperature oxidation and surface recession when re-entering Earth’s atmosphere. Ablation is the result of many coupled and competing thermal, mechanical, and chemical phenomena, and it is difficult to isolate the role of each on the overall degradation of the TPS. Here we develop an ablation model for material recession coupled explicitly to finite rate carbon oxidation in complex microstructures. In this work, Stochastic PArallel Rarified-gas Time-accurate Analyzer (SPARTA), a direct-simulation Monte Carlo (DSMC) code, is modified to allow oxidation-driven ablation of implicitly defined carbon surfaces. In SPARTA, implicit surfaces are generated from the grid corner point values via a marching cubes algorithm, therefore creating a new set of surface elements every time ablation is performed. The finite-rate oxidation model developed by Gopalan et. al, was adapted to tally surface reactions and other surface data on a per-grid cell basis. The ablation functionality was also adjusted so once the reactions have occurred, the number of reactions leading to CO formation can be converted to corner point reduction values; therefore, carbon removal is directly proportional to surface recession. We also develop robust algorithms which handle the evolution of the flow cells and solid material regions, including split cells (flow cell divided in two by a solid surface). Finally, we demonstrate our implicit chemistry model for 2D and 3D geometries by producing reaction statistics and detailed visualization of oxidation-induced material recession at the microscale.

V Arias↗

Automation and optimization of stopping and range of ions in matter simulation runtime

Prior to every ion implantation experiment a simulation of the ion range and other relevant parameters is performed using Monte-Carlo based codes. Although increasing computing power has improved the speed of these calculations, the demands on Monte-Carlo codes are also increasing, requiring evaluation of the optimal number of simulations while ensuring accuracy within threshold bounds. We evaluate the “Stopping and Range of Ions in Matter” (SRIM) code due to its widespread usage. We show how dividing simulations into multiple parallel simulations with different random seeds can lead to calculation speedup and find lower bounds for the required number of ion traces simulated based on an exemplar system of a Ga focused ion beam and a high energy C beam as used in high linear energy transfer testing. Here our results indicate simulations can yield results within the underlying data accuracy of SRIM at 10X and 100X shorter simulation time than the SRIM default values.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Saturation of the kinetic ballooning instability due to the electron parallel nonlinearity

The electron parallel nonlinearity (EPN) is implemented in the gyrokinetic particle-in-cell turbulence code GEM [Y. Chen and S. E. Parker, J. Comp. Phys. 220, 839 (2007)]. Application to the Cyclone Base Case reveals a strong effect of EPN on the saturated heat transport above the kinetic ballooning mode (KBM) threshold. Evidence is provided to show that the strong effect is associated with the electron radial motion due to magnetic fluttering, which turns fine structures of the KBM eigenmode in radius into fine structures in velocity and increases the magnitude of the EPN term in the kinetic equation.

Gyrokinetic simulations↗

Simulations of edge and SOL turbulence in diverted negative and positive triangularity plasmas

Optimizing the performance of magnetic confinement fusion devices is critical to achieving an attractive fusion reactor design. Negative triangularity (NT) scenarios have been shown to achieve excellent levels of energy confinement, while avoiding edge localized modes. Modeling turbulent transport in the edge and SOL is key in understanding the impact of NT on turbulence and extrapolating the results to future devices and regimes. Previous gyrokinetic turbulence studies have reported beneficial effects of NT across a broad range of parameters. However, most simulations have focused on the inner plasma region, neglecting the impact of NT on the outermost edge. In this work, we investigate the effect of NT in edge and scrape-off layer simulations, including the magnetic X-point and separatrix. For the first time, we employ a multi-fidelity approach, combining global, non-linear gyrokinetic simulations with drift-reduced fluid simulations, to gain a deeper understanding of the underlying physics at play. First-principles simulations using the GENE-X code demonstrate that in comparable NT and PT geometries, similar profiles are achieved, while the turbulent heat flux is reduced by more than 50% in NT. Comparisons with results from the drift-reduced fluid turbulence code GRILLIX suggest that the turbulence is driven by trapped electron modes. The parallel heat flux width on the divertor targets is reduced in NT, primarily due to a lower spreading factor S.

GENE-X↗

python binder for libROM

pylibROM introduces python binder for libROM through pybind11. Through python interface, users who are familiar with python can take advantage of capabilities available in libROM that is fully parallelized C++ library for reduced order modeling. libROM itself is an open source code developed at LLNL. libROM is available at https://github.com/LLNL/libROM .

Choi, Youngsoo↗

2D reactive transport model of shale chemical weathering and biogeochemical fluxes along a mountainous hillslope, East River Watershed, Colorado: Input files and simulation results

This data package contains input files and simulation results for a two-dimensional (2D) reactive transport model used to quantitatively analyze the coupled hydrological and biogeochemical processes governing shale weathering and associated biogeochemical fluxes under realistic environmental conditions in the high-elevation East River Watershed. These data support the conclusions presented in Stolze et al. (Water Resources Research, under review), "Model-based interpretation of solute exports and carbon partitioning during shale weathering in a mountainous hillslope". The model simulates atmospheric-subsurface gas exchange, subsurface water flow, and shale weathering processes under dynamic, year-scale conditions along a shale-underlain hillslope located in the East River watershed. The simulations were performed using the PFLOTRAN flow and reactive transport code and executed on the Perlmutter supercomputer to leverage its large-scale parallel computing capabilities. The data package contains two zipped folders, "model_input_files" and "simulation_results", and one readme.txt file. "model_input_files" contains the necessary input files to run the calibrated base-base model presented in Stolze et al. (Water Resources Research, under review). "simulation_results" contains a single hdf5 file ("Output_2D_hillslope_model.h5") which includes the results of simulation performed using the base-case model. This file can be opened with HDFView 3.1.4, Python, or MATLAB. "readme.txt" contains relevant information about the base-case model and provides guidelines on how to run the associated input files provided in the folder "model_input_files". Furthermore, readme.txt provides information regarding the model results provided in "Output_2D_hillslope_model.h5" such as matrix dimensionality and output units. Field datasets used to evaluate model performance were collected at three monitoring wells located along a hillslope transect (PLM1, PLM2, and PLM3). Dissolved ion concentration data were collected from November 2016 to October 2021 for Ca, Mg, DIC, Na, K, SO4 (Dong et al., 2025 - dic_npoc_data_2014_2024.zip - DOI:10.15485/1660459; Williams et al., 2025 - anion_data_2014_2024.zip - DOI:10.15485/1668054; Dong et al., 2025 - cation_data_2014_2024.zip - DOI:10.15485/1668055). Note that we used the files named er_PLM1_xx_yy, er_PLM2_xx_yy, and er_PLM3_xx_yy where xx stands for the name of the aqueous species and yy stands for the depth where the measurements were performed. Soil water content ([0 - 1] m) and water table depth were collected from November 2016 to October 2021 (Wan et al., 2024 - Dynamic_water_table__depthsFig2b.csv and Soil_water_content_Fig4e.csv - DOI:10.15485/2322567). Gaseous CO2 concentration were collected from October 2020 to December 2021(Wan et al., 2024 - Soil_CO2_concentrations_Fig4h.csv - DOI:10.15485/2322567) Gaseous CO2 flux from the subsurface to the atmosphere were collected in the vicinity of PLM2 from October 2019 to May 2022 (Wu et al., 2025). Soil microbial biomass concentration was measured from August 2016 to June 2017 (Sorensen et al., 2019 - 2017_East_River_Pumphouse_Microbial_Biomass__1_.csv - DOI:10.15485/1577267) All field data are published as CSV files compatible with Microsoft Excel, MATLAB, and Python, or as text files. The coordinates of the monitoring wells and the CO2(g) flux sensor in the coordinate system WGS84 are: -PLM1: [38.9197710 ; -106.9492750] -PLM2: [38.9201580 ; -106.9487170] -PLM3: [38.9207843 ; -106.9483668] -PLM4: 38.9210060 ; -106.9479528] -CO2(g) flux sensor: [38.9199180 ; -106.9489906] ------------------------------------------------------------------------------------------- This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research used resources of the National Energy Research Scientific Computing Center (NERSC), a Department of Energy User Facility using NERSC award BER-ERCAP 23980, BER-ERCAP 28550, and BER-ERCAP 33789.

54 ENVIRONMENTAL SCIENCES↗

SCALE Activities in FY23 [Slides]

This presentation shows that the SCALE 6.3 is available from RSICC. The lecture touches on new features including the new ENDF/B-VIII.0 data including covariances. It shows that the updated parallel infrastructure enables parallel capability on Windows. Production release with maintenance until 2026 at minimum addressing Code or data bugs, Performance, Ease of installation. There is a New Government Use Agreement (GUA) for SCALE 7.0 beta access. With site licenses available for non-commercial testing and feedback, handled through ORNL technology transfer.

97 MATHEMATICS AND COMPUTING↗

Acquisition Of Spread-Spectrum Code

Effects of Doppler shift and data modulation taken into account. Two advanced schemes for acquisition of direct-sequence spread-spectrum codes proposed. M1-Lag correlator in each strip of spread-spectrum-code detector operates at different offset code-chip time. Each offset represents assumed (tentative) Doppler shift. Schemes have highly parallel architecture implemented with currently available technology. Possible to use hybrid parallel/serial architecture in which acquisition time varies in inverse proportion to number of correlators and fast-Fourier-transform processors.

Cheng, Unjeng↗

Collisional electrostatic ion cyclotron waves as a possible source of energetic heavy ions in the magnetosphere

A new mechanism is proposed for the source of energetic heavy ions (NO/+/, O2/+/, and O/+/) found in the magnetosphere. Simulations using a multispecies particle simulation code for resistive current-driven electrostatic ion cyclotron waves show transverse and parallel bulk heating of bottomside ionospheric heavy ion populations. The dominant mechanism for the transverse bulk heating is resonant ion heating by wave-particle ion trapping. Using a linear kinetic dispersion relation for a magnetized, collisional, homogenous, and multiion plasma, it is found that collisional electrostatic ion cyclotron waves near the NO(+), O2(+), and O(+) gyrofrequencies are unstable to field-aligned currents of 50 microA/sq m for a typical bottomside ionosphere.

Providakes, Jason↗

The engine design engine. A clustered computer platform for the aerodynamic inverse design and analysis of a full engine

An application for parallel computation on a combined cluster of powerful workstations and supercomputers was developed. A Parallel Virtual Machine (PVM) is used as message passage language on a macro-tasking parallelization of the Aerodynamic Inverse Design and Analysis for a Full Engine computer code. The heterogeneous nature of the cluster is perfectly handled by the controlling host machine. Communication is established via Ethernet with the TCP/IP protocol over an open network. A reasonable overhead is imposed for internode communication, rendering an efficient utilization of the engaged processors. Perhaps one of the most interesting features of the system is its versatile nature, that permits the usage of the computational resources available that are experiencing less use at a given point in time.

Sanz, J.↗

Development of Message Passing Routines for High Performance Parallel Computations

Computational Fluid Dynamics (CFD) calculations require a great deal of computing power for completing the detailed computations involved. In an effort shorten the time it takes to complete such calculations they are implemented on a parallel computer. In the case of a parallel computer some sort of message passing structure must be used to communicate between the computers because, unlike a single machine, each computer in a parallel computing cluster does not have access to all the data or run all the parts of the total program. Thus, message passing is used to divide up the data and send instructions to each machine. The nature of my work this summer involves programming the "message passing" aspect of the parallel computer. I am working on modifying an existing program, which was written with OpenMP, and does not use a multi-machine parallel computing structure, to work with Message Passing Interface (MPI) routines. The actual code is being written in the FORTRAN 90 programming language. My goal is to write a parameterized message passing structure that could be used for a variety of individual applications and implement it on Silicon Graphics Incorporated s (SGI) IRIX operating system. With this new parameterized structure engineers would be able to speed up computations for a wide variety of purposes without having to use larger and more expensive computing equipment from another division or another NASA center.

Summers, Edward K.↗

Plume-Free Stream Interaction Heating Effects During Orion Crew Module Reentry

During reentry of the Orion Crew Module (CM), vehicle attitude control will be performed by firing reaction control system (RCS) thrusters. Simulation of RCS plumes and their interaction with the oncoming flow has been difficult for the analysis community due to the large scarf angles of the RCS thrusters and the unsteady nature of the Orion capsule backshell environments. The model for the aerothermal database has thus relied on wind tunnel test data to capture the heating effects of thruster plume interactions with the freestream. These data are only valid for the continuum flow regime of the reentry trajectory. A Direct Simulation Monte Carlo (DSMC) analysis was performed to study the vehicle heating effects that result from the RCS thruster plume interaction with the oncoming freestream flow at high altitudes during Orion CM reentry. The study was performed with the DSMC Analysis Code (DAC). The inflow boundary conditions for the jets were obtained from Data Parallel Line Relaxation (DPLR) computational fluid dynamics (CFD) solutions. Simulations were performed for the roll, yaw, pitch-up and pitch-down jets at altitudes of 105 km, 125 km and 160 km as well as vacuum conditions. For comparison purposes (see Figure 1), the freestream conditions were based on previous DAC simulations performed without active RCS to populate the aerodynamic database for the Orion CM. Other inputs to the analysis included a constant Orbital reentry velocity of 7.5 km/s and angle of attack of 160 degrees. The results of the study showed that the interaction effects decrease quickly with increasing altitude. Also, jets with highly scarfed nozzles cause more severe heating compared to the nozzles with lower scarf angles. The difficulty of performing these simulations was based on the maximum number density and the ratio of number densities between the freestream and the plume for each simulation. The lowest altitude solutions required a substantial amount of computational resources (up to 1800 processors) to simulate approximately 2 billion molecules for the refined (adapted) solutions.

Marichalar, J.↗

Thermodynamic Limits of Redox-Based Thermochemical Processes (REDOTHERM)

Solar thermochemical fuel production is a potential pathway for the production of sustain liquid drop-in fuels, which can help decarbonize the aviation and maritime sectors. In an attempt to analyze the commercial viability of this technology, several studies have been conducted, including system and technoeconomic analysis (TEA) modeling. However, most studies to date simply assume a given redox reactor efficiency, which is significantly higher than demonstrated values to date. While it is widely recognized that utilizing a counter-current flow (CF) configuration could increase the redox reactor efficiency, an over-simplification in the thermodynamic modeling may lead to unphysical results which has been included in multiple publications. The fact that the solar redox reactor is the least developed component in the process chain makes it hard to identify technology gaps and evaluate pathways to deployment at scale using this approach. In this work, a thermodynamic model for a moving oxide system has been developed, in a general form that allows to analyze the system for different redox-active materials, under a wide range of operating conditions, for both parallel and countercurrent flows. The model capabilites are demonstrated, and the model's code will be shared as an open-source on GitHub in the next few months.

chemical looping↗