Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel projection algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Three-Dimensional Grid Visualization for Planning Activities: A Dubai Case Study

National Laboratory of the Rockies (NLR), in collaboration with the Dubai Electricity and Water Authority (DEWA) and Infra-X, has undertaken the Energy Visualization Analysis Project. The aim of this project is to enhance analytical and 3D visualization capabilities for distribution network planning and renewable energy integration. As modern grid continues to evolve with large-scale solar PV deployment and emerging distributed energy resources (DERs), the ability to effectively analyze, visualize, and communicate complex grid behaviors has become increasingly critical. The project focuses on developing empirical use cases based on real distribution feeder data and engineering workflows, ensuring the outcomes are directly aligned with operational environment. Through time-series power flow simulations and nodal hosting capacity analysis, the study quantifies the impacts of high PV penetration on voltage and thermal limits within representative 11 kV feeders. These analyses identify specific nodes and conditions where DER integration challenges arise. Furthermore, a Battery Energy Storage System (BESS) optimization algorithm was applied to determine the optimal size and placement of storage systems that can mitigate network constraints and enhance hosting capacity. The comparative results between base-case and BESS-augmented scenarios clearly demonstrate improvements in network stability and load management efficiency. In parallel, the NLR team developed an immersive 3D visualization framework, enabling interactive exploration of grid simulations using commodity head-mounted display (HMD) systems. This framework transforms conventional 2D simulation data into spatially intuitive visual environments - allowing engineers to analyze feeder conditions, PV hosting potential, and BESS effects in real time. This report represents the first foundational phase in establishing a visualization-driven analytical ecosystem. It provides a methodological foundation for data integration, visualization architecture, and simulation-based decision support, paving the way for large-scale adoption of immersive visualization across DEWA's Smart Grid Initiative, R&D activities, and future network resilience studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Projective Hedging Algorithms for Multistage Stochastic Programming, Supporting Distributed and Asynchronous Implementation

Here we propose a decomposition algorithm for multistage stochastic programming that resembles the progressive hedging method of Rockafellar and Wets but is provably capable of several forms of asynchronous operation. We derive the method from a class of projective operator splitting methods fairly recently proposed by Combettes and Eckstein, significantly expanding the known applications of those methods. Our derivation assures convergence for convex problems whose feasible set is compact, subject to some standard regularity conditions and a mild “fairness” condition on subproblem selection. The method’s convergence guarantees are deterministic and do not require randomization, in contrast to other proposed asynchronous variations of progressive hedging. Computational experiments described in an online appendix show the method to outperform progressive hedging on large-scale problems in a highly parallel computing environment.

97 MATHEMATICS AND COMPUTING↗

MFIX DEM Enhancement for Industry-Relevant Flows (Final Report)

The overall goal of this two-phase project is to implement performance improvements of the Multiphase Flow with Interphase Exchanges (MFIX) Discrete Element Model (DEM) code that enable a transformative shift for industrial use. Prior to this effort, the largest simulations performed using MFIX are O(10 7 ) particles. This falls short of the O(10 9 ) particle simulations that must be completed on a timescale of days or weeks (vs. months or years) to enable simulations with physically-relevant domain sizes to be incorporated into industrial design cycles within five years. This was accomplished by tailoring best-in-class practices to bear on the unique challenges posed by the MFIX-DEM algorithm and code base. Scientific simulations (e.g., in cosmology, turbulent combustion) routinely use massively parallel computing to update far more particles in short wall clock times. Results from Phase 1 (1.5 years in duration) indicated significant gains in speed were possible for a wide range of benchmark cases. Moreover, a survey sent to >35 companies indicates that the timing is ideal for such an enhanced tool, with >80% of the respondents indicating that DEM is already value-added or will be within the next 5 years, and >70% of the respondents indicating that improved speed is the top computational priority. In Phase 2 (3.5 years in duration), the two major barriers that hinder industry from effectively using multiphase Computational Fluid Dynamics (CFD) to cut costs and improve performance, namely computational overhead and confidence in predictions, continued to be addressed. Regarding the former, the results from Phase 1 to guide the effort, with enhancements focused on an improved time-stepping algorithm and particle sorting. Four target problems of 1 billion particles each and increasing complexity were identified: homogeneous cooling, tumbler with continuous particle size distribution, discharge from a rectangular hopper and a cylindrical riser. Each of these were successfully simulated for relevant time scales (on order of seconds) using less than 24 hours of wall clock time. These represent the first 1-billion particle DEM simulations performed with MFIX, namely using the MFIX-Exa code. This code is currently under development at NETL in collaboration with Lawrence Berkeley National Laboratory. Regarding the second barrier on predictive uncertainty, experiments from Phase 1 (interacting nozzles - hydrodynamics only) and Phase 2 (very small-scale segregation experiments) were used to demonstrate the ability of two simplified approaches to uncertainty quantification (UQ). By limiting the number of particles, UQ based on the simplified treatment was compared to standard UQ, which was shown to have much higher computational demands. Experiments were also performed on a pilot-scale stripper unit to provide validation data for future CFD-DEM simulations and UQ.

20 FOSSIL-FUELED POWER PLANTS↗

HPC-enabled computation of demand models at scale

The purpose of this project is to examine the energy impact of urban-scale traffic for the Los Angeles Basin by developing and implementing a scalable traffic assignment model. An energy optimization function will be posed and when integrated into the optimization code for travel assignment it can be mathematically proven to converge. The energy optimization function can then be compared to the typical travel time optimization that is traditionally used in traffic assignment models. The analysis will begin with static traffic assignment models with the routing for all origin and destinations computed in parallel on high performance computing facilities. Convergence of the numerical methods rely on the solution of convex programs (or extensions of these). This step will mostly consist of demonstrating the ability to parallelize the Frank Wolfe algorithm on various platforms.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

HPC4Mobilty w/ UCB

The purpose of this project is to examine the energy impact of urban-scale traffic for the Los Angeles Basin by developing and implementing a scalable traffic assignment model. An energy optimization function will be posed and when integrated into the optimization code for travel assignment it can be mathematically proven to converge. The energy optimization function can then be compared to the typical travel time optimization that is traditionally used in traffic assignment models. The analysis will begin with static traffic assignment models with the routing for all origin and destinations computed in parallel on high performance computing facilities. Convergence of the numerical methods rely on the solution of convex programs (or extensions of these). This step will mostly consist of demonstrating the ability to parallelize the Frank Wolfe algorithm on various platforms. This work will contribute to LBNL’s efforts to develop new processes, analytical tools, program designs, and business models to advance the state of the art in next-generation sustainable transportation solutions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Scaled ILU Smoothers for Navier-Stokes Pressure Projection

Incomplete LU (ILU) smoothers are effective in the algebraic multigrid (AMG) V-cycle for reducing high-frequency components of the error. However, the requisite direct triangular solves are comparatively slow on GPUs. Previous work has demonstrated the advantages of Jacobi iteration as an alternative to direct solution of these systems. Depending on the threshold and fill-level parameters chosen, the factors can be highly nonnormal and Jacobi is unlikely to converge in a low number of iterations. We demonstrate that row scaling can reduce the departure from normality, allowing us to replace the inherently sequential solve with a rapidly converging Richardson iteration. There are several advantages beyond the lower compute time. Scaling is performed locally for a diagonal block of the global matrix because it is applied directly to the factor. Further, an ILUT Schur complement smoother maintains a constant GMRES iteration count as the number of MPI ranks increases, and thus parallel strong-scaling is improved. Our algorithms have been incorporated into hypre, and we demonstrate improved time to solution for linear systems arising in the Nalu-Wind and PeleLM pressure solvers. For large problem sizes, GMRES+AMG executes at least five times faster when using iterative triangular solves compared with direct solves on massively parallel GPUs.

algebraic multigrid↗

IPRT Polarized Radiative Transfer Model Intercomparison Project-Phase A

The polarization state of electromagnetic radiation scattered by atmospheric particles such as aerosols, cloud droplets, or ice crystals contains much more information about the optical and microphysical properties than the total intensity alone. For this reason an increasing number of polarimetric observations are performed from space, from the ground and from aircraft. Polarized radiative transfer models are required to interpret and analyse these measurements and to develop retrieval algorithms exploiting polarimetric observations. In the last years a large number of new codes have been developed, mostly for specific applications. Benchmark results are available for specific cases, but not for more sophisticated scenarios including polarized surface reflection and multi-layer atmospheres. The International Polarized Radiative Transfer (IPRT) working group of the International Radiation Commission (IRC) has initiated a model intercomparison project in order to fill this gap. This paper presents the results of the first phase A of the IPRT project which includes ten test cases, from simple setups with only one layer and Rayleigh scattering to rather sophisticated setups with a cloud embedded in a standard atmosphere above an ocean surface. All scenarios in the first phase A of the intercomparison project are for a one-dimensional plane-parallel model geometry. The commonly established benchmark results are available at the IPRT website

radiative transfer↗

Broadband Characterization and Circuit Model Development of Transmission-Scale Transformers

This report describes broadband measurements of transmission-scale transformers typical in the electric power grid. This work was performed as part of the EMP Resilient Grid LDRD project at Sandia National Laboratories to generate circuit models that can be used for high-altitude electromagnetic pulse (HEMP) coupling simulations and response predictions. The objective of the work was to obtain characterization data of substation yard equipment across a frequency range relevant to HEMP. Vector network analyzer measurements up to 100 MHz were performed on two power transformers at ABB-Hitachi and a single ITEC potential transformer. Custom cable breakouts were designed to interface with the transformer terminals and provide ground connections to the chassis at the base of the transformer bushings. The three-phase terminals of the power transformers were measured as a common mode impedance using a parallel resistive splitter, and the single-phase terminals of the potential transformer were measured directly. A vector fitting algorithm was used to empirically fit circuit models to the resulting two-port networks and input impedances of the measured objects. Simplified circuit representations of the input impedances were also generated to assess the degree of precision needed for high-altitude electromagnetic pulse response predictions, which were performed in Sandia's XYCE circuit simulator platform. HEMP coupling simulations using the transformer models showed significant reduction in the voltage peak and broadening in the pulse width seen at the power transformer compared to the traveling wave voltage. This indicated the importance of the load condition when defining the coupled insult in an electric power substation. Simplified circuit models showed a similar voltage at the transformer with a smoothed waveform. The presence of potential transformers in the simulation did not significantly change the simulated voltage at the power transformer. Single-port input impedance models were also developed to define load conditions when transfer characteristics were not necessary.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Constrained Multipoint Aerodynamic Shape Optimization Using an Adjoint Formulation and Parallel Computers

An aerodynamic shape optimization method that treats the design of complex aircraft configurations subject to high fidelity computational fluid dynamics (CFD), geometric constraints and multiple design points is described. The design process will be greatly accelerated through the use of both control theory and distributed memory computer architectures. Control theory is employed to derive the adjoint differential equations whose solution allows for the evaluation of design gradient information at a fraction of the computational cost required by previous design methods. The resulting problem is implemented on parallel distributed memory architectures using a domain decomposition approach, an optimized communication schedule, and the MPI (Message Passing Interface) standard for portability and efficiency. The final result achieves very rapid aerodynamic design based on a higher order CFD method. In order to facilitate the integration of these high fidelity CFD approaches into future multi-disciplinary optimization (NW) applications, new methods must be developed which are capable of simultaneously addressing complex geometries, multiple objective functions, and geometric design constraints. In our earlier studies, we coupled the adjoint based design formulations with unconstrained optimization algorithms and showed that the approach was effective for the aerodynamic design of airfoils, wings, wing-bodies, and complex aircraft configurations. In many of the results presented in these earlier works, geometric constraints were satisfied either by a projection into feasible space or by posing the design space parameterization such that it automatically satisfied constraints. Furthermore, with the exception of reference 9 where the second author initially explored the use of multipoint design in conjunction with adjoint formulations, our earlier works have focused on single point design efforts. Here we demonstrate that the same methodology may be extended to treat complete configuration designs subject to multiple design points and geometric constraints. Examples are presented for both transonic and supersonic configurations ranging from wing alone designs to complex configuration designs involving wing, fuselage, nacelles and pylons.

Reuther, James↗

Dynamic Mode Decomposition of Unsteady Pressure-Sensitive Paint Measurements for the NASA Unitary Plan Wind Tunnel Tests

This paper discusses the Dynamic Mode Decomposition (DMD) of the Unsteady Pressure-Sensitive Paint (uPSP) measurements, which were collected with four Phantom high-speed cameras at a constant sample frequency in the Ascent Transient Aerodynamics Test (ATAT) of the Space Launch System (SLS) Block 1 cargo vehicle with the Unitary Plan Wind Tunnel (UPWT) 11-by-11-foot Transonic Wind Tunnel in September 2019 at NASA Ames Research Center. The conventional DMD algorithm is based on the Singular Value Decomposition (SVD). For the data with zero mean, the DMD is equivalent to the Discrete Fourier Transform (DFT). Since the uPSP is mainly used to determine the unsteady property of the aerodynamic flow, the DMD of the uPSP measurements is implemented in two steps: (1) subtract the mean value from the uPSP measurement; (2) apply the Fast Fourier Transform (FFT) on the resulting data with zero mean. The DMD of the uPSP measurements with FFT has two advantages: (1) the FFT algorithm is well known for its computational efficiency, therefore, compared to the SVD-based DMD algorithm, the DMD with FFT reduces the computation time; (2) the DMD with FFT can be easily implemented in parallel processing. The DMD outputs were generated with the execution in parallel of a code in C, with libraries of FFTW for FFT and MPI/OpenMP for parallel processing, on the NASA Pleiades supercomputer. In this paper, the results of DMD of the uPSP measurements in the tests of Mach sweep runs of the SLS ATAT are presented, and the effectiveness of the DMD of the uPSP measurements in the diagnosis of the unsteady, aerodynamic phenomena is demonstrated. The work described in this paper is a part of NASA’s development of a new state-of-the-art uPSP capability in production wind tunnels. Funding for this research was provided by the NASA Aeroscience Evaluation and Test Capabilities Project.

Pressure-Sensitive Paint↗

Stage Separation Performance Analysis Project

Stage separation process is an important phenomenon in multi-stage launch vehicle operation. The transient flowfield coupled with the multi-body systems is a challenging problem in design analysis. The thermodynamics environment with burning propellants during the upper-stage engine start in the separation processes adds to the complexity of the-entire system. Understanding the underlying flow physics and vehicle dynamics during stage separation is required in designing a multi-stage launch vehicle with good flight performance. A computational fluid dynamics model with the capability to coupling transient multi-body dynamics systems will be a useful tool for simulating the effects of transient flowfield, plume/jet heating and vehicle dynamics. A computational model using generalize mesh system will be used as the basis of this development. The multi-body dynamics system will be solved, by integrating a system of six-degree-of-freedom equations of motion with high accuracy. Multi-body mesh system and their interactions will be modeled using parallel computing algorithms. Adaptive mesh refinement method will also be employed to enhance solution accuracy in the transient process.

Chen, Yen-Sen↗

Distributed Prognostics based on Structural Model Decomposition

Within systems health management, prognostics focuses on predicting the remaining useful life of a system. In the model-based prognostics paradigm, physics-based models are constructed that describe the operation of a system and how it fails. Such approaches consist of an estimation phase, in which the health state of the system is first identified, and a prediction phase, in which the health state is projected forward in time to determine the end of life. Centralized solutions to these problems are often computationally expensive, do not scale well as the size of the system grows, and introduce a single point of failure. In this paper, we propose a novel distributed model-based prognostics scheme that formally describes how to decompose both the estimation and prediction problems into independent local subproblems whose solutions may be easily composed into a global solution. The decomposition of the prognostics problem is achieved through structural decomposition of the underlying models. The decomposition algorithm creates from the global system model a set of local submodels suitable for prognostics. Independent local estimation and prediction problems are formed based on these local submodels, resulting in a scalable distributed prognostics approach that allows the local subproblems to be solved in parallel, thus offering increases in computational efficiency. Using a centrifugal pump as a case study, we perform a number of simulation-based experiments to demonstrate the distributed approach, compare the performance with a centralized approach, and establish its scalability. Index Terms-model-based prognostics, distributed prognostics, structural model decomposition ABBREVIATIONS

centrifugal pump↗

3-D CFD in a day - The laser digitizer project

The computation of airflow over complex configurations requires a complete description of the geometry. This can be obtained from CAD data, from blueprints, or from actual models. In any case, the time required is currently estimated at 4 to 6 months. It is proposed to shorten this time by a factor of 10 to 100 through the use of automated software, a fast, highly parallel computer and a three-dimensional laser digitizer. This device can provide (x,y,z) coordinates of surface points at rates exceeding 14,500/sec. Thus, it is possible to digitize an entire model in a few minutes. The accuracy of measurement on a flat white surface is better than 0.005 inches. Higher accuracy is available at higher cost. This work discusses the challenges which remain to be addressed. In particular, the surface point data need to be converted into a surface description, the surface description needs to be made into a surface grid, and the surface grid used to make a volume grid for the flow solver. Algorithms are kept in place or in mind for all of these problems. Integration of the more mature flow solution and visualization algorithms then allows generation of solution graphics directly from a wind tunnel model.

Merriam, Marshal↗

Transport Analysis & Optimization in a MW-Scale CO2 Electrolyzer (Final Report)

As Twelve continues to scale up their CO2 electrolyzers, both in the size of a single cell and in the number of cells used in a stack, thermal management becomes a growing concern, since excess heat can affect reaction yield and accelerate degradation. In this project, we aim to computationally explore how the anode flow fields used in Twelve’s CO2 electrolyzers function as heat exchangers. In particular, using a homogenized model of a CO2 electrolyzer, we first estimate the amount of heat generated in a cell. Then, we develop a computational fluid dynamics (CFD) model of the so-called “flow field”, i.e. a flow manifold, based on Twelve’s CAD drawings, to evaluate how these flow fields perform as a heat exchanger for the generated heat. We explore both a single cell and a 3-cell stack operating in parallel, where heat generated in one cell can now be transferred to another cell. We evaluate how performance is affected when environmental heat losses are taken into account. Finally, we leverage topology optimization to explore the types of design features a computational optimization algorithm would suggest to supplement our intuition. Overall, our work aims to provide design recommendations for CO2 electrolyzer flow fields and provides a foundation for future studies of flow field optimization.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE↗

Porting Classical Approaches for Quantum Simulations to Quantum Computers

Simulating quantum many-body systems is one of the most promising problems in which we might anticipate that quantum computers should show quantum advantage. Unfortunately, there is still a gap between this promise and actual practice. New quantum algorithms need to be developed and the current quantum algorithms have various difficulties - e.g efficient state preparation - which must be overcome and improved upon. In many cases, classical approaches need to be ported over to quantum devices. In this project we have developed a suite of new quantum algorithms which makes progress in this regard. We developed a new optimization scheme for variational quantum eigensolvers, UBOS, which mitigates problems with local minimas and barren plateaus while improving convergence to the ground state by an order of magnitude. We developed a new way to utilize qubitization to find ground states of nearly frustration-free Hamiltonians faster than all previous methods. We developed a series of state preparation techniques which helps initialize parameterized quantum circuits into reasonable starting points on which quantum algorithms are then applied. In addition to the development of novel algorithms, it is critical to have classical simulation techniques for approximately simulating quantum circuits which can be used to benchmark and understand quantum algorithms. Toward that end, we developed a novel POVM formalism to simulate quantum circuits as well as exemplify the massive parallelization of tensor network methodologies. Finally, we developed physical understanding of entanglement phase transitions such as many-body localization and random tensor networks.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Qu8its for quantum simulations of lattice quantum chromodynamics

We explore the utility of d = 8 qudits, qu8its, for quantum simulations of the dynamics of 1+1⁢D SU(3) lattice quantum chromodynamics, including a mapping for arbitrary number of flavors and lattice size and a reorganization of the Hamiltonian for efficient time evolution. Recent advances in parallel gate applications, along with the shorter application times of single-qudit operations compared with two-qudit operations, lead to significant projected advantages in quantum simulation fidelities and circuit depths using qu8its rather than qubits. The number of two-qudit entangling gates required for time evolution using qu8its is found to be more than a factor of 5 fewer than for qubits. Here, we anticipate that the developments presented in this work will enable improved quantum simulations to be performed using emerging quantum hardware.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗