Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81

Fluid dynamic and thermal performance of a slotted cylinder at low Reynolds number

The fluid dynamic and thermal performance of a circular cylinder with a slot parallel to the flow is numerically investigated. The study utilized the semi-implicit finite volume multi-material algorithm MPM-ICE, a component of the Uintah framework. The normalized slot width s/D ranges from 0.1 - 0.3, introducing an additional heat transfer surface area between ~ 10 and ~ 50%, and a mass reduction between ~ 13 and ~ 38% in the cylinder. We assumed two-dimensional incompressible flow and simulated a Reynolds number Re D between 100 and 1000. The slotted cylinders are found to have a total drag force reduction up to ~ 45%, compared to a solid cylinder despite the additional viscous drag force in the slot. Convection heat transfer is enhanced up to ~ 70%. Further, the slotted cylinder performance index, defined as the ratio of the heat rate to the drag force, increases up to maximum of ~ 3, indicating better overall thermal fluid performance. An entropy analysis showed the best performance index occurs at the highest Re D . Correlations for drag coefficient and Nusselt number are proposed along with an entropy optimization method.

42 ENGINEERING↗

Quantum Electrodynamics Coupled-Cluster at Scale: High-Performance Implementation for Complex Systems

Coupled-cluster theory (CC) is a highly accurate and versatile method for simulating complex interactions within quantum systems. The extension of CC theory to model mixed electron-photon processes with quantum electrodynamics (QED) has improved our capability to predict cavity-modified chemistry, a field where photons are used as cost-effective and eco-friendly alternatives to catalyze/inhibit chemical reactions. However, calculations with CC methods, even without incorporating QED effects, are often prohibitively expensive. Simulations of larger systems require scalable infrastructures that exist for traditional CC methods but not for QED-CC methods. As such, we present a GPU-enabled, high-performance, open-source implementation of the quantum electrodynamics coupled-cluster method with single and double excitations (QED-CCSD) within the ExaChem quantum chemistry software package. ExaChem relies on the Tensor Algebra for Many-body Methods (TAMM) infrastructure: a parallel heterogeneous tensor library designed to achieve scalable performance on modern heterogeneous supercomputing platforms. Furthermore, we discuss theoretical foundations, algorithmic details, and numerical benchmarks to showcase the larger systems that ExaChem can simulate and how the integration of photonic degrees-of-freedom alters their ground-state properties.

Basis sets↗

Predictive Inverse Model for Advective Heat Transfer in a Short–Circuited Fracture: Dimensional Analysis, Machine Learning, and Field Demonstration

Identifying fluid flow maldistribution in planar geometries is a well–established problem in subsurface science/engineering. Of particular importance to the thermal performance of enhanced (or “engineered”) geothermal systems is identifying the existence of nonuniform (i.e., heterogeneous) permeability and subsequently predicting advective heat transfer. Here, machine learning via a genetic algorithm (GA) identifies the spatial distribution of an unknown permeability field in a two–dimensional Hele–Shaw geometry (i.e., parallel plates). The inverse problem is solved by minimizing the L2 norm between simulated residence time distribution (RTD) and measurements of an inert tracer breakthrough curve (BTC) (C–Dot nanoparticle). Principal component analysis (PCA) of spatially correlated permeability fields enabled reduction of the parameter space by more than a factor of 10 and restricted the inverse search to reservoir–scale permeability variations. Thermal experiments and tracer tests conducted at the mesoscale Altona Field Laboratory (AFL) demonstrate that the method accurately predicts the effects of extreme flow channeling on heat transfer in a single bedding–plane rock fracture. However, this is only true when the permeability distributions provide adequate matches to both tracer RTD and frictional pressure loss. Without good agreement to frictional pressure loss, it is still possible to match a simulated RTD to measurements, but subsequent predictions of heat transfer are grossly inaccurate. Here, the results of this study suggest that it is possible to anticipate the thermal effects of flow maldistribution, but only if both simulated RTDs and frictional pressure loss between fluid inlets and outlets are in good agreement with measurements.

42 ENGINEERING↗

2020 IEEE PES Innovative Smart Grid Technologies Europe (ISGT-Europe)

Recent proliferation of distributed energy sources in distribution or sub-transmission systems necessitates close monitoring of these three-phase power grids which typically operate under unbalanced loading conditions. Unlike the transmission systems where the network equations are commonly based on the positive sequence component models, a detailed three phase model will have to be used in implementing network applications for these systems. In the specific case of the state estimator, where measurement and parameter errors may bias the solution, bad data and parameter error detection algorithms should also be incorporated. Implementing the state estimator and error detection algorithms for three-phase systems impose additional computational burden and modifications to the state estimation code. This paper proposes a practical solution to avoid these issues by using synchronized phasor measurements and modal decoupling. The previously developed parameter error detection algorithm based on the normalized Lagrange multipliers (NLM) test is applied to the measurements independently in each mode in parallel, not only saving CPU time but also avoiding new code development for a three-phase estimator. Different parameter error scenarios are created and tested to verify the effectiveness of the proposed error detection approach.

Khalili, Ramtin↗

Restoration Strategy for Active Distribution Systems Considering Endogenous Uncertainty in Cold Load Pickup

Cold load pickup (CLPU) phenomenon is identified as the persistent power inrush upon a sudden load pickup after an outage. Under the active distribution system (ADS) paradigm, where distributed energy resources (DERs) are extensively installed, the decreased outage duration can induce a strong interdependence between CLPU pattern and load pickup decisions. In this paper, we propose a novel modelling technique to tractably capture the decision-dependent uncertainty (DDU) inherent in the CLPU process. Subsequently, a two-stage stochastic decision-dependent service restoration (SDDSR) model is constructed, where first stage searches for the optimal switching sequences to decide step-wise network topology, and the second stage optimizes the detailed generation schedule of DERs as well as the energization of switchable loads. Further, to tackle the computational burdens introduced by mixed-integer recourse, the progressive hedging algorithm (PHA) is utilized to decompose the original model into scenario-wise subproblems that can be solved in parallel. The numerical test on modified IEEE 123-node test feeders has verified the efficiency of our proposed SDDSR model and provided fresh insights into the monetary and secure values of DDU quantification.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark↗

Broadband Characterization and Circuit Model Development of Transmission-Scale Transformers

This report describes broadband measurements of transmission-scale transformers typical in the electric power grid. This work was performed as part of the EMP Resilient Grid LDRD project at Sandia National Laboratories to generate circuit models that can be used for high-altitude electromagnetic pulse (HEMP) coupling simulations and response predictions. The objective of the work was to obtain characterization data of substation yard equipment across a frequency range relevant to HEMP. Vector network analyzer measurements up to 100 MHz were performed on two power transformers at ABB-Hitachi and a single ITEC potential transformer. Custom cable breakouts were designed to interface with the transformer terminals and provide ground connections to the chassis at the base of the transformer bushings. The three-phase terminals of the power transformers were measured as a common mode impedance using a parallel resistive splitter, and the single-phase terminals of the potential transformer were measured directly. A vector fitting algorithm was used to empirically fit circuit models to the resulting two-port networks and input impedances of the measured objects. Simplified circuit representations of the input impedances were also generated to assess the degree of precision needed for high-altitude electromagnetic pulse response predictions, which were performed in Sandia's XYCE circuit simulator platform. HEMP coupling simulations using the transformer models showed significant reduction in the voltage peak and broadening in the pulse width seen at the power transformer compared to the traveling wave voltage. This indicated the importance of the load condition when defining the coupled insult in an electric power substation. Simplified circuit models showed a similar voltage at the transformer with a smoothed waveform. The presence of potential transformers in the simulation did not significantly change the simulated voltage at the power transformer. Single-port input impedance models were also developed to define load conditions when transfer characteristics were not necessary.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The solution of the three-dimensional viscous-compressible Navier-Stokes equations on a vector computer

The development of a vectorized computer code for the solution of the three-dimensional viscous-compressible Navier-Stokes equations is described. The code is applied on the CDC STAR-100 vector computer which is capable of achieving high result rates when a high degree of parallelism is present in the computations. The computational technique is an explicit time-split MacCormack predictor-corrector algorithm. Since a large volume of data is processed and virtual memory utilized, a data management scheme based on interleaving is used. The program has been applied to obtain the solution of the laminar supersonic flow about a family of three-dimensional corners. The equations of motion are expressed in a generalized form relative to a uniform rectangular computational domain. The metric coefficient and boundary conditions must be supplied for the corresponding physical domain. For calculations with 30,000 grid points, a computational rate of 0.00015 seconds per grid point per time step is observed.

Smith, R. E.↗

Array distribution in data-parallel programs

We consider distribution at compile time of the array data in a distributed-memory implementation of a data-parallel program written in a language like Fortran 90. We allow dynamic redistribution of data and define a heuristic algorithmic framework that chooses distribution parameters to minimize an estimate of program completion time. We represent the program as an alignment-distribution graph. We propose a divide-and-conquer algorithm for distribution that initially assigns a common distribution to each node of the graph and successively refines this assignment, taking computation, realignment, and redistribution costs into account. We explain how to estimate the effect of distribution on computation cost and how to choose a candidate set of distributions. We present the results of an implementation of our algorithms on several test problems.

Chatterjee, Siddhartha↗

A High-Order Direct Solver for Helmholtz Equations with Neumann Boundary Conditions

In this study, a compact finite-difference discretization is first developed for Helmholtz equations on rectangular domains. Special treatments are then introduced for Neumann and Neumann-Dirichlet boundary conditions to achieve accuracy and separability. Finally, a Fast Fourier Transform (FFT) based technique is used to yield a fast direct solver. Analytical and experimental results show this newly proposed solver is comparable to the conventional second-order elliptic solver when accuracy is not a primary concern, and is significantly faster than that of the conventional solver if a highly accurate solution is required. In addition, this newly proposed fourth order Helmholtz solver is parallel in nature. It is readily available for parallel and distributed computers. The compact scheme introduced in this study is likely extendible for sixth-order accurate algorithms and for more general elliptic equations.

Sun, Xian-He↗

The Geostationary Lighting Mapper (GLM) for GOES-R: A New Operational Capability to Improve Storm Forecasts and Warnings

The next generation Geostationary Operational Environmental Satellite (GOES-R) series is a follow on to the existing GOES system currently operating over the Western Hemisphere. Superior spacecraft and instrument technology will support expanded detection of environmental phenomena, resulting in more timely and accurate forecasts and warnings. Advancements over current GOES capabilities include a new capability for total lightning detection (cloud and cloud-to-ground flashes) from the Geostationary Lightning Mapper (GLM), and improved spectral (3x), spatial (4x), and temporal (5x) resolution for the Advanced Baseline Imager (ABI). The GLM, an optical transient detector and imager operating in the near-IR at 777.4 nm will map all (in-cloud and cloud-to-ground) lighting flashes continuously day and night with near-uniform spatial resolution of 8 km with a product refresh rate of less than 20 sec over the Americas and adjacent oceanic regions, from the west coast of Africa (GOES-E) to New Zealand (GOES-W) when the constellation is fully operational. This will aid in forecasting severe storms and tornado activity, and convective weather impacts on aviation safety and efficiency. In parallel with the instrument development (a prototype and 4 flight models), a GOES-R Risk Reduction Team and Algorithm Working Group Lightning Applications Team have begun to develop the Level 2 algorithms and applications. Proxy total lightning data from the NASA Lightning Imaging Sensor on the Tropical Rainfall Measuring Mission (TRMM) satellite and regional test beds are being used to develop the pre-launch algorithms and applications, and also improve our knowledge of thunderstorm initiation and evolution. Real time lightning mapping data are being provided in an experimental mode to selected National Weather Service (NWS) national centers and forecast offices via the GOES-R Proving Ground to help improve our understanding of the application of these data in operational settings and facilitate Day-1 user readiness for this new capability.

Goodman, Steven J.↗

The GOES-R GeoStationary Lightning Mapper (GLM)

The Geostationary Operational Environmental Satellite (GOES-R) is the next series to follow the existing GOES system currently operating over the Western Hemisphere. Superior spacecraft and instrument technology will support expanded detection of environmental phenomena, resulting in more timely and accurate forecasts and warnings. Advancements over current GOES capabilities include a new capability for total lightning detection (cloud and cloud-to-ground flashes) from the Geostationary Lightning Mapper (GLM), and improved capability for the Advanced Baseline Imager (ABI). The Geostationary Lighting Mapper (GLM) will map total lightning activity (in-cloud and cloud-to-ground lighting flashes) continuously day and night with near-uniform spatial resolution of 8 km with a product refresh rate of less than 20 sec over the Americas and adjacent oceanic regions. This will aid in forecasting severe storms and tornado activity, and convective weather impacts on aviation safety and efficiency among a number of potential applications. In parallel with the instrument development (a prototype and 4 flight models), a GOES-R Risk Reduction Team and Algorithm Working Group Lightning Applications Team have begun to develop the Level 2 algorithms (environmental data records), cal/val performance monitoring tools, and new applications using GLM alone, in combination with the ABI, merged with ground-based sensors, and decision aids augmented by numerical weather prediction model forecasts. Proxy total lightning data from the NASA Lightning Imaging Sensor on the Tropical Rainfall Measuring Mission (TRMM) satellite and regional test beds are being used to develop the pre-launch algorithms and applications, and also improve our knowledge of thunderstorm initiation and evolution. An international field campaign planned for 2011-2012 will produce concurrent observations from a VHF lightning mapping array, Meteosat multi-band imagery, Tropical Rainfall Measuring Mission (TRMM) Lightning Imaging Sensor (LIS) overpasses, and related ground and in-situ lightning and meteorological measurements in the vicinity of Sao Paulo. These data will provide a new comprehensive proxy data set for algorithm and application development.

Goodman, Steven J.↗

Dynamic Mode Decomposition of Unsteady Pressure-Sensitive Paint Measurements for the NASA Unitary Plan Wind Tunnel Tests

This paper discusses the Dynamic Mode Decomposition (DMD) of the Unsteady Pressure-Sensitive Paint (uPSP) measurements, which were collected with four Phantom high-speed cameras at a constant sample frequency in the Ascent Transient Aerodynamics Test (ATAT) of the Space Launch System (SLS) Block 1 cargo vehicle with the Unitary Plan Wind Tunnel (UPWT) 11-by-11-foot Transonic Wind Tunnel in September 2019 at NASA Ames Research Center. The conventional DMD algorithm is based on the Singular Value Decomposition (SVD). For the data with zero mean, the DMD is equivalent to the Discrete Fourier Transform (DFT). Since the uPSP is mainly used to determine the unsteady property of the aerodynamic flow, the DMD of the uPSP measurements is implemented in two steps: (1) subtract the mean value from the uPSP measurement; (2) apply the Fast Fourier Transform (FFT) on the resulting data with zero mean. The DMD of the uPSP measurements with FFT has two advantages: (1) the FFT algorithm is well known for its computational efficiency, therefore, compared to the SVD-based DMD algorithm, the DMD with FFT reduces the computation time; (2) the DMD with FFT can be easily implemented in parallel processing. The DMD outputs were generated with the execution in parallel of a code in C, with libraries of FFTW for FFT and MPI/OpenMP for parallel processing, on the NASA Pleiades supercomputer. In this paper, the results of DMD of the uPSP measurements in the tests of Mach sweep runs of the SLS ATAT are presented, and the effectiveness of the DMD of the uPSP measurements in the diagnosis of the unsteady, aerodynamic phenomena is demonstrated. The work described in this paper is a part of NASA’s development of a new state-of-the-art uPSP capability in production wind tunnels. Funding for this research was provided by the NASA Aeroscience Evaluation and Test Capabilities Project.

Pressure-Sensitive Paint↗

Cabana: A Performance Portable Library for Particle-Based Simulations

Particle-based simulations are ubiquitous throughout many fields of computational science and engineering, spanning the atomistic level with molecular dynamics (MD), to mesoscale particle-in-cell (PIC) simulations for solid mechanics, device-scale modeling with PIC methods for plasma physics, and massive N-body cosmology simulations of galaxy structures, with many other methods in between (Hockney & Eastwood, 1989). While these methods use particles to represent significantly different entities with completely different physical models, many low-level details are shared including performant algorithms for short- and/or long-range particle interactions, multi-node particle communication patterns, and other data management tasks such as particle sorting and neighbor list construction. Cabana is a performance portable library for particle-based simulations, developed as part of the Co-Design Center for Particle Applications (CoPA) within the Exascale Computing Project (ECP) (Alexander et al., 2020). The CoPA project and its full development scope, including ECP partner applications, algorithm development, and similar software libraries for quantum MD, is described in (Mniszewski et al., 2021). Cabana uses the Kokkos library for on-node parallelism (Edwards et al., 2014; Trott et al., 2022), enabling simulation on multi-core CPU and GPU architectures, and MPI for GPU-aware, multi-node communication. Cabana provides particle simulation capabilities on almost all current Kokkos backends, including serial execution, OpenMP (including OpenMP-Target for GPUs), CUDA (NVIDIA GPUs), HIP (AMD GPUs), and SYCL (Intel GPUs), providing a clear path for the coming generation of accelerator-based exascale hardware. Cabana builds on Kokkos by providing new particle data structures and particle algorithms resulting in a similar execution policy-based, node-level programming model that is intended to be used in addition to the core Kokkos library within an application. Cabana is designed as an application and physics agnostic, but particle-specific toolkit which can either be used to generate a new application, or to be used as needed in existing applications at various levels of invasiveness including through interfaces that wrap user memory in existing data structures.

97 MATHEMATICS AND COMPUTING↗

Algorithm and code development for unsteady three-dimensional Navier-Stokes equations

Aeroelastic tests require extensive cost and risk. An aeroelastic wind-tunnel experiment is an order of magnitude more expensive than a parallel experiment involving only aerodynamics. By complementing the wind-tunnel experiments with numerical simulations, the overall cost of the development of aircraft can be considerably reduced. In order to accurately compute aeroelastic phenomenon it is necessary to solve the unsteady Euler/Navier-Stokes equations simultaneously with the structural equations of motion. These equations accurately describe the flow phenomena for aeroelastic applications. At ARC a code, ENSAERO, is being developed for computing the unsteady aerodynamics and aeroelasticity of aircraft, and it solves the Euler/Navier-Stokes equations. The purpose of this cooperative agreement was to enhance ENSAERO in both algorithm and geometric capabilities. During the last five years, the algorithms of the code have been enhanced extensively by using high-resolution upwind algorithms and efficient implicit solvers. The zonal capability of the code has been extended from a one-to-one grid interface to a mismatching unsteady zonal interface. The geometric capability of the code has been extended from a single oscillating wing case to a full-span wing-body configuration with oscillating control surfaces. Each time a new capability was added, a proper validation case was simulated, and the capability of the code was demonstrated.

Obayashi, Shigeru↗

An Open-Source Parallel EMT Simulation Framework: Preprint

As the integration level of inverter-based resources (IBR) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗

Development, Verification and Validation of Parallel, Scalable Volume of Fluid CFD Program for Propulsion Applications

There are many instances involving liquid/gas interfaces and their dynamics in the design of liquid engine powered rockets such as the Space Launch System (SLS). Some examples of these applications are: Propellant tank draining and slosh, subcritical condition injector analysis for gas generators, preburners and thrust chambers, water deluge mitigation for launch induced environments and even solid rocket motor liquid slag dynamics. Commercially available CFD programs simulating gas/liquid interfaces using the Volume of Fluid approach are currently limited in their parallel scalability. In 2010 for instance, an internal NASA/MSFC review of three commercial tools revealed that parallel scalability was seriously compromised at 8 cpus and no additional speedup was possible after 32 cpus. Other non-interface CFD applications at the time were demonstrating useful parallel scalability up to 4,096 processors or more. Based on this review, NASA/MSFC initiated an effort to implement a Volume of Fluid implementation within the unstructured mesh, pressure-based algorithm CFD program, Loci-STREAM. After verification was achieved by comparing results to the commercial CFD program CFD-Ace+, and validation by direct comparison with data, Loci-STREAM-VoF is now the production CFD tool for propellant slosh force and slosh damping rate simulations at NASA/MSFC. On these applications, good parallel scalability has been demonstrated for problems sizes of tens of millions of cells and thousands of cpu cores. Ongoing efforts are focused on the application of Loci-STREAM-VoF to predict the transient flow patterns of water on the SLS Mobile Launch Platform in order to support the phasing of water for launch environment mitigation so that vehicle determinantal effects are not realized.

West, Jeff↗