Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

USING ER@CEBAF TO SHOW THAT A MULTIPASS ERL CAN DRIVE AN XFEL

A multi-pass recirculating superconducting CW linac offers a cost effective path to a multi-user facility with unprecedented scientific and industrial reach over a wide range of disciplines. We propose such a facility as an option for a potential UK-XFEL. Energy Recovery enables multi-MHz FEL sources, for example, an X-ray FEL oscillator or regenerative amplifier FEL. Additionally, combining with external lasers and/or self-interaction would provide access to MeV and GeV gamma-rays via inverse Compton scattering at high average power for nuclear and particle physics applications. An opportunity exists to demonstrate the necessary point-to-parallel longitudinal matches to drive an XFEL and successfully energy recover at the upcoming 5-pass up, 5-pass down Energy Recovery experiment on CEBAF at JLab termed ER@CEBAF. We show candidate matches and simulations supporting the minimal necessary modifications to CEBAF this will require. This includes linearisation of the longitudinal phase space in the injector and a reduction in the dispersion of the arcs, both of which increase the energy acceptance of CEBAF. We expect to commence initial tests of these adaptations on CEBAF during 2021.

Perez-Segurana, G.↗

ExaSGD: 2022 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-407 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY22 the main objective of the Kernel Thrust was (i) provide sparse optimization solver that runs efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations, (ii) strengthen the reliability and increase the performance of the mixed-dense sparse (MDS) solver of HiOp for deployment on the FY22 target architectures, Summit and Crusher, and (iii) increase performance by improving the mathematical algorithm and refining the parallel MPI-based implementation of the coarse-grain parallel solver HiOp-PriDec for capabilities deployment on the FY22 target architectures, Summit and Crusher. This document presents the developments and contributions done by the Kernels Thrust Team in FY22 toward completion of the above-mentioned objectives. These contributions progressed along four main development (sub)thrusts: (1) Design and implementation of a sparse optimization solver for use on hardware accelerators; (2) Improvement of the mathematical algorithm and of the parallel implementation of HiOp-PriDec to ensure readiness and efficient coarse-grain parallelism for FY23 target exascale machine; and (3) Support Software and Application Development Thrusts of the exaSGD project in their deployment of the project’s software stack on AMD- and NVIDIA-based architectures. The development of the sparse optimization solver (thrust 1 above) was new in FY22 and resulted in a new sparse solver in HiOp (available as of version 0.6). The second development thrust was a continuation of the efforts from FY21 and improved the mathematical algorithm and the communication strategy of the HiOp-PriDec solver. The last developement thrust is a large collaborative effort. Namely, the project’s teams from multiple labs (LLNL, PNNL, ORNL, and NREL) performed large-scale demonstration of the ExaSGD software stack, namely the optimization solvers of HiOp interfaced with the modeling front-end ExaGO and the stochastic sampler PowerScenarios. These demonstration efforts solved large-scale instances of the SC-ACOPF challenge problem of medium network sizes (10, 000-bus system) and large number of contingencies on Summit (NVIDIA accelerators) and Crusher (AMD accelerators) systems at ORNL.

97 MATHEMATICS AND COMPUTING↗

Status report on HFIR irradiation of optimized alumina forming alloys

Properties of FeCrAl alloys under neutron irradiation are of interest because of these materials’ potential application as accident-tolerant fuel cladding in nuclear systems. In parallel, alumina-forming austenitic (AFA) alloys are of interest for use as structural materials in advanced nuclear systems for their potential higher resistance to embrittlement and high-temperature steam oxidation resistance. An irradiation campaign for fiscal year 2024 has been developed under the Advanced Fuels Campaign to perform irradiation testing of various FeCrAl and AFA alloys in Oak Ridge National Laboratory’s High Flux Isotope Reactor (HFIR). The goals of this irradiation campaign are to (1) study the impact of minor alloying elements on the neutron-irradiated mechanical properties of FeCrAl alloys and (2) collect neutron-irradiated mechanical properties on AFA alloys for comparison with those of FeCrAl alloys. This campaign will include both tensile and fracture toughness specimens tested following HFIR irradiation at temperatures representative of normal operating conditions in light-water reactors. The pre-irradiation characterization to date, the irradiation plan for the FeCrAl and AFA specimens, and the subsequent post-irradiation experimental test plan are presented in this report, along with the status of HFIR builds and scheduled insertion dates.

36 MATERIALS SCIENCE↗

From Raw to Curated Data: A Lakehouse Approach for Scientific Workflows

This report provides a technical overview of how to go from raw to curated data in three stages using a lakehouse approach. We focus on the application of open source tools in scientific use cases (while noting parallels to enterprise and commercial alternatives). Our goal is to provide scientific data managers and infrastructure providers with a common frame of reference for understanding and applying modern lakehouse technologies and approaches.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

BOXKIT

SF-23-067 BoxKit is a library that provides building blocks to parallelize and scale data science, high performance computing, and machine learning applications for block-structured datasets. Spatial data from simulations and experiments can be accessed and managed using tools available in this library when working with more data analysis oriented packages like SciKit (https://github.com/scikit-learn/scikit-learn) and FlowNet (https://github.com/NVIDIA/flownet2-pytorch)

DHRUV, AKASH↗

Extended Physics-Informed Neural Networks (XPINNs): A Generalized Space-Time Domain Decomposition Based Deep Learning Framework for Nonlinear Partial Differential Equations

Here we propose a generalized space-time domain decomposition approach for the physics-informed neural networks (PINNs) to solve nonlinear partial differential equations (PDEs) on arbitrary complex-geometry domains. The proposed framework, named eXtended PINNs ( X P I N N s ), further pushes the boundaries of both PINNs as well as conservative PINNs (cPINNs), which is a recently proposed domain decomposition approach in the PINN framework tailored to conservation laws. Compared to PINN, the XPINN method has large representation and parallelization capacity due to the inherent property of deployment of multiple neural networks in the smaller subdomains. Unlike cPINN, XPINN can be extended to any type of PDEs. Moreover, the domain can be decomposed in any arbitrary way (in space and time), which is not possible in cPINN. Thus, XPINN offers both space and time parallelization, thereby reducing the training cost more effectively. In each subdomain, a separate neural network is employed with optimally selected hyperparameters, e.g., depth/width of the network, number and location of residual points, activation function, optimization method, etc. A deep network can be employed in a subdomain with complex solution, whereas a shallow neural network can be used in a subdomain with relatively simple and smooth solutions. We demonstrate the versatility of XPINN by solving both forward and inverse PDE problems, ranging from one-dimensional to three-dimensional problems, from time-dependent to time-independent problems, and from continuous to discontinuous problems, which clearly shows that the XPINN method is promising in many practical problems. The proposed XPINN method is the generalization of PINN and cPINN methods, both in terms of applicability as well as domain decomposition approach, which efficiently lends itself to parallelized computation. The XPINN code is available on h t t p s : / / g i t h u b . c o m / A m e y a J a g t a p / X P I N N s .

97 MATHEMATICS AND COMPUTING↗

Enhanced accuracy through ensembling of randomly initialized auto-regressive models for dynamical systems

Computational mechanics simulations using traditional finite element methods (FEM) require prohibitively expensive computational resources for real-time engineering applications, design optimization, and digital twin implementations. While machine learning (ML) surrogate models offer significant computational speedups, autoregressive ML models for time-dependent mechanical systems suffer from error accumulation that compromises long-term prediction reliability - a critical concern for engineering applications where accuracy over extended time horizons is essential for safety and performance assessments. Here, we propose a deep ensemble framework specifically designed to address this challenge in computational mechanics applications, where multiple ML surrogate models with random weight initializations are trained in parallel and their predictions aggregated during inference. This approach leverages statistical diversity to maximize information gain from a fixed set of training data and to mitigate error propagation, while maintaining the computational efficiency that makes ML surrogates attractive for engineering practice. We validate the framework on three representative problems spanning critical areas of computational mechanics: stress field evolution in heterogeneous microstructures under complex loading (relevant to advanced materials design and composite analysis), planetary-scale shallow water dynamics (applicable to environmental and geotechnical engineering), and Gray-Scott reaction-diffusion systems (relevant to mass transport and chemical process engineering). Across all test cases, the ensemble approach demonstrates consistent error reduction of 15-33% compared to individual models. The codes for this work are available on GitHub (https://github.com/Graham-Brady-Research-Group/AutoregressiveEnsemble_SpatioTemporal_Evolution).

autoregressive prediction↗

Packaging a 650V/400A GaN Half-bridge Power Module with Ultra-low Parasitics for Electric Vehicle Drive Applications

This paper proposes a compact and efficient half-bridge power module with three 650 V / 150 A GaN dies in parallel. The power module incorporates a main power printed circuit board (PCB), an interface PCB, and a flex PCB to achieve low parasitics in both power loop and gate-side connection, resolving the issue of high parasitics typically encountered with wire bonding in high-current applications. Additionally, the interface PCB decouples the design constraints between the power loop and the gate loops. The proposed design is optimized with a vertical loop configuration to reduce power loop inductance through magnetic flux cancellation. Finite element analysis indicates that the power loop inductance is 0.58 nH at 100 MHz, while the maximum die junction temperature reaches 131 °C under an ambient temperature of 65 °C and a load current of 385 A. The proposed multi-piece PCB structure reduces the inductance of the drive circuit to minimize EMI and to mitigate false triggering. At the same time, it reduces impedance mismatches across different driver circuits, thereby achieving dynamic current sharing in multi-chip parallel configurations. Under simulation conditions of 400 V / 385 A, the current imbalance among chips was limited to 5 A. A 400 V / 385 A double-pulse test was conducted to experimentally validate the performance of the proposed power module.

30 DIRECT ENERGY CONVERSION↗

Scalability of high-performance PDE solvers

Performance tests and analyses are critical to effective high-performance computing software development and are central components in the design and implementation of computational algorithms for achieving faster simulations on existing and future computing architectures for large-scale application problems. In this article, we explore performance and space-time trade-offs for important compute-intensive kernels of large-scale numerical solvers for partial differential equations (PDEs) that govern a wide range of physical applications. We consider a sequence of PDE-motivated bake-off problems designed to establish best practices for efficient high-order simulations across a variety of codes and platforms. We measure peak performance (degrees of freedom per second) on a fixed number of nodes and identify effective code optimization strategies for each architecture. In addition to peak performance, we identify the minimum time to solution at 80% parallel efficiency. The performance analysis is based on spectral and p-type finite elements but is equally applicable to a broad spectrum of numerical PDE discretizations, including finite difference, finite volume, and h-type finite elements.

97 MATHEMATICS AND COMPUTING↗

High-Fidelity Ion State Detection Using Trap-Integrated Avalanche Photodiodes

Integrated technologies greatly enhance the prospects for practical quantum information processing and sensing devices based on trapped ions. High-speed and high-fidelity ion state readout is critical for any such application. Integrated detectors offer significant advantages for system portability and can also greatly facilitate parallel operations if a separate detector can be incorporated at each ion-trapping location. Here, we demonstrate ion quantum state detection at room temperature utilizing single-photon avalanche diodes (SPADs) integrated directly into the substrate of silicon ion trapping chips. Furthermore, we detect the state of a trapped Sr + ion via fluorescence collection with the SPAD, achieving 99.92(1)% average fidelity in 450 μs, opening the door to the application of integrated state detection to quantum computing and sensing utilizing arrays of trapped ions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Spectral quadrature for the first principles study of crystal defects: Application to magnesium

In this work, we present an accurate and efficient finite-difference formulation and parallel implementation of Kohn-Sham Density (Operator) Functional Theory (DFT) for non periodic systems embedded in a bulk environment. Specifically, employing non-local pseudopotentials, local reformulation of electrostatics, and truncation of the spatial Kohn-Sham Hamiltonian, and the Linear Scaling Spectral Quadrature method to solve for the pointwise electronic fields in real-space and the non-local component of the atomic force, we develop a parallel finite difference framework suitable for distributed memory computing architectures to simulate non-periodic systems embedded in a bulk environment. Choosing examples from magnesium-aluminum alloys, we first demonstrate the convergence of energies and forces with respect to spectral quadrature polynomial order, and the width of the spatially truncated Hamiltonian. Next, we demonstrate the parallel scaling of our framework, and show that the computation time and memory scale linearly with respect to the number of atoms. Next, we use the developed framework to simulate isolated point defects and their interactions in magnesium-aluminum alloys. Our findings conclude that the binding energies of divacancies, Al solute-vacancy and two Al solute atoms are anisotropic and are dependent on cell size. Furthermore, the binding is favorable in all three cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

SPACE: 3D parallel solvers for Vlasov-Maxwell and Vlasov-Poisson equations for relativistic plasmas with atomic transformations

A parallel, relativistic, three-dimensional particle-in-cell code SPACE has been developed for the simulation of electromagnetic fields, relativistic particle beams, and plasmas. In addition to the standard second-order Particle-in-Cell (PIC) algorithm, SPACE includes efficient novel algorithms to resolve atomic physics processes such as multi-level ionization of plasma atoms, recombination, and electron attachment to dopants in dense neutral gases. SPACE also contains a highly adaptive particle-based method, called Adaptive Particle-in-Cloud (AP-Cloud), for solving the Vlasov-Poisson problems. It eliminates the traditional Cartesian mesh of PIC and replaces it with an adaptive octree data structure. The code's algorithms, structure, capabilities, parallelization strategy, and performance have been discussed. Additionally, typical examples of SPACE applications to accelerator science and engineering problems are described.

43 PARTICLE ACCELERATORS↗

Optimizing Error-Bounded Lossy Compression for Scientific Data on GPUs

Error-bounded lossy compression is a critical technique for significantly reducing scientific data volumes. With ever-emerging heterogeneous high-performance computing (HPC) architecture, GPU-accelerated error-bounded compressors (such as CUSZ and cuZFP) have been developed. However, they suffer from either low performance or low compression ratios. To this end, we propose CUSZ+ to target both high compression ratios and throughputs. We identify that data sparsity and data smoothness are key factors for high compression throughputs. Our key contributions in this work are fourfold: (1) We propose an efficient compression workflow to adaptively perform run-length encoding and/or variable-length encoding. (2) We derive Lorenzo reconstruction in decompression as multidimensional partial-sum computation and propose a fine-grained Lorenzo reconstruction algorithm for GPU architectures. (3) We carefully optimize each of CUSZ kernels by leveraging state-of-the-art CUDA parallel primitives. (4) We evaluate CUSZ+ using seven real-world HPC application datasets on V100 and A100 GPUs. Experiments show CUSZ+ improves the compression throughputs and ratios by up to 18.4x and 5.3x, respectively, over CUSZ on the tested datasets.

Tian, Jiannan↗

Manipulation of Scattering Spectra with Topology of Light and Matter

Structured lights, including beams carrying spin and orbital angular momenta, radially and azimuthally polarized vector beams, as well as spatiotemporal optical vortices, have attracted significant interest due to their unique amplitude, phase front, polarization, and temporal structures, enabling a variety of applications in optical and quantum communications, micromanipulation, and super-resolution imaging. In parallel, structured optical materials, metamaterials, and metasurfaces consisting of engineered unit cells—meta-atoms, opened new avenues for manipulating the flow of light and optical sensing. While several studies explored structured light effects on the individual meta-atoms, their shapes are largely limited to simple spherical geometries. However, the synergy of the structured light and complex-shaped meta-atoms has not been fully explored. Here, in this paper, the role of the helical wavefront of Laguerre–Gaussian beams in the excitation and suppression of higher-order resonant modes inside all-dielectric meta-atoms of various shapes, aspect ratios, and orientations, is demonstrated and the excitation of various multipolar moments that are not accessible via unstructured light illumination is predicted. The presented study elucidates the role of the complex phase distribution of the incident light in shape-dependent resonant scattering, which is of utmost importance in a wide spectrum of applications ranging from remote sensing to spectroscopy.

42 ENGINEERING↗

Parallelized multiple nozzle system and method to produce layered droplets and fibers for microencapsulation

The present disclosure relates to a nozzle system for use in a microfluidic production application for producing at least one of particles, capsules or fibers. The system has a main body portion having a compressed fluid inlet and a core fluid inlet, and a plurality of parallel arranged core fluid nozzles that receive the core fluid and create a plurality of core fluid streams. At least one compressed fluid inlet associated with the main body channels compressed fluid to areas adjacent ends of the core fluid nozzles. An apertured plate having a plurality of apertures is arranged near the ends of the core fluid nozzles, with each aperture being uniquely associated with a single one of the core fluid nozzles. The compressed fluid acts on the core fluid streams exiting the core fluid nozzles to help create, with the apertures, at least one of core fluid droplets or core fluid fibers from the core fluid streams.

Ye, Congwang↗

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

Scalable line and plane relaxation in a parallel structured multigrid solver

The efficient solution of sparse, linear systems that arise through the discretization of partial differential equations remains a key challenge for a range of high performance scientific simulations. One approach for reducing data movement and improving performance is by exposing and exploiting structure in a problem through the use of robust structured multilevel solvers. By choosing coarsening that preserves the structure of the problem, these methods maintain efficient structured computation and communication throughout the multigrid hierarchy. However, when coarsening is not permitted to be dependent on the operator, anisotropy must be addressed by the smoother — producing error compatible for coarse-grid correction with structured coarsening. Here, the components required in a scalable parallel structured solver are described with a focus on memory and communication efficiency of robust smoothers. While the implementation of communication and memory reduction techniques in smoothers integrated in a complete 3D solver present a significant engineering challenge, a novel approach is proposed that addresses these challenges systematically through a change to the solver’s execution model. Enabled by user-level threading paired with a set of data and communication abstractions, this approach permits seamless aggregation of communication in plane smoothers — directly reusing code for a 2D distributed multilevel cycle. Results show an effective reduction in communication costs for coarse-grid problems, and result in a speedup of 8.7x in smoothing routines shown in Fig. 12 using this approach. This produces a significant improvement to strong scalability while maintaining favorable weak scaling behavior. Finally, a parallel scaling study using a series of refined meshes is included that demonstrates the effectiveness of this approach in an application of interest.

97 MATHEMATICS AND COMPUTING↗

Capacity Investment under Bayesian Information Updates at Reporting Periods: Model and Application

We consider capacity addition decisions by a new product manufacturer faced with uncertain technology alternatives. The manufacturing capacity addition and technology development occurs in parallel, with preliminary results from a technology project's success providing valuable information to the manufacturer in adding capacity. We solve a stochastic dynamic program with Bayesian updates to obtain the manufacturer's expected profit‐maximizing capacity investment decision. Our model and applications are motivated by the Critical Materials Institute (CMI) (funded by the Department of Energy), which manages research projects focused on mitigating critical material constraints, vital to renewable energy technologies such as direct‐drive wind turbines, electric vehicles, and energy‐efficient lighting. We capture three unique aspects of the problem: first, the learning from project progress depends on task‐based stochastic outcomes and a project's percent‐done. Second, the underlying technology's profitability is based on a model of competition. Third, we evaluate the impact of progress across a portfolio of projects based on a manufacturer's capacity addition. We develop a heuristic that produces results that are close to optimal and can thus be used for large problem sizes. The managerial insights from an application of our model to CMI projects include: (i) technology projects that report the percent‐done of a project earlier increase expected manufacturer profit; (ii) careful choice of “safe bets,” that is, technologies with low profitability but a high probability of success, can increase expected manufacturer profit; (iii) a portfolio of projects can increase profits significantly over separate project evaluation; and (iv) dynamic management of project resources can increase overall profit.

Vedantam, Aditya↗