Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

IDAES-PSE 2.7.0 Release

The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.7.0 Release Highlights New features: AutoScaler and CustomScalerBase classes: Such tools are the core of the new scaling framework being implemented in IDAES. Wider adoption of scaling tools among users will result in quicker and more robust model solutions. Scaler for equilibrium reactor and saponification properties: These scaler models are examples to follow for how to use the new scaling tools. ONNX Surrogate support from Optimization & Machine Learning Toolkit (OMLT): ONNX is an open standard format to save and load ML/AI models that is widely supported by all major frameworks. This capability makes it easier for IDAES users to create surrogate models and use them without having to support each framework individually. 1D Membrane Model for CO2 Capture and Utilization: Supports ongoing efforts for modeling and optimizing polymer membrane processes for CO2 capture and conversion into formic acid. StreamScaler unit model: Unrelated to the CustomScalerBase, this unit model allows a stream’s extensive variables to be scaled by a fixed factor. This allows streams being processed by multiple units in parallel to be scaled down to unit scale and scaled back up to process scale. Bug fixes or improvements: Scaling, EoS, Diagnostics tool, Modular Properties, tests & documentation Deprecations: Old Cubic EoS

AS↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Supercomputing Pipelines Search for Therapeutics Against COVID-19

The urgent search for drugs to combat SARS-CoV-2 has included the use of supercomputers. The use of general-purpose graphical processing units (GPUs), massive parallelism, and new software for high-performance computing (HPC) has allowed researchers to search the vast chemical space of potential drugs faster than ever before. We developed a new drug discovery pipeline using the Summit supercomputer at Oak Ridge National Laboratory to help pioneer this effort, with new platforms that incorporate GPU-accelerated simulation and allow for the virtual screening of billions of potential drug compounds in days compared to weeks or months for their ability to inhibit SARS-COV-2 proteins. Here, this effort will accelerate the process of developing drugs to combat the current COVID-19 pandemic and other diseases.

60 APPLIED LIFE SCIENCES↗

Characterization and Rationalization of Microstructural Evolution in GRCop-84 Processed by Laser-Powder Bed Fusion (L-PBF)

In this study, prismatic geometries of GRCop-84 [Cu-8Cr-4Nb (at. pct)] were built with laser-powder bed fusion (L-PBF) process. The samples were sectioned parallel or perpendicular to the build direction and characterized in the as-built and after post-processing with a hot-isostatically pressing (HIP) treatment. The microstructure and phase evolutions were evaluated with optical microscopy, scanning electron microscopy (SEM), electron backscattered diffraction (EBSD), and high-temperature X-ray diffraction (HTXRD) up to 1223 K. The samples in the as-built conditions exhibited predominantly columnar epitaxial and misoriented Cu-FCC grains. The microstructure evolutions are discussed based on locations within the overall build geometry, the dynamics of small melt pool shape and sectioning effects. The above grain structure did not change significantly during post-process HIP treatment. The stability of this FCC grain structure is attributed to the formation of primary stable Cr 2 Nb (Laves phase) during L-PBF, even before the emergence of FCC grains from liquid. The stability of Cr 2 Nb in both as-built and HIPed samples were evaluated using high-temperature X-ray diffraction measurements and compared with that of gas-atomized powder. The significance of these results is discussed with reference to aerospace applications.

36 MATERIALS SCIENCE↗

An open source fast fluid dynamics model for data center thermal management

Although computational fluid dynamics (CFD) has been widely adopted to improve data center thermal management, the high computational demand limits its applications, such as multivariate optimal design and operation. Fast fluid dynamics (FFD), which has been applied for fast airflow simulation, shows great potential. However, few research applied FFD for optimal design and operation of data center thermal management. This research improves the FFD model for data centers and conducts a comprehensive evaluation and demonstration. First, the FFD model is improved by solving the advection and diffusion equations together using an upwind scheme instead of a semi-Lagrangian advection solver in the conventional FFD model. Second, new features for data centers are added, such as a pressure correction method to simulate plenum airflow and dynamic boundary conditions for IT racks. The new FFD model is first validated with two indoor environment cases and the results show that the new FFD model has slightly better overall prediction accuracy and faster speed compared to the conventional FFD model. It is also observed that both FFD models achieve acceptable accuracy, except for a few localized disparities with experimental data, which might be due to simplified handling of turbulence viscosity near the boundaries. Furthermore, validation with a real data center shows that the FFD model achieves a similar level of accuracy as CFD when compared to the experimental measurements with some level of uncertainties. It is then demonstrated for data center optimal design and operation, which saves 53.4–58.8% of annual energy while still meeting the thermal requirements. In conclusion, with a much faster speed and comparable accuracy compared to CFD, the FFD model parallelized on a graphics processing unit is promising for practical model-based data center early design and operation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nanoscale Probing of Electrical Memory Effects in van der Waals Layered PdSe 2

Tunable electronic materials that can be switched between different impedance states are fundamental to the hardware elements for neuromorphic computing architectures. This “brain-like” computing paradigm uses highly paralleled and colocated data processing, leading to greatly improved energy efficiency and performance compared to traditional architectures in which data have to be frequently transferred between processor and memory. In this work, we use scanning microwave impedance microscopy for nanoscale electrical and electronic characterization of two-dimensional layered semiconductor PdSe 2 to probe neuromorphic properties. The local resolution of tens of nanometers reveals significant differences in electronic behavior between and within PdSe 2 nanosheets (NSs). In particular, we detected both n-type and p-type behaviors, although previous reports only point to ambipolar n-type dominating characteristics. Nanoscale capacitance–voltage curves and subsequent calculation of characteristic maps revealed a hysteretic behavior originating from the creation and erasure of Se vacancies as well as the switching of defect charge states. In addition, stacks consisting of two NSs show enhanced resistive and capacitive switching, which is attributed to trapped charge carriers at the interfaces between the stacked NSs. Stacking n- and p-type NSs results in a combined behavior that allows one to tune electrical characteristics. In conclusion, as local inhomogeneities of electrical and electronic behavior can have a significant impact on the overall device performance, the demonstrated nanoscale characterization and analysis will be applicable to a wide range of semiconducting materials.

2D material↗

Formation of late-generation atmospheric compounds inhibited by rapid deposition

Reactive organic carbon species are important fuel for atmospheric chemical reactions, including the formation of secondary organic aerosol. However, in parallel to atmospheric oxidation processes, deposition can remove compounds from the atmosphere and impact downstream environments. To understand the impact of deposition on atmospheric oxidation, we present a framework for predicting and visualizing the fate of a molecule on the basis of the physicochemical properties of compounds (Henry’s law constant, vapour pressure and reaction rate constants), which are used to estimate timescales for oxidation and deposition. Further, by implementing our deposition rates in chemical models, we show that deposition substantially suppresses atmospheric reactivity and aerosol formation by removing early-generation products and preventing the formation of large fractions (up to 90%) of downstream, late-generation compounds. Deposition is frequently missing in the laboratory experiments and detailed chemical modelling, which probably biases our understanding of atmospheric composition.

54 ENVIRONMENTAL SCIENCES↗

Laboratory study of the PFRC-2's initial plasma densification stages

Initial plasma densification by odd-parity rotating magnetic fields (RMF o ) applied to the linear magnetized Princeton field-reversed configuration (PFRC-2) device with fill gases at pressures near 1 mTorr proceeds through two phases: a slow one, characterized by a rise time $τ_{slow}$ ~ 100 $μ$s, followed by a fast one, characterized by $τ_{fast}$ ~ 10 $μ$s. The transition from slow to fast occurs at a line-integral-averaged electron density, t n e , near 2$\times$ 10 11 cm –3 , independent of magnetic field. Here, over most of the range of experimental parameters investigated, as the PFRC-2 axial magnetic field strength was increased, RMF o power decreased, gas fill pressure lowered, or lower atomic mass unit (AMU) fill gas used, the duration of the slow phase lengthened from 50 $μ$s to longer than 10 ms after the RMF o power began. The post-fast-phase maximum n e increases with the fill-gas AMU, exceeding 5 × 10 13 cm –3 for Ar. The slow phase is consistent with atomic physics processes and field-parallel sound-speed losses. The fast phase may be explained by improved axial confinement, possibly augmented by radial or axial contraction of the plasma. Another possible explanation, a large increase in electron temperature, is inconsistent with x-ray emission. The n e behavior is discussed in relation to the E to H transition.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Understanding detachment of the W7-X island divertor

The fundamental behavior of the W7-X island divertor under detached conditions, which has been theoretically predicted with the EMC3-Eirene code, is re-examined here under the experimental conditions achieved so far and compared with the first experimental results. Both simulations and experiments cover a range of divertor configurations and plasma parameters, and show the following common trends: (1) with rising impurity radiation, the target heat load decreases 'uniformly' over the entire target surface in the sense that both the peak and average heat loads can drop by an order of magnitude. Impurity radiation (mainly from intrinsic carbon) occurs primarily at the plasma edge and the resulting negative impact on the stored energy is less than 10%. (2) When the total radiation exceeds a critical level, the target particle flux (the recycling flux Γrecy) begins to fall and can drop by a factor of 3–5 at high radiation levels without an obvious indication of significant volume recombination. (3) While Γ recy decreases, the divertor neutral pressure continues to build up and reaches a maximum, at which point Γ recy has declined significantly. (4) During detachment, the electron temperature at the last closed flux surface falls in a way that is not quantitatively understandable from parallel classical heat conduction processes. This paper presents a physical explanation of the numerical/experimental results described above. Furthermore, using the EMC3-Eirene code as a diagnostic tool, we are able, apparently for the first time, to provide a full quantitative analysis of each transport channel in the island divertor, aiming to clarify how the island divertor plasma self-regulates to maintain particle, energy, and momentum balance under detached conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Solar Cell Metallization Wear is Sensitive to Loading Frequency

Here, we performed cyclic loading of photovoltaic laminates with precracked silicon cells to explore if and how loading frequency and contact pressure influence the ensuing gridline wear-out process. A measurement of parallel resistance across cracked gridlines on a laminated cell coupon was used as the metric for gridline electrical contact degradation. A statistical analysis of variance (ANOVA) analysis of the experimental results discerned that loading frequency is a more significant factor than contact pressure for gridline degradation.

14 SOLAR ENERGY↗

A Two-Stage Decomposition Approach for AC Optimal Power Flow

The alternating current optimal power flow (AC-OPF) problem is critical to power system operations and planning, but it is generally hard to solve due to its nonconvex and large-scale nature. Furthermore, this paper proposes a scalable decomposition approach in which the power network is decomposed into a master network and a number of subnetworks, where each network has its own AC-OPF subproblem. This formulates a two-stage optimization problem and requires only a small amount of communication between the master and subnetworks. The key contribution is a smoothing technique that renders the response of a subnetwork differentiable with respect to the input from the master problem, utilizing properties of the barrier problem formulation that naturally arises when subproblems are solved by a primal-dual interior-point algorithm. Consequently, existing efficient nonlinear programming solvers can be used for both the master problem and the subproblems. The advantage of this framework is that speedup can be obtained by processing the subnetworks in parallel, and it has convergence guarantees under reasonable assumptions. The formulation is readily extended to instances with stochastic subnetwork loads. Numerical results show favorable performance and illustrate the scalability of the algorithm which is able to solve instances with more than 11 million buses.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High energy density picoliter-scale zinc-air microbatteries for colloidal robotics

The recent interest in microscopic autonomous systems, including microrobots, colloidal state machines, and smart dust, has created a need for microscale energy storage and harvesting. However, macroscopic materials for energy storage have noted incompatibilities with microfabrication techniques, creating substantial challenges to realizing microscale energy systems. Here, we photolithographically patterned a microscale zinc/platinum/SU-8 system to generate the highest energy density microbattery at the picoliter (10 −12 liter) scale. The device scavenges ambient or solution-dissolved oxygen for a zinc oxidation reaction, achieving an energy density ranging from 760 to 1070 watt-hours per liter at scales below 100 micrometers lateral and 2 micrometers thickness in size. The parallel nature of photolithography processes allows 10,000 devices per wafer to be released into solution as colloids with energy stored on board. Within a volume of only 2 picoliters each, these primary microbatteries can deliver open circuit voltages of 1.05 ± 0.12 volts, with total energies ranging from 5.5 ± 0.3 to 7.7 ± 1.0 microjoules and a maximum power near 2.7 nanowatts. We demonstrated that such systems can reliably power a micrometer-sized memristor circuit, providing access to nonvolatile memory. We also cycled power to drive the reversible bending of microscale bimorph actuators at 0.05 hertz for mechanical functions of colloidal robots. Additional capabilities, such as powering two distinct nanosensor types and a clock circuit, were also demonstrated. The high energy density, low volume, and simple configuration promise the mass fabrication and adoption of such picoliter zinc-air batteries for micrometer-scale, colloidal robotics with autonomous functions.

Robotics↗

Cardinal: Seismic and Geoacoustic Array Processing

Data collected via seismic and infrasound array deployments are leveraged in the geosciences to detect and characterize a myriad of natural and anthropogenic sources. These deployments consist of numerous sensors placed in a predetermined configuration to amplify signal strength and improve the efficacy of array processing techniques used to measure signal directionality and waveform coherence. High‐fidelity feature extraction is often predicated on interstation distance as well as the frequency content and wavelength of an incident signal. Numerous array processing softwares analyze data in sequential frequency bands to obtain a more detailed characterization of a signal. However, current algorithms are limited in their ability to determine optimal array configuration for each band. We introduce an open‐source Python code, called Cardinal, to process seismic and infrasound array data in discretized time–frequency space with the option of applying an adaptive array design to determine optimal subarray configuration for each frequency band. To reduce computational time, the array processing step can be run in parallel using multithreading. Furthermore, the software has the capability to aggregate array processing results from different time–frequency pixels to produce separate sets of detections, or families, with added utility via the application of an adaptive semblance threshold, which aids in isolating signals‐of‐interest from coherent background noise. Upon appropriate configuration, Cardinal exhibits the potential to combine distinct seismic and infrasound phases into separate families.

Adaptive Array↗

Machine Learning-Driven Conservative-to-Primitive Conversion in Hybrid Piecewise Polytropic and Tabulated Equations of State

We present a novel machine learning (ML)-based method to accelerate conservative-to-primitive inversion, focusing on hybrid piecewise polytropic and tabulated equations of state. Traditional root-finding techniques are computationally expensive, particularly for large-scale relativistic hydrodynamics simulations. To address this, we employ feedforward neural networks (NNC2PS and NNC2PL), trained in PyTorch (2.0+) and optimized for GPU inference using NVIDIA TensorRT (8.4.1), achieving significant speedups with minimal accuracy loss. The NNC2PS model achieves 𝐿 1 and 𝐿 ∞ errors of 4.54 × 10 −7 and 3.44 × 10−6, respectively, while the NNC2PL model exhibits even lower error values. TensorRT optimization with mixed-precision deployment substantially accelerates performance compared to traditional root-finding methods. Specifically, the mixed-precision TensorRT engine for NNC2PS achieves inference speeds approximately 400 times faster than a traditional single-threaded CPU implementation for a dataset size of 1,000,000 points. Ideal parallelization across an entire compute node in the Delta supercomputer (dual AMD 64-core 2.45 GHz Milan processors and 8 NVIDIA A100 GPUs with 40 GB HBM2 RAM and NVLink) predicts a 25-fold speedup for TensorRT over an optimally parallelized numerical method when processing 8 million data points. Moreover, the ML method exhibits sub-linear scaling with increasing dataset sizes. We release the scientific software developed, enabling further validation and extension of our findings. By exploiting the underlying symmetries within the equation of state, these findings highlight the potential of ML, combined with GPU optimization and model quantization, to accelerate conservative-to-primitive inversion in relativistic hydrodynamics simulations.

conservative-to-primitive conversion↗

aphBO-2GP-3B: a budgeted asynchronous parallel multi-acquisition functions for constrained Bayesian optimization on high-performing computing architecture

High-fidelity complex engineering simulations are often predictive, but also computationally expensive and often require substantial computational efforts. The mitigation of computational burden is usually enabled through parallelism in high-performance cluster (HPC) architecture. Optimization problems associated with these applications is a challenging problem due to the high computational cost of the high-fidelity simulations. In this paper, an asynchronous parallel constrained Bayesian optimization method is proposed to efficiently solve the computationally expensive simulation-based optimization problems on the HPC platform, with a budgeted computational resource, where the maximum number of simulations is a constant. The advantage of this method are three-fold. Firstly, the efficiency of the Bayesian optimization is improved, where multiple input locations are evaluated parallel in an asynchronous manner to accelerate the optimization convergence with respect to physical runtime. This efficiency feature is further improved so that when each of the inputs is finished, another input is queried without waiting for the whole batch to complete. Second, the proposed method can handle both known and unknown constraints. Third, the proposed method samples several acquisition functions based on their rewards using a modified GP-Hedge scheme. The proposed framework is termed aphBO-2GP-3B, which means asynchronous parallel hedge Bayesian optimization with two Gaussian processes and three batches. The numerical performance of the proposed framework aphBO-2GP-3B is comprehensively benchmarked using 16 numerical examples, compared against other 6 parallel Bayesian optimization variants and 1 parallel Monte Carlo as a baseline, and demonstrated using two real-world high-fidelity expensive industrial applications. The first engineering application is based on finite element analysis (FEA) and the second one is based on computational fluid dynamics (CFD) simulations.

97 MATHEMATICS AND COMPUTING↗

Leveraging FPGA Advantages for Quicker Data Processing for LBNF

The Long Baseline Neutrino Facility (LBNF) will deliver a 2.4 MW muon neutrino beam from Fermilab to the Deep Underground Neutrino Experiment (DUNE), requiring unprecedented precision in beamline alignment to achieve DUNE's neutrino oscillation measurement goals. Vertical misalignments of beamline components as small as 0.5 mm can contribute 6-7\% uncertainty in predicted neutrino flux, necessitating sub-0.1 mm alignment monitoring capabilities. The Horn Location Sensor (HLS) system employs frequency sweep interferometry (FSI) in a distributed hydrostatic leveling network to achieve the required precision under harsh radiation conditions up to 5000 kRad/year. Traditional FSI implementations suffer from laser sweep nonlinearities that degrade resolution and require computationally intensive post-processing corrections using gas reference cells. This work presents a real-time FPGA-based implementation of the HLS data acquisition and processing system using a sweep tracker interferometer for dynamic sweep linearization. The system utilizes a PYNQ-Z2 FPGA with programmable logic implementing parallel 16k-point FFT processing across four channels, synchronized by the sweep tracker signal to eliminate post-processing requirements. Spectral performance testing demonstrates significant improvements in peak sharpness compared to traditional fixed-frequency digitization. The FPGA implementation enables real-time displacement monitoring with processing speeds orders of magnitude faster than software-based approaches, essential for the operational requirements of LBNF's eventual distributed sensor network. This advancement in real-time FSI processing directly supports DUNE's precision neutrino physics program by providing the rapid feedback necessary for maintaining stringent beamline alignment tolerances during high-power beam operations.

Rossel, Jacob↗

Real-Time FPGA Implementation For Frequency Sweep Interferometry In The LBNF Complex

The Long Baseline Neutrino Facility (LBNF) will deliver a 2.4 MW muon neutrino beam from Fermilab to the Deep Underground Neutrino Experiment (DUNE), requiring unprecedented precision in beamline alignment to achieve DUNE's neutrino oscillation measurement goals. Vertical misalignments of beamline components as small as 0.5 mm can contribute 6-7\% uncertainty in predicted neutrino flux, necessitating sub-0.1 mm alignment monitoring capabilities. The Horn Location Sensor (HLS) system employs frequency sweep interferometry (FSI) in a distributed hydrostatic leveling network to achieve the required precision under harsh radiation conditions up to 5000 kRad/year. Traditional FSI implementations suffer from laser sweep nonlinearities that degrade resolution and require computationally intensive post-processing corrections using gas reference cells. This work presents a real-time FPGA-based implementation of the HLS data acquisition and processing system using a sweep tracker interferometer for dynamic sweep linearization. The system utilizes a PYNQ-Z2 FPGA with programmable logic implementing parallel 16k-point FFT processing across four channels, synchronized by the sweep tracker signal to eliminate post-processing requirements. Spectral performance testing demonstrates significant improvements in peak sharpness compared to traditional fixed-frequency digitization. The FPGA implementation enables real-time displacement monitoring with processing speeds orders of magnitude faster than software-based approaches, essential for the operational requirements of LBNF's eventual distributed sensor network. This advancement in real-time FSI processing directly supports DUNE's precision neutrino physics program by providing the rapid feedback necessary for maintaining stringent beamline alignment tolerances during high-power beam operations.

Rossel, A. Jacob [Fermilab; Unlisted]↗

A Performance Portable, Fully Implicit Landau Collision Operator with Batched Linear Solvers

Modern accelerators use hierarchical parallel programming models that enable massive multithreading within a processing element (PE), with multiple PEs per device driven by traditional processes. Batching is a technique for exposing PE-level parallelism in algorithms that have traditionally run on MPI processes or multiple threads within a single process. Opportunities for batching arise in, for example, kinetic discretizations of magnetized plasmas where collisions are advanced in velocity space at each spatial point independently. This paper builds on previous work on a high-performance, fully nonlinear, Landau collision operator by batching the linear solver, as well as batching the spatial point problems and adding new support for multiple grids for multiscale, multispecies problems. An anisotropic relaxation verification test that agrees well with previously published results and analytical models is presented. The performance results from NVIDIA A100 and AMD MI250X nodes are presented with hardware utilization analysis for each architecture. Finally, the entire implicit Landau operator time advance is implemented in Kokkos for performance portability, running entirely on the device and is available in the PETSc numerical library.

97 MATHEMATICS AND COMPUTING↗