Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel projection algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Cold Plasma Measurements

We have continued the simulation campaign in support of our ongoing magnetospheric cold plasma research project. This project aims to develop the next-generation particle instruments to measure the properties of the cold particle populations in the Earth’s magnetosphere. For this purpose, simulations have been performed with a Particle-In-Cell (PIC) code called the Curvilinear PIC (CPIC). The code is formulated in curvilinear geometry and couples the standard PIC algorithm with algorithms for the generation and adaptation of the underlaying computational mesh. It conforms to complex objects like spacecraft and it can place more grid points in regions where higher resolution is needed. The code also features a scalable solver based on the multigrid algorithm and it is fully parallelized via domain decomposition and MPI.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Scalable and accurate multi-GPU-based image reconstruction of large-scale ptychography data

Abstract While the advances in synchrotron light sources, together with the development of focusing optics and detectors, allow nanoscale ptychographic imaging of materials and biological specimens, the corresponding experiments can yield terabyte-scale volumes of data that can impose a heavy burden on the computing platform. Although graphics processing units (GPUs) provide high performance for such large-scale ptychography datasets, a single GPU is typically insufficient for analysis and reconstruction. Several works have considered leveraging multiple GPUs to accelerate the ptychographic reconstruction. However, most of these works utilize only the Message Passing Interface to handle the communications between GPUs. This approach poses inefficiency for a hardware configuration that has multiple GPUs in a single node, especially while reconstructing a single large projection, since it provides no optimizations to handle the heterogeneous GPU interconnections containing both low-speed (e.g., PCIe) and high-speed links (e.g., NVLink). In this paper, we provide an optimized intranode multi-GPU implementation that can efficiently solve large-scale ptychographic reconstruction problems. We focus on the maximum likelihood reconstruction problem using a conjugate gradient (CG) method for the solution and propose a novel hybrid parallelization model to address the performance bottlenecks in the CG solver. Accordingly, we have developed a tool, called PtyGer ( Pty chographic G PU(multipl e )-based r econstruction), implementing our hybrid parallelization model design. A comprehensive evaluation verifies that PtyGer can fully preserve the original algorithm’s accuracy while achieving outstanding intranode GPU scalability.

97 MATHEMATICS AND COMPUTING↗

TEAM Project Review, Year 2

This report summarizes our research activities within the TEAM project between December 2020 and December 2021, funded by the ASCR Advanced Research in Quantum Computing program. During the reporting period the LLNL-MSU team has made progress on several fronts. An overarching goal of the team is to provide a comprehensive suite of software tools that can be used for the Characterize-Optimize-Compute loop needed to implement and execute algorithms on quantum devices. We are concurrently developing lightweight solvers that can be used on desktop computers to find optimal control pulses and to characterize small quantum systems (consisting of a few transmons and cavities). However, desktop computers are insufficient for simulating and characterizing larger quantum systems. We have therefore also developed parallel, distributed memory, simulators and optimization solvers, both for open and closed quantum systems. These parallel solvers have, for example, been used to study quantum optimal control for pure-state preparation, utilizing 1000’s of cores on a modern high-performance computing (HPC) platform.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scaling the SciDAC QuantOm Workflow

As part of the Scientific Discovery through Advanced Computing (SciDAC) program, the Quantum Chromodynamics Nuclear Tomography (QuantOM) project aims to analyze data from Deep Inelastic Scattering (DIS) experiments conducted at Jefferson Lab and the upcoming Electron Ion Collider. The DIS data analysis is performed on an event-level by combining the input from theoretical and experimental nuclear physics into a single, composable workflow. The optimization itself (I.e. fitting the experimental data with theoretical predictions) is carried out by a machine / deep learning algorithm. The size of the acquired DIS data as well as the complexity of the workflow itself require that the analysis is performed across multiple GPUs on high performance computing systems, such as Polaris at Argonne National Laboratory. This presentation discusses the novelties and challenges that came along with parallelizing this workflow. Recent results are compared to common distributed training techniques.

Lersch, Daniel↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Three-Dimensional Grid Visualization for Planning Activities: A Dubai Case Study

National Laboratory of the Rockies (NLR), in collaboration with the Dubai Electricity and Water Authority (DEWA) and Infra-X, has undertaken the Energy Visualization Analysis Project. The aim of this project is to enhance analytical and 3D visualization capabilities for distribution network planning and renewable energy integration. As modern grid continues to evolve with large-scale solar PV deployment and emerging distributed energy resources (DERs), the ability to effectively analyze, visualize, and communicate complex grid behaviors has become increasingly critical. The project focuses on developing empirical use cases based on real distribution feeder data and engineering workflows, ensuring the outcomes are directly aligned with operational environment. Through time-series power flow simulations and nodal hosting capacity analysis, the study quantifies the impacts of high PV penetration on voltage and thermal limits within representative 11 kV feeders. These analyses identify specific nodes and conditions where DER integration challenges arise. Furthermore, a Battery Energy Storage System (BESS) optimization algorithm was applied to determine the optimal size and placement of storage systems that can mitigate network constraints and enhance hosting capacity. The comparative results between base-case and BESS-augmented scenarios clearly demonstrate improvements in network stability and load management efficiency. In parallel, the NLR team developed an immersive 3D visualization framework, enabling interactive exploration of grid simulations using commodity head-mounted display (HMD) systems. This framework transforms conventional 2D simulation data into spatially intuitive visual environments - allowing engineers to analyze feeder conditions, PV hosting potential, and BESS effects in real time. This report represents the first foundational phase in establishing a visualization-driven analytical ecosystem. It provides a methodological foundation for data integration, visualization architecture, and simulation-based decision support, paving the way for large-scale adoption of immersive visualization across DEWA's Smart Grid Initiative, R&D activities, and future network resilience studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Projective Hedging Algorithms for Multistage Stochastic Programming, Supporting Distributed and Asynchronous Implementation

Here we propose a decomposition algorithm for multistage stochastic programming that resembles the progressive hedging method of Rockafellar and Wets but is provably capable of several forms of asynchronous operation. We derive the method from a class of projective operator splitting methods fairly recently proposed by Combettes and Eckstein, significantly expanding the known applications of those methods. Our derivation assures convergence for convex problems whose feasible set is compact, subject to some standard regularity conditions and a mild “fairness” condition on subproblem selection. The method’s convergence guarantees are deterministic and do not require randomization, in contrast to other proposed asynchronous variations of progressive hedging. Computational experiments described in an online appendix show the method to outperform progressive hedging on large-scale problems in a highly parallel computing environment.

97 MATHEMATICS AND COMPUTING↗

MFIX DEM Enhancement for Industry-Relevant Flows (Final Report)

The overall goal of this two-phase project is to implement performance improvements of the Multiphase Flow with Interphase Exchanges (MFIX) Discrete Element Model (DEM) code that enable a transformative shift for industrial use. Prior to this effort, the largest simulations performed using MFIX are O(10 7 ) particles. This falls short of the O(10 9 ) particle simulations that must be completed on a timescale of days or weeks (vs. months or years) to enable simulations with physically-relevant domain sizes to be incorporated into industrial design cycles within five years. This was accomplished by tailoring best-in-class practices to bear on the unique challenges posed by the MFIX-DEM algorithm and code base. Scientific simulations (e.g., in cosmology, turbulent combustion) routinely use massively parallel computing to update far more particles in short wall clock times. Results from Phase 1 (1.5 years in duration) indicated significant gains in speed were possible for a wide range of benchmark cases. Moreover, a survey sent to >35 companies indicates that the timing is ideal for such an enhanced tool, with >80% of the respondents indicating that DEM is already value-added or will be within the next 5 years, and >70% of the respondents indicating that improved speed is the top computational priority. In Phase 2 (3.5 years in duration), the two major barriers that hinder industry from effectively using multiphase Computational Fluid Dynamics (CFD) to cut costs and improve performance, namely computational overhead and confidence in predictions, continued to be addressed. Regarding the former, the results from Phase 1 to guide the effort, with enhancements focused on an improved time-stepping algorithm and particle sorting. Four target problems of 1 billion particles each and increasing complexity were identified: homogeneous cooling, tumbler with continuous particle size distribution, discharge from a rectangular hopper and a cylindrical riser. Each of these were successfully simulated for relevant time scales (on order of seconds) using less than 24 hours of wall clock time. These represent the first 1-billion particle DEM simulations performed with MFIX, namely using the MFIX-Exa code. This code is currently under development at NETL in collaboration with Lawrence Berkeley National Laboratory. Regarding the second barrier on predictive uncertainty, experiments from Phase 1 (interacting nozzles - hydrodynamics only) and Phase 2 (very small-scale segregation experiments) were used to demonstrate the ability of two simplified approaches to uncertainty quantification (UQ). By limiting the number of particles, UQ based on the simplified treatment was compared to standard UQ, which was shown to have much higher computational demands. Experiments were also performed on a pilot-scale stripper unit to provide validation data for future CFD-DEM simulations and UQ.

20 FOSSIL-FUELED POWER PLANTS↗

Scaled ILU Smoothers for Navier-Stokes Pressure Projection

Incomplete LU (ILU) smoothers are effective in the algebraic multigrid (AMG) V-cycle for reducing high-frequency components of the error. However, the requisite direct triangular solves are comparatively slow on GPUs. Previous work has demonstrated the advantages of Jacobi iteration as an alternative to direct solution of these systems. Depending on the threshold and fill-level parameters chosen, the factors can be highly nonnormal and Jacobi is unlikely to converge in a low number of iterations. We demonstrate that row scaling can reduce the departure from normality, allowing us to replace the inherently sequential solve with a rapidly converging Richardson iteration. There are several advantages beyond the lower compute time. Scaling is performed locally for a diagonal block of the global matrix because it is applied directly to the factor. Further, an ILUT Schur complement smoother maintains a constant GMRES iteration count as the number of MPI ranks increases, and thus parallel strong-scaling is improved. Our algorithms have been incorporated into hypre, and we demonstrate improved time to solution for linear systems arising in the Nalu-Wind and PeleLM pressure solvers. For large problem sizes, GMRES+AMG executes at least five times faster when using iterative triangular solves compared with direct solves on massively parallel GPUs.

algebraic multigrid↗

Broadband Characterization and Circuit Model Development of Transmission-Scale Transformers

This report describes broadband measurements of transmission-scale transformers typical in the electric power grid. This work was performed as part of the EMP Resilient Grid LDRD project at Sandia National Laboratories to generate circuit models that can be used for high-altitude electromagnetic pulse (HEMP) coupling simulations and response predictions. The objective of the work was to obtain characterization data of substation yard equipment across a frequency range relevant to HEMP. Vector network analyzer measurements up to 100 MHz were performed on two power transformers at ABB-Hitachi and a single ITEC potential transformer. Custom cable breakouts were designed to interface with the transformer terminals and provide ground connections to the chassis at the base of the transformer bushings. The three-phase terminals of the power transformers were measured as a common mode impedance using a parallel resistive splitter, and the single-phase terminals of the potential transformer were measured directly. A vector fitting algorithm was used to empirically fit circuit models to the resulting two-port networks and input impedances of the measured objects. Simplified circuit representations of the input impedances were also generated to assess the degree of precision needed for high-altitude electromagnetic pulse response predictions, which were performed in Sandia's XYCE circuit simulator platform. HEMP coupling simulations using the transformer models showed significant reduction in the voltage peak and broadening in the pulse width seen at the power transformer compared to the traveling wave voltage. This indicated the importance of the load condition when defining the coupled insult in an electric power substation. Simplified circuit models showed a similar voltage at the transformer with a smoothed waveform. The presence of potential transformers in the simulation did not significantly change the simulated voltage at the power transformer. Single-port input impedance models were also developed to define load conditions when transfer characteristics were not necessary.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Transport Analysis & Optimization in a MW-Scale CO2 Electrolyzer (Final Report)

As Twelve continues to scale up their CO2 electrolyzers, both in the size of a single cell and in the number of cells used in a stack, thermal management becomes a growing concern, since excess heat can affect reaction yield and accelerate degradation. In this project, we aim to computationally explore how the anode flow fields used in Twelve’s CO2 electrolyzers function as heat exchangers. In particular, using a homogenized model of a CO2 electrolyzer, we first estimate the amount of heat generated in a cell. Then, we develop a computational fluid dynamics (CFD) model of the so-called “flow field”, i.e. a flow manifold, based on Twelve’s CAD drawings, to evaluate how these flow fields perform as a heat exchanger for the generated heat. We explore both a single cell and a 3-cell stack operating in parallel, where heat generated in one cell can now be transferred to another cell. We evaluate how performance is affected when environmental heat losses are taken into account. Finally, we leverage topology optimization to explore the types of design features a computational optimization algorithm would suggest to supplement our intuition. Overall, our work aims to provide design recommendations for CO2 electrolyzer flow fields and provides a foundation for future studies of flow field optimization.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE↗

Porting Classical Approaches for Quantum Simulations to Quantum Computers

Simulating quantum many-body systems is one of the most promising problems in which we might anticipate that quantum computers should show quantum advantage. Unfortunately, there is still a gap between this promise and actual practice. New quantum algorithms need to be developed and the current quantum algorithms have various difficulties - e.g efficient state preparation - which must be overcome and improved upon. In many cases, classical approaches need to be ported over to quantum devices. In this project we have developed a suite of new quantum algorithms which makes progress in this regard. We developed a new optimization scheme for variational quantum eigensolvers, UBOS, which mitigates problems with local minimas and barren plateaus while improving convergence to the ground state by an order of magnitude. We developed a new way to utilize qubitization to find ground states of nearly frustration-free Hamiltonians faster than all previous methods. We developed a series of state preparation techniques which helps initialize parameterized quantum circuits into reasonable starting points on which quantum algorithms are then applied. In addition to the development of novel algorithms, it is critical to have classical simulation techniques for approximately simulating quantum circuits which can be used to benchmark and understand quantum algorithms. Toward that end, we developed a novel POVM formalism to simulate quantum circuits as well as exemplify the massive parallelization of tensor network methodologies. Finally, we developed physical understanding of entanglement phase transitions such as many-body localization and random tensor networks.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Qu8its for quantum simulations of lattice quantum chromodynamics

We explore the utility of d = 8 qudits, qu8its, for quantum simulations of the dynamics of 1+1⁢D SU(3) lattice quantum chromodynamics, including a mapping for arbitrary number of flavors and lattice size and a reorganization of the Hamiltonian for efficient time evolution. Recent advances in parallel gate applications, along with the shorter application times of single-qudit operations compared with two-qudit operations, lead to significant projected advantages in quantum simulation fidelities and circuit depths using qu8its rather than qubits. The number of two-qudit entangling gates required for time evolution using qu8its is found to be more than a factor of 5 fewer than for qubits. Here, we anticipate that the developments presented in this work will enable improved quantum simulations to be performed using emerging quantum hardware.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The DECADE cosmic shear project I: A new weak lensing shape catalog of 107 million galaxies

We present the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. This catalog was assembled from public DECam data including survey and standard observing programs. These data were consistently processed with the Dark Energy Survey Data Management pipeline as part of the DECADE campaign and serve as the basis of the DECam Local Volume Exploration survey (DELVE) Early Data Release 3 (EDR3). We apply the Metacalibration measurement algorithm to generate and calibrate galaxy shapes. After cuts, the resulting cosmology-ready galaxy shape catalog covers a region of $5,\!412 \,\,{\rm deg}^2$ with an effective number density of $4.59\,\, {\rm arcmin}^{-2}$. The coadd images used to derive this data have a median limiting magnitude of $r = 23.6$, $i = 23.2$, and $z = 22.6$, estimated at ${\rm S/N} = 10$ in a 2 arcsecond aperture. We present a suite of detailed studies to characterize the catalog, measure any residual systematic biases, and verify that the catalog is suitable for cosmology analyses. In parallel, we build an image simulation pipeline to characterize the remaining multiplicative shear bias in this catalog, which we measure to be $m = (-2.454 \pm 0.124) \times10^{-2}$ for the full sample. Despite the significantly inhomogeneous nature of the data set, due to it being an amalgamation of various observing programs, we find the resulting catalog has sufficient quality to yield competitive cosmological constraints.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Cybersecurity for Grid Connected eXtreme Fast Charging (XFC) Station (CyberX) (Final Scientific/Technical Report)

This report summarizes the activities conducted under the DOE VTO funded project DE- EE0008451, where ABB Inc. (ABB), in collaboration with Idaho National Laboratory (INL), APS Global (APS), and XOS Trucks (XOS) pursued the development of a cyber-resilient extreme fast charging (XFC) management system. This project entitled Cybersecurity for Grid Connected eXtreme Fast Charging (XFC) Station (CyberX) focuses on a resilient architecture for smart charging EV Supply Equipment (EVSE) device control and Coordinated Anomaly Detection System (CADS) features that can be added at the charging site depot level to increase cybersecurity. The project was split into two budget periods focused first on developing the threat model and resilient control concepts and second on testing, improving, and validating those developed resilient control algorithms and features with a focus on key vulnerabilities identified during the threat assessment portion of the project. During the first budget period of the CyberX project, the ABB led team focused on activities to identify, model, and quantitatively prioritize high-impact attack scenarios with potential cyber-physical effects while also modeling and developing concepts for a resilient control system that could securely address integration of DERs and other resources with EV charging. Development of the security focused XFC management system (XMS) was accomplished first by offline simulation using a developed XFC station or depot with 480V input level and simulating measurement inputs to monitoring and control systems in concept development. A representative distribution grid model was developed supporting an EV charging site model with BESS and 6 general EV charging models. These EV charging models allowed multiple configurations of charging level, multiple connected protection and measurement devices, and simulation function to show general compromise of EV, BESS, and protection features based on parallel threat analysis. During the second budget period, the EV site and supporting systems model was developed in more detail and converted from offline model to real-time to real-time with EV charging hardware in the loop (HIL). The resilient control architecture developed as concept in the first part of the project was further tested and validated for integration of local energy resources and XFC charging station site equipment while maintaining cybersecure operating principles. The proposed resilient architecture for smart charging and cybersecurity features consists of two main concepts developed and tested within the project. The first concept is an XFC management system (XMS) consisting of a hardware gateway, software platform, and Supervisory Control and Data Acquisition (SCADA) or Distribution Management System integration components. The second concept is a Coordinated Anomaly Detection System (CADS) which forms a primarily software-related subsystem of the total CyberX solution focused on monitoring system measurements, estimation of measurement states, and predicting current at the utility point of interaction based on machine learning for anomaly detection.

33 ADVANCED PROPULSION SYSTEMS↗

Information content of and the ability to reconstruct dichroic X-ray tomography and laminography

Dichroic tomography is a 3D imaging technique in which the polarization of the incident beam is used to induce contrast due to the magnetization or orientation of a sample. The aim is to reconstruct not only the optical density but the dichroism of the sample. The theory of dichroic tomographic and laminographic imaging in the parallel-beam case is discussed as well as the problem of reconstruction of the sample’s optical properties. The set of projections resulting from a single tomographic/laminographic measurement is not sufficient to reconstruct the magnetic moment for magnetic circular dichroism unless additional constraints are applied or data are taken at two or more tilt angles. For linear dichroism, three polarizations at a common tilt angle are insufficient for unconstrained reconstruction. However, if one of the measurements is done at a different tilt angle than the other, or the measurements are done at a common polarization but at three distinct tilt angles, then there is enough information to reconstruct without constraints. Possible means of applying constraints are discussed. Furthermore, it is shown that for linear dichroism, the basic assumption that the absorption through a ray path is the integral of the absorption coefficient, defined on the volume of the sample, along the ray path, is not correct when dichroism or birefringence is strong. This assumption is fundamental to tomographic methods. An iterative algorithm for reconstruction of linear dichroism is demonstrated on simulated data.

Marcus, Matthew A. (ORCID:0000000325277586)↗

Enabling particle applications for exascale computing platforms

The Exascale Computing Project (ECP) is invested in co-design to assure that key applications are ready for exascale computing. Within ECP, the Co-design Center for Particle Applications (CoPA) is addressing challenges faced by particle-based applications across four “sub-motifs”: short-range particle–particle interactions (e.g., those which often dominate molecular dynamics (MD) and smoothed particle hydrodynamics (SPH) methods), long-range particle–particle interactions (e.g., electrostatic MD and gravitational N-body), particle-in-cell (PIC) methods, and linear-scaling electronic structure and quantum molecular dynamics (QMD) algorithms. Our crosscutting co-designed technologies fall into two categories: proxy applications (or “apps”) and libraries. Proxy apps are vehicles used to evaluate the viability of incorporating various types of algorithms, data structures, and architecture-specific optimizations and the associated trade-offs; examples include ExaMiniMD, CabanaMD, CabanaPIC, and ExaSP2. Libraries are modular instantiations that multiple applications can utilize or be built upon; CoPA has developed the Cabana particle library, PROGRESS/BML libraries for QMD, and the SWFFT and fftMPI parallel FFT libraries. Success is measured by identifiable “lessons learned” that are translated either directly into parent production application codes or into libraries, with demonstrated performance and/or productivity improvement. The libraries and their use in CoPA’s ECP application partner codes are also addressed.

97 MATHEMATICS AND COMPUTING↗