Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Enabling particle applications for exascale computing platforms

The Exascale Computing Project (ECP) is invested in co-design to assure that key applications are ready for exascale computing. Within ECP, the Co-design Center for Particle Applications (CoPA) is addressing challenges faced by particle-based applications across four “sub-motifs”: short-range particle–particle interactions (e.g., those which often dominate molecular dynamics (MD) and smoothed particle hydrodynamics (SPH) methods), long-range particle–particle interactions (e.g., electrostatic MD and gravitational N-body), particle-in-cell (PIC) methods, and linear-scaling electronic structure and quantum molecular dynamics (QMD) algorithms. Our crosscutting co-designed technologies fall into two categories: proxy applications (or “apps”) and libraries. Proxy apps are vehicles used to evaluate the viability of incorporating various types of algorithms, data structures, and architecture-specific optimizations and the associated trade-offs; examples include ExaMiniMD, CabanaMD, CabanaPIC, and ExaSP2. Libraries are modular instantiations that multiple applications can utilize or be built upon; CoPA has developed the Cabana particle library, PROGRESS/BML libraries for QMD, and the SWFFT and fftMPI parallel FFT libraries. Success is measured by identifiable “lessons learned” that are translated either directly into parent production application codes or into libraries, with demonstrated performance and/or productivity improvement. The libraries and their use in CoPA’s ECP application partner codes are also addressed.

97 MATHEMATICS AND COMPUTING↗

Axom

Axom is an open-source library of "building block" software components that provide core infrastructure capabilities for HPC applications. Such capabilities include: input file parsing and verification for various file formats, format specification for material shape input to multi-material simulations and tools for placing those shapes on meshes, parallel distributed coordination of diagnostic messages, geometric primitives, spatial queries and spatial search acceleration data structures, in memory key-value data store for managing simulation data and parallel file I/O. Axom components are used widely across LLNL Advanced Simulation and Computing program applications. It is also part of the LLNL institutionally-supported RADIUSS project, which promotes and funds its adoption by projects across LLNL.

Zagaris, George↗

Characterizing Output Bottlenecks of a Production Supercomputer: Analysis and Implications

This article studies the I/O write behaviors of the Titan supercomputer and its Lustre parallel file stores under production load. The results can inform the design, deployment, and configuration of file systems along with the design of I/O software in the application, operating system, and adaptive I/O libraries.We propose a statistical benchmarking methodology to measure write performance across I/O configurations, hardware settings, and system conditions. Moreover, we introduce two relative measures to quantify the write-performance behaviors of hardware components under production load. In addition to designing experiments and benchmarking on Titan, we verify the experimental results on one real application and one real application I/O kernel, XGC and HACC IO, respectively. These two are representative and widely used to address the typical I/O behaviors of applications.In summary, we find that Titan’s I/O system is variable across the machine at fine time scales. This variability has two major implications. First, stragglers lessen the benefit of coupled I/O parallelism (striping). Peak median output bandwidths are obtained with parallel writes to many independent files, with no striping or write sharing of files across clients (compute nodes). I/O parallelism is most effective when the application—or its I/O libraries—distributes the I/O load so that each target stores files for multiple clients and each client writes files on multiple targets in a balanced way with minimal contention. Second, our results suggest that the potential benefit of dynamic adaptation is limited. In particular, it is not fruitful to attempt to identify “good locations” in the machine or in the file system: component performance is driven by transient load conditions and past performance is not a useful predictor of future performance. For example, we do not observe diurnal load patterns that are predictable.

97 MATHEMATICS AND COMPUTING↗

Quantum Simulators and Applications on Quantum Framework

Simulating quantum circuits is essential for validating quantum algorithms. However, no single simulator consistently performs best - efficiency depends on circuit structure, entanglement, and depth. In this work, we integrate Qiskit-Aer (state-vector and matrix product state) and QTensor, a tree-tensor-network based simulator, into the Quantum Framework (QFw), a modular platform that supports multiple quantum backends via a unified interface. We also enable distributed quantum approximate optimization algorithm (DQAOA) application compatibility with QFw, allowing sub-problems to be solved in parallel at scale. We then benchmark DQAOA and TFIM (transverse field Ising model) circuits across supported simulators, showing how performance varies significantly with problem type. All simulations are deployed on the Frontier supercomputer using QFw's MPI-based orchestration for distributed, multinode execution. These results underscore the need for simulatoragnostic infrastructure to enable systematic evaluation and highperformance scaling of quantum workloads. QFw provides a practical and extensible path toward reproducible quantum algorithm development across diverse application domains.

Chundury, Srikar [ORNL] (ORCID:0009000183359259)↗

Game Theoretic Orchestration for Cooperation among Power Distribution System Applications

The evolving transformation with the proliferation of distributed energy resources and advanced metering, necessitates advanced distribution systems to integrate and orchestrate a large number of grid-edge devices while also serving multiple system-level objectives such as resilience, decarbonization, equity and other system mandates. The parallel deployment and control of resources towards achieving diverse objectives may lead to conflicts between applications that want to control overlapping sets of device setpoints, potentially leading to oscillatory behavior and suboptimal performance. This work aims at leveraging game theoretic framework to drive cooperative behavior among competitive applications. The work proposes a weighted-consensus based game design to facilitate conflict resolution through consensus-building iterations for modular platform. Simulation-based evaluation on a sample test system demonstrates the performance the proposed deconfliction strategy in resolving operational conflicts and achieving close-to-optimal trade off among the applications. Results also compare the proposed strategy with a distribution optimization approach and illustrate it effectiveness in diverse apps regardless of their design while also incentivizing apps with flexible design.

Advanced distribution operations, cooperation, app↗

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)↗

Advanced Characterization of Particulate Flows for Concentrating Solar Power Applications

The overarching goal of this project was to gain a greater fundamental understanding of heat and mass transfer in particulate-media in concentrated solar power (CSP) applications for thermal energy storage (TES) and to disseminate results to support parallel Gen3 research initiatives. Project objectives were achieved through three planned project phases to systematically characterize the heat transfer and flow properties for particulate (granular) flows at elevated temperatures up to 800 °C. These objectives were accomplished using a combination of fundamental experimental measurements, modeling, and simplified flow experiments over a range of temperatures. This work addressed a serious gap within the field related to the understanding and modeling of granular flow behavior and the related heat transfer at different temperatures, which directly correspond to the operating points of CSP applications which use particles for heat storage.

14 SOLAR ENERGY↗

Deep learning for electron and scanning probe microscopy: From materials design to atomic fabrication

Machine learning and artificial intelligence (ML/AI) are rapidly becoming an indispensable part of physics research, with applications ranging from theory and materials prediction to high-throughput data analysis. In parallel, the recent successes in applying ML/AI methods for autonomous systems from robotics through self-driving cars to organic and inorganic synthesis are generating enthusiasm for the potential of these techniques to enable automated and autonomous experiment in imaging. Here, we discuss recent progress in application of machine learning methods in scanning transmission electron microscopy and scanning probe microscopy, from applications such as data compression and exploratory data analysis to physics learning to atomic fabrication.

36 MATERIALS SCIENCE↗

PyOMP: Multithreaded Parallel Programming in Python

We know that Python is a widely used language in scientific computing. When the goal is high performance, however, Python lags far behind low-level languages such as C and Fortran. To support applications that stress performance, Python needs to access the full capabilities of modern CPUs. That means support for parallel multithreading. In this paper, we describe PyOMP, a system that enables OpenMP in Python. Programmers write code in Python with OpenMP, Numba generates code that compiles to LLVM, and the resulting programs run with performance that approaches that from code written with C and OpenMP. In this paper we provide an update on the PyOMP project and explain how to install it and use it to write parallel multithreaded code in Python.

97 MATHEMATICS AND COMPUTING↗

High-Level Synthesis of Irregular Applications: A Case Study on Influence Maximization

The Influence Maximization problem is the problem of identifying a small cohort of actors from a broader population that, when initially activated in a diffusion process, are expected to result in a large number of activations in the population. While the problem is known to be NP-hard, several approximation algorithms have been devised by leveraging its submodular structure. While these algorithms are theoretically efficient, they are computationally very expensive in practice. This work advances the current state-of-the-art parallelization scheme for the IMM algorithm by devising the adoption of custom hardware accelerators implemented on FPGAs by leveraging High Level Synthesis from OpenCL. We study the performance of our proposed approach by exploring optimizations tailored at improving the parallel efficiency of the accelerators and highlight their effects and limitations in accelerating complex graph analytic applications. Our experimental evaluation shows that FPGA acceleration can improve the performance of the LT diffusion model up to 1.72x for the entire application and up to 2.90x for its most important kernel with respect to a CPU only parallel execution. The FPGA acceleration of the LT model shows also a 1.54x reduction in energy consumption when compared to a parallel CPU only run.

Neff, Reece W.↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

COMPOFF: A Compiler Cost model using Machine Learning to predict the Cost of OpenMP Offloading

The HPC industry is inexorably moving towards an era of extremely heterogeneous architectures, with more devices configured on any given HPC platform and potentially more kinds of devices, some of them highly specialized. Writing a separate code suitable for each target system for a given HPC application is not practical. The better solution is to use directive-based parallel programming models such as OpenMP. OpenMP provides a number of options for offloading a piece of code to devices like GPUs. To select the best option from such options during compilation, most modern compilers use analytical models to estimate the cost of executing the original code and the different offloading code variants. Building such an analytical model for compilers is a difficult task that necessitates a lot of effort on the part of a compiler engineer. Recently, machine learning techniques have been successfully applied to build cost models for a variety of compiler optimization problems. In this paper, we present COMPOFF, a cost model which uses the multi-layer perceptrons to statically estimates the Cost of OpenMP OFFloading. We used six different transformations on a parallel code of Wilson Dslash Operator to support GPU offloading, and we predicted their cost of execution on different GPUs using COMPOFF during compile time. Our results show that this model can predict offloading costs with a root mean squared error in prediction of less than 0.5 seconds. Our preliminary findings indicate that this work will make it much easier and faster for scientists and compiler developers to port legacy HPC applications that use OpenMP to new heterogeneous computing environment.

97 MATHEMATICS AND COMPUTING↗

Single‐Domain Multiferroic Array‐Addressable Terfenol‐D (SMArT) Micromagnets for Programmable Single‐Cell Capture and Release

Abstract Programming magnetic fields with microscale control can enable automation at the scale of single cells ≈10 µm. Most magnetic materials provide a consistent magnetic field over time but the direction or field strength at the microscale is not easily modulated. However, magnetostrictive materials, when coupled with ferroelectric material (i.e., strain‐mediated multiferroics), can undergo magnetization reorientation due to voltage‐induced strain, promising refined control of magnetization at the micrometer‐scale. This work demonstrates the largest single‐domain microstructures (20 µm) of Terfenol‐D (Tb 0.3 Dy 0.7 Fe 1.92 ), a material that has the highest magnetostrictive strain of any known soft magnetoelastic material. These Terfenol‐D microstructures enable controlled localization of magnetic beads with sub‐micrometer precision. Magnetically labeled cells are captured by the field gradients generated from the single‐domain microstructures without an external magnetic field. The magnetic state on these microstructures is switched through voltage‐induced strain, as a result of the strain‐mediated converse magnetoelectric effect, to release individual cells using a multiferroic approach. These electronically addressable micromagnets pave the way for parallelized multiferroics‐based single‐cell sorting under digital control for biotechnology applications.

Khojah, Reem↗

Stochastic AC optimal power flow: A data-driven approach

There is an emerging need for efficient solutions to stochastic AC Optimal Power Flow (AC-OPF) to ensure optimal and reliable grid operations in the presence of increasing demand and generation uncertainty. Herein this paper presents a highly scalable data-driven algorithm for stochastic AC-OPF that has extremely low sample requirement. The novelty behind the algorithm’s performance involves an iterative scenario design approach that merges information regarding constraint violations in the system with data-driven sparse regression. Compared to conventional methods with random scenario sampling, our approach is able to provide feasible operating points for realistic systems with much lower sample requirements. Furthermore, multiple sub-tasks in our approach can be easily paralleled and based on historical data to enhance its performance and application. We demonstrate the computational improvements of our approach through simulations on different test cases in the IEEE PES PGLib-OPF benchmark library.

42 ENGINEERING↗

Grid-based minimization at scale: Feldman-Cousins corrections for light sterile neutrino search

High Energy Physics (HEP) experiments generally employ sophisticated statistical methods to present results in searches of new physics. In the problem of searching for sterile neutrinos, likelihood ratio tests are applied to short-baseline neutrino oscillation experiments to construct confidence intervals for the parameters of interest. The test statistics of the form Δχ2 is often used to form the confidence intervals, however, this approach can lead to statistical inaccuracies due to the small signal rate in the region-of-interest. In this paper, we present a computational model for the computationally expensive Feldman-Cousins corrections to construct a statistically accurate confidence interval for neutrino oscillation analysis. The program performs a grid-based minimization over oscillation parameters and is written in C++. Our algorithms make use of vectorization through Eigen3, yielding a single-core speed-up of 350 compared to the original implementation, and achieve MPI data parallelism by employing DIY. We demonstrate the strong scaling of the application at High-Performance Computing (HPC) sites. We utilize HDF5 along with HighFive to write the results of the calculation to file.

Wospakrik, Marianette↗

Fokker-Planck simulations of fast ion ICRF and electron EC heating in a mirror plasma using CQL3D-m

The CQL3D-m continuum bounce-average Fokker-Planck code is adapted for magnetic mirror plasmas [1] and is now routinely used in no-free-parameter classical integrated modeling of mirror devices [2, 3]. In the present effort, we report on two RF methods of plasma heating in mirror machine. The fast ions (FI) are heated by Fast waves at 2nd-4th harmonic, where FIs originate from neutral beam injection at 45 degrees to the magnetic field. The scenario shows an efficient ion heating near the FI bouncing point. The electrons are heated by X-mode launched from the high magnetic field side towards the resonance. Different from the tokamak applications, CQL3D-m provides an evolving self-consistent ambipolar parallel electric field, which determines the shape of the loss cone and hence an accurate confinement time of both ions and electrons. Also, it includes a description of ion and electron sources and sinks (related to charge exchange and impact ionization) which are updated at every time step. CQL3D-m utilizes a fully nonlinear Coulomb collision operator that is important for the significantly non-Maxwellian ion distributions typically established in mirror plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Closure theory for high-collisionality multi-ion plasmas

A general formalism is developed to construct and solve a system of linearized moment equations for parallel and perpendicular closures in high-collisionality plasmas. It is applicable for multiple ion species with arbitrary masses, temperatures, charges, and densities. The convergence of closure coefficients is evaluated by increasing the number of moments from 2 to 32 for scalar, vector, and rank-2 tensor moments. As an example, the complete set of closure coefficients for a deuterium-carbon plasma over the entire Hall parameter range is presented. Furthermore, the closure coefficients at various temperature ratios show that the one-temperature closure coefficients can differ significantly from the two-temperature coefficients.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗