Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel application”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

PNNL-CompBio/BoltzmannMFX

BoltzmannMFX is a biological simulation code that solves chemical reaction networks using maximum entropy methods. It uses modules from MFiX-Exa and is based on the AMReX framework for massively parallel block-structured adaptive mesh applications.

Palmer, Bruce↗

Simulating microgalvanic corrosion in alloys using the PRISMS phase-field framework

In this prospective paper, we first review the existing simulation tools to simulate microgalvanic corrosion during free immersion. Then, we describe a recently developed application that employs PRISMS-PF, an open-source, high-performance phase-field modeling framework. The model employed in the application accounts for the electrochemical reaction at the metal/electrolyte interface and ionic migration in the electrolyte to determine the evolution of the corrosion front. We present the implementation details for the application and discuss its features such as super-linear parallel scaling performance for a sufficiently large system. Finally, we demonstrate the capability of the application by simulating corrosion of the matrix phase of an alloy near a secondary phase particle in two and three dimensions.

36 MATERIALS SCIENCE↗

Description of Sensor Assignment Optimization Method as Deployed on a Multi-Node Cluster

Data analytic methods are being developed to address the problem of how to assign a sensor set in a nuclear facility such that a requisite level of process monitoring capability is realized and that the sensor set is sufficiently rich to determine the status of the individual sensors with respect to need for calibration. There is an awareness in the nuclear industry that data analytics combined with rich sensor sets represent a means to improve operations and reduce costs. In the industry the calibration problem has been previously approached as an empirical data-driven problem with several methods having been developed. However, the experience of the utilities over the past ten years with these methods indicates that the absence of physics-based information renders the data-driven approach less reliable. Complicating factors such as the inherent variability of operation (both equipment alignment and operating condition) can confound a pure data-driven approach while there are no rigorous guidelines for determining what constitutes an adequate sensor set. The solution under development to overcome these shortcomings supplements the data analytic method with process information in a so-called process-constrained data-analytic approach. Simple balance equations are written for generic components (e.g., mechanical pump, valve, and heat exchanger). These do not require a priori knowledge of process parameters, such as heat transfer coefficients or friction factors. All that is needed on the part of the utility user is to identify the components and how they are connected. This report describes the development of a parallel computing capability for determining the optimal sensor set. The optimal sensor set problem suffers from the curse of dimensionality. Computation time increases exponentially as the size of the system grows. To overcome this difficulty a pre-conditioner algorithm is developed to find an approximate solution close the actual solution. This serves as a seed for the full-blown algorithm and acts to constrain the space that must searched. The optimization algorithms are described and the implementation on a parallel computing platform is described. The application of the method to a use case we are solving in collaboration with our utility partner served to illustrate how the default sensor set in a nuclear plant may not provide sufficient coverage to infer sensor calibration status.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Charliecloud 101

This workshop will provide participants with background and hands-on experience to use basic containers for HPC applications. We will discuss what containers are, why they matter for HPC, and how they work. We’ll give an overview of Charliecloud, the unprivileged container solution from HPC Division, and walk participants through installing it on their own compute resource. Participants will build toy containers and a real HPC application, and then run them in parallel on an HPC Division cluster. This will be a highly interactive workshop with lots of Q&A.

97 MATHEMATICS AND COMPUTING↗

Practical applications of machine-learned flows on gauge fields

Normalizing flows are machine-learned maps between different lattice theories which can be used as components in exact sampling and inference schemes. Ongoing work yields increasingly expressive flows on gauge fields, but it remains an open question how flows can improve lattice QCD at state-of-the-art scales. We discuss and demonstrate two applications of flows in replica exchange (parallel tempering) sampling, aimed at improving topological mixing, which are viable with iterative improvements upon presently available flows.

Abbott, Ryan↗

pyAMReX v23.08

The Python binding for AMReX, pyAMReX, bridges the worlds of block-structured codes and data science: it provides zero-copy application GPU data access for AI/ML, in situ analysis, application coupling and enables rapid, massively parallel prototyping.

Huebl, Axel↗

Devastator Parallel Discrete Event Simulation Runtime (Devastator) v1.0

The Devastator runtime is a modern C++ implementation of optimistic parallel discrete event simulation methods. Devastator allows simulation application code to productively specify their component and event functionality with C++14 constructs. It utilizes GASNet-EX for distributed memory communication and includes parallel performance optimizations such as light-weight thread message queues and asynchronous GVT. Furthermore, it supports efficient event broadcasts and pause-rewind-resume functionality to support periodic load balancing and outer loop optimization algorithms.

Chan, Cy↗

Molecular underpinnings of ssDNA specificity by Rep HUH-endonucleases and implications for HUH-tag multiplexing and engineering

Abstract Replication initiator proteins (Reps) from the HUH-endonuclease superfamily process specific single-stranded DNA (ssDNA) sequences to initiate rolling circle/hairpin replication in viruses, such as crop ravaging geminiviruses and human disease causing parvoviruses. In biotechnology contexts, Reps are the basis for HUH-tag bioconjugation and a critical adeno-associated virus genome integration tool. We solved the first co-crystal structures of Reps complexed to ssDNA, revealing a key motif for conferring sequence specificity and for anchoring a bent DNA architecture. In combination, we developed a deep sequencing cleavage assay, termed HUH-seq, to interrogate subtleties in Rep specificity and demonstrate how differences can be exploited for multiplexed HUH-tagging. Together, our insights allowed engineering of only four amino acids in a Rep chimera to predictably alter sequence specificity. These results have important implications for modulating viral infections, developing Rep-based genomic integration tools, and enabling massively parallel HUH-tag barcoding and bioconjugation applications.

59 BASIC BIOLOGICAL SCIENCES↗

Understanding the use of message passing interface in exascale proxy applications

The Exascale Computing Project (ECP) focuses on the development of future exascale-capable applications. Most ECP applications use the message passing interface (MPI) as their parallel programming model with mini-apps serving as proxies. This paper explores the explicit usage of MPI in such ECP proxy applications. We empirically analyze 14 proxy applications from the ECP Proxy Apps Suite. We use the MPI profiling interface (PMPI) to collect MPI usage patterns in ECP proxy apps. Our analysis shows that a small subset of features from MPI is commonly used in the proxies of exascale-capable applications, even when they reference third-party libraries. Overall, this study is intended to provide a better understanding of the use of MPI in current exascale applications. The findings can help focus software investments made for exascale systems in the MPI middleware including optimization, fault-tolerance, tuning, and hardware-offload.

97 MATHEMATICS AND COMPUTING↗

Vidyut3d: A Non-Equilibrium Plasma Modeling Tool [SWR-24-101]

Vidyut3d is a massively-parallel plasma-fluid solver for low-temperature plasmas (LTPs) that supports both local field (LFA) and local mean energy (LMEA) approximations, as well as complex gas and surface-phase chemistry. The solver supports 2D and 3D domains, and uses AMReX's adaptive mesh refinement capabilities to increase the grid resolution around complex structures (e.g. streamer heads and sheaths) while maintaining a tractable problem size. Vidyut specializes in simulating various types of gas-phase discharges, as well as plasma/surface interactions and surface chemistry (e.g. for plasma-mediated catalysis applications). The solver also supports hybrid CPU/GPU parallelization strategies, and has demonstrated excellent scaling on various HPC architectures for problem sizes consisting of O(100 M) control volumes.

Sitaraman, Hariswaran↗

Sprain energy consequences for damage localization and fracture mechanics

The 2023 smooth Lagrangian Crack-Band Model (slCBM), inspired by the 2020 invention of the gap test, prevented spurious damage localization during fracture growth by introducing the second gradient of the displacement field vector, named the “sprain,” as the localization limiter. The key idea was that, in the finite element implementation, the displacement vector and its gradient should be treated as independent fields with the lowest ( C 0 ) continuity, constrained by a second-order Lagrange multiplier tensor. Coupled with a realistic constitutive law for triaxial softening damage, such as microplane model M7, the known limitations of the classical Crack Band Model were eliminated. Here, we show that the slCBM closely reproduces the size effect revealed by the gap test at various crack-parallel stresses. To describe it, we present an approximate corrective formula, although a strong loading-path dependence limits its applicability. Except for the rare case of zero crack-parallel stresses, the fracture predictions of the line crack models (linear elastic fracture mechanics, phase-field, extended finite element method (XFEM), cohesive crack models) can be as much as 100% in error. We argue that the localization limiter concept must be extended by including the resistance to material rotation gradients. We also show that, without this resistance, the existing strain-gradient damage theories may predict a wrong fracture pattern and have, for Mode II and III fractures, a load capacity error as much as 55%. Finally, we argue that the crack-parallel stress effect must occur in all materials, ranging from concrete to atomistically sharp cracks in crystals.

Science & Technology - Other Topics↗

All-Atom Simulation of 3D Hot Spot Formation in Shocked TATB Explosive

TATB is an insensitive high explosive (IHE) critical to the stockpile that is challenging to model at the continuum scale. Advanced detonation models in the Cheetah high explosive chemistry code require validation though subscale simulations. High explosive initiation is determined by micron-scale physics of hot spots formed a shock-collapsed pores. Pore sizes between 100 nm and 1 μm are believed to be the most important for determining the shock sensitivity of TATB. This range of pore sizes is difficult to access at the atomic scale through allatom molecular dynamics (MD) simulations, even with Sierra-class computers. Quasi-2D simulations are widely used and allow much larger pore sizes (up to 400 nm) to be studied, but the applicability of 2D simulations to the actual 3D pore response is not understood. Resolving these uncertainties through “full physics” MD modeling is key for generalizing, parameterizing, and validating the kinds of continuum models used to inform design, safety, and performance. This work was a continuation of FY20 efforts pushing simulations to full 3D with the largest-ever all-atom simulations of an explosive. These were the first all-atom full-3D simulations of large hot spots thought to govern explosive detonation and required over a billion atoms. Simulations were performed using LAMMPS, an open SNL science code. MD explosive models present unique challenges, even for established codes such as LAMMPS. Their model forms are more complex than typical models for metals, while simulating high temperature-pressure conditions is demanding and increases computational cost. Scaling problems in GPU-enabled MD algorithms initially limited simulations to <100 million atoms but were resolved through collaboration with SNL. An overall 24x speedup was obtained relative to CPU machines. Specialized analysis of these simulations required a bottom-up refactoring and algorithm parallelization of in-house codes and application of computer vision algorithms to extract meaningful information.

36 MATERIALS SCIENCE↗

Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems

The semantics of HPC storage systems are defined by the consistency models to which they abide. Storage consistency models have been less studied than their counterparts in memory systems, with the exception of the POSIX standard and its strict consistency model. The use of POSIX consistency imposes a performance penalty that becomes more significant as the scale of parallel file systems increases and the access time to storage devices, such as node-local solid storage devices, decreases. While some efforts have been made to adopt relaxed storage consistency models, these models are often defined informally and ambiguously as by-products of a particular implementation. Here in this work, we establish a connection between memory consistency models and storage consistency models and revisit the key design choices of storage consistency models from a high-level perspective. Further, we propose a formal and unified framework for defining storage consistency models and a layered implementation that can be used to easily evaluate their relative performance for different I/O workloads. Finally, we conduct a comprehensive performance comparison of two relaxed consistency models on a range of commonly seen parallel I/O workloads, such as checkpoint/restart of scientific applications and random reads of deep learning applications. We demonstrate that for certain I/O scenarios, a weaker consistency model can significantly improve the I/O performance. For instance, in small random reads that are typically found in deep learning applications, session consistency achieved a 5x improvement in I/O bandwidth compared to commit consistency, even at small scales.

97 MATHEMATICS AND COMPUTING↗

A direct transcription-based multiple shooting formulation for dynamic optimization

The growing need for fast and efficient solution techniques for solving dynamic optimization problems is driven by a broad spectrum of applications in scheduling and control. We suggest a novel framework for dynamic optimization that utilizes a multiple shooting “backbone” with discrete rather than continuous subproblems, thereby eliminating need for repeated time-integration. A Lagrangian relaxation (LR)-based decomposition scheme is proposed, which dualizes the state continuity requirements between subproblems and enables parallel solution of the problem. We demonstrate the applicability of our method on two case studies: the Van der Pol oscillator and a batch reactor.

42 ENGINEERING↗

Characterization and identification of HPC applications at leadership computing facility

High Performance Computing (HPC) is an important method for scientific discovery via large-scale simulation, data analysis, or artificial intelligence. Leadership-class supercomputers are expensive, but essential to run large HPC applications. The Petascale era of supercomputers began in 2008, with the first machines achieving performance in excess of one petaflops, and with the advent of new supercomputers in 2021 (e.g., Aurora, Frontier), the Exascale era will soon begin. However, the high theoretical computing capability (i.e., peak FLOPS) of a machine is not the only meaningful target when designing a supercomputer, as the resources demand of applications varies. A deep understanding of the characterization of applications that run on a leadership supercomputer is one of the most important ways for planning its design, development and operation. In order to improve our understanding of HPC applications, user demands and resource usage characteristics, we perform correlative analysis of various logs for different subsystems of a leadership supercomputer. This analysis reveals surprising, sometimes counter-intuitive patterns, which, in some cases, conflicts with existing assumptions, and have important implications for future system designs as well as supercomputer operations. For example, our analysis shows that while the applications spend significant time on MPI, most applications spend very little time on file I/O. Combined analysis of hardware event logs and task failure logs show that the probability of a hardware FATAL event causing task failure is low. Combined analysis of control system logs and file I/O logs reveals that pure POSIX I/O is used more widely than higher level parallel I/O. Based on holistic insights of the application gained through combined and co-analysis of multiple logs from different perspectives and general intuition, we engineer features to "fingerprint" HPC applications. We use t-SNE (a machine learning technique for dimensionality reduction) to validate the explainability of our features and finally train machine learning models to identify HPC applications or group those with similar characteristic. To the best of our knowledge, this is the first work that combines logs on file I/O, computing, and inter-node communication for insightful analysis of HPC applications in production.

Liu, Zhengchun↗

Sculpting the Plasmonic Responses of Nanoparticles by Directed Electron Beam Irradiation

Spatial confinement of matter in functional nanostructures has propelled these systems to the forefront of nanoscience, both as a playground for exotic physics and quantum phenomena and in multiple applications including plasmonics, optoelectronics, and sensing. In parallel, the emergence of monochromated electron energy loss spectroscopy (EELS) has enabled exploration of local nanoplasmonic functionalities within single nanoparticles and the collective response of nanoparticle assemblies, providing deep insight into associated mechanisms. However, modern synthesis processes for plasmonic nanostructures are often limited in the types of accessible geometry, and materials and are limited to spatial precisions on the order of tens of nm, precluding the direct exploration of critical aspects of the structure-property relationships. Here, the atomic-sized probe of the scanning transmission electron microscope is used to perform precise sculpting and design nanoparticle configurations. Using low-loss EELS, dynamic analyses of the evolution of the plasmonic response are provided. It is shown that within self-assembled systems of nanoparticles, individual nanoparticles can be selectively removed, reshaped, or patterned with nanometer-level resolution, effectively modifying the plasmonic response in both space and energy. This process significantly increases the scope for design possibilities and presents opportunities for unique structure development, which are ultimately the key for nanophotonic design.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

High-throughput determination of Hubbard $U$ and Hund $J$ values for transition metal oxides via the linear response formalism

DFT+U provides a convenient, cost-effective correction for the self-interaction error (SIE) that arises when describing correlated electronic states using conventional approximate density functional theory (DFT). The success of a DFT+U(+J) calculation hinges on the accurate determination of its Hubbard U and Hund J parameters, and the linear response (LR) methodology has proven to be computationally effective and accurate for calculating these parameters. This study provides a high-throughput computational analysis of the U and J values for transition metal d-electron states in a representative set of over 1000 magnetic transition metal oxides (TMOs), providing a frame of reference for researchers who use DFT+U to study transition metal oxides. In order to perform this high-throughput study, an ATOMATE workflow is developed for calculating U and J values automatically on massively parallel supercomputing architectures. Here, to demonstrate an application of this workflow, the spin-canting magnetic structure and unit cell parameters of the multiferroic olivine LiNiPO4 are calculated using the computed Hubbard U and Hund J values for Ni-d and O-p states, and are compared with experiment. Both the Ni-d U and J corrections have a strong effect on the Ni-moment canting angle. Additionally, including a O-pU value results in a significantly improved agreement between the computed lattice parameters and experiment

36 MATERIALS SCIENCE↗