Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Generalizable coordination of large multiscale workflows: challenges and learnings at scale

The advancement of machine learning techniques and the heterogeneous architectures of most current supercomputers are propelling the demand for large multiscale simulations that can automatically and autonomously couple diverse components and map them to relevant resources to solve complex problems at multiple scales. Nevertheless, despite the recent progress in workflow technologies, current capabilities are limited to coupling two scales. In the first-ever demonstration of using three scales of resolution, we present a scalable and generalizable framework that couples pairs of models using machine learning and in situ feedback. We expand upon the massively parallel Multiscale Machine-Learned Modeling Infrastructure (MuMMI), a recent, award-winning workflow, and generalize the framework beyond its original design. We discuss the challenges and learnings in executing a massive multiscale simulation campaign that utilized over 600,000 node hours on Summit and achieved more than 98% GPU occupancy for more than 83% of the time. We present innovations to enable several orders of magnitude scaling, including simultaneously coordinating 24,000 jobs, and managing several TBs of new data per day and over a billion files in total. Finally, we describe the generalizability of our framework and, with an upcoming open-source release, discuss how the presented framework may be used for new applications.

Bhatia, Harsh↗

Implementing Arbitrary/Common Concurrent Writes of CRCW PRAM

The Parallel Random Access Machines (PRAM) abstraction is the simplest and most elegant algorithmic model for the design and analysis of parallel algorithms. It consists of different models categorized based on the underlying memory access mode used, the most powerful of which is the Concurrent Read Concurrent Write (CRCW) model. A PRAM algorithm describes a series of rounds, each of which consists of a collection of operations that can be executed concurrently within the same time step. However, the lack of support for concurrent memory accesses and the prevalence of asynchronous programming models led to the belief that implementing CRCW PRAM algorithms is unattainable and prompted many to avoid this model except for theoretical studies of optimal performance.In this work, we study the arbitrary and common concurrent writes in the CRCW PRAM model and explore implementation challenges on general-purpose systems. Moreover, we examine current practices for implementing common/arbitrary concurrent writes and propose a new efficient lightweight and thread-safe method to implement concurrent writes through leveraging atomic instructions. To demonstrate the efficacy of our method, we developed OpenMP kernels for classical CRCW PRAM algorithms and provide experimental results and comparisons based on run time performance measured over the x86 multicore architecture. Our results show a performance speedup compared to current practices up to 4.5x across all our benchmarks.

Ghanim, Fady↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Towards fast, accurate predictions of RF simulations via data-driven modeling: Forward and lateral models

Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. A single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time to complete, which is acceptable for many purposes, but too slow for integrated modeling and real-time control applications. More accurate simulations with fast electron diffusion are even slower, requiring multiple hours of run time with parallel processing. The machine learning models use a database of 16,000+ GEN-RAY/CQL3D simulations for training, validation, and testing. Latin hypercube sampling methods implemented in πScope ensure that the database covers the range of 9 input parameters (n e0 , T e0 , I p , B t , R 0 , n ∥︀ , Z e f f , V loop , P LHCD ) with sufficient density in all regions of parameter space. The surrogate models reduce the computation time from minutes-hours to ms with high accuracy across the input parameter space. Data-driven surrogate models also allow for solving inverse and “lateral” problems. A surrogate model for the inverse problem maps from a desired current drive or power deposition profile to a set of input parameters that would result in such a profile, while a surrogate model for the lateral problem maps from a measured experimental quantity such as hard x-ray emission to a current drive or power deposition profile. In conclusion, the πScope database creation workflow is flexible and applicable to other RF simulation codes such as TORIC.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Muon Collider Forum report

A multi-TeV muon collider offers a spectacular opportunity in the direct exploration of the energy frontier. Offering a combination of unprecedented energy collisions in a comparatively clean leptonic environment, a high energy muon collider has the unique potential to provide both precision measurements and the highest energy reach in one machine that cannot be paralleled by any currently available technology. The topic generated a lot of excitement in Snowmass meetings and continues to attract a large number of supporters, including many from the early career community. In light of this very strong interest within the US particle physics community, Snowmass Energy, Theory and Accelerator Frontiers created a cross-frontier Muon Collider Forum in November of 2020. The Forum has been meeting on a monthly basis and organized several topical workshops dedicated to physics, accelerator technology, and detector R&D. Findings of the Forum are summarized in this report.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Receptance coupling substructure analysis and chatter frequency-informed machine learning for milling stability

This paper describes a milling stability identification approach that simultaneously considers: physics-based models for the tool tip frequency response functions and stability predictions; the binary result from a milling test (automatically labeled as stable or unstable based on frequency content); chatter frequency when an unstable result is obtained; and user risk tolerance. The algorithm applies probabilistic Bayesian machine learning with adaptive, parallelized Markov Chain Monte Carlo sampling to update the probability of stability with each milling test. Furthermore, the result is a robust solution for rapid convergence to optimized milling parameters for maximum metal removal rate using all available information.

42 ENGINEERING↗

A general approach to seismic inversion with automatic differentiation

Imaging Earth structure or seismic sources from seismic data involves minimizing a target misfit function, and is commonly solved through gradient-based optimization. The adjoint-state method has been developed to compute the gradient efficiently; however, its implementation can be time-consuming and difficult. We develop a general seismic inversion framework to calculate gradients using reverse-mode automatic differentiation. The central idea is that adjoint-state methods and reverse-mode automatic differentiation are mathematically equivalent. Here, the mapping between numerical PDE simulation and deep learning allows us to build a seismic inverse modeling library, ADSeismic, based on deep learning frameworks, which supports high performance reverse-mode automatic differentiation on CPUs and GPUs. We demonstrate the performance of ADSeismic on inverse problems related to velocity model estimation, rupture imaging, earthquake location, and source time function retrieval. ADSeismic has the potential to solve a wide variety of inverse modeling applications within a unified framework.

58 GEOSCIENCES↗

High-Throughput Discovery Illuminates Design Principles and Limits for Long-Lived Charged Species in Organic Electrolytes

The chemical stability of charged molecules in all-organic redox flow batteries (RFBs) is required for the prolonged operation of these devices. Molecular engineering and electrolyte optimization are used to mitigate parasitic reactions and extend the lifetimes of the charge carriers. However, how much can structural variation extend the lifetime? To probe this query, we designed a high-throughput kinetic study of the radical cation of N-methylphenothiazinium, guided by statistical sampling and learning algorithms. Using Argonne’s autonomous discovery facility, we conducted over 6,000 kinetic experiments with robotic sample preparation, parallel kinetic measurements, and machine learning inputs, testing 188 solvent molecules selected from a space of over 540 candidates from 11 chemical classes. Algorithmic selections guided us to stable solvent candidates, which were further tested in high concentration with and without supporting electrolyte. Our findings reveal the inherent difficulty of exceeding the current state of the art through solvent variation. The desired stability is statistically rare and poorly predictable. Among the many tested, only three solvents significantly outperformed our baseline, acetonitrile─and none by more than a factor of 3─suggesting a general challenge in achieving the necessary techno-economic targets. Furthermore, we suggest that self-discharge through solvent homolysis is the cause of the observed limitations. Several structural motifs contribute to >1,000 h half-life stability including molecular simplicity, symmetry, oxidation complement, and strategic fluorination. Importantly, this workflow establishes effective assays for diagnosing and predicting oxidative stress for highly stable liquid electrolytes in all batteries.

Batteries↗

Automated model-predictive design of synthetic promoters to control transcriptional profiles in bacteria

Abstract Transcription rates are regulated by the interactions between RNA polymerase, sigma factor, and promoter DNA sequences in bacteria. However, it remains unclear how non-canonical sequence motifs collectively control transcription rates. Here, we combine massively parallel assays, biophysics, and machine learning to develop a 346-parameter model that predicts site-specific transcription initiation rates for any σ 70 promoter sequence, validated across 22132 bacterial promoters with diverse sequences. We apply the model to predict genetic context effects, design σ 70 promoters with desired transcription rates, and identify undesired promoters inside engineered genetic systems. The model provides a biophysical basis for understanding gene regulation in natural genetic systems and precise transcriptional control for engineering synthetic genetic systems.

42 ENGINEERING↗

The US effort towards a muon collider

A multi-TeV muon collider has the unique potential to provide both precision measurements and the highest energy reach in one machine that cannot be paralleled by any currently available technology. There has been significant physics interest on Muon Colliders recently as indicated by the number of publications, relevant workshops, Snowmass activities but also the P5 report. This study describes a possible set of R&D and deliverables of the muon collider accelerator R&D program in the U.S. We describe high-priority studies to be performed in the first phase that will address critical questions for deciding the future plan for a muon collider design. The goal of these studies is to firm up choices for the most challenging components of a muon collider design, and to propose and begin testing and prototyping of components and systems that are needed to have confidence in and inform our specification choices. Key areas wherein the US can provide critical contributions to the newly formed international muon collider collaboration will be discussed as well.

43 PARTICLE ACCELERATORS↗

AXEAP (ARGONNE X-RAY EMISSION PACKAGE)

Argonne X-ray Emission Package (AXEAP), a singular purpose software package for processing X-ray emission (XES) images collected with a 2-dimensional position sensitive pixel array detector, has been developed. AXEAP can rapidly convert XES image files into a spectral form by applying parallel computation and unsupervised machine learning to compute vast amount of image data. Special focus has been placed on designing user-friendly-interface for processing multiple edges, non-resonant and resonant x-ray emission image analysis, in order to make data processing quick and easy. AXEAP is free software and is written in MATLAB, armed with powerful libraries and toolboxes. The software runs on all common operating systems such as Linux, Window, and Mac.

SUN, CHENGJUN↗

Sentinel

Network intrusion detection systems (NIDS) are commonplace in network security but they frequently employ algorithms that are computational demanding requiring hardware and software with significant power requirements. Two examples of such resource-intensive algorithms used for network security are regular expression matching and broader signature pattern matching which are commonly used in deep packet inspection (DPI). Network security algorithms that have large power requirements may be a challenge for low-power internet-of-things (IoT) environments, which generally lack the power resources to implement complex security measures like computationally expensive DPI at the edge. Furthermore, IoT environments incorporating 5G standalone networks have network latency constraints beyond just power that make DPI at the edge even more difficult. Programmable logic is ideally suited for machine learning inference for DPI because of its deep instruction level parallelism and single-cycle memory access. Machine learning approaches for DPI have been explored before using the programmable logic of field programmable gate arrays (FPGA) as a potential solution for NIDS approaches that would be power-suitable for IoT. However, those previous programmable logic NIDS approaches utilize either a supervised or unsupervised learning model. Sentinel utilizes the ensemble of these two machine learning approaches known as a semi-supervised approach which has shown promise in NIDS implementations. Sentinel provides a programmable logic implementation of a semi-supervised approach for DPI which operates at much lower power and latency than a GPU implementation with negligible loss of accuracy due to quantization through a logistic regressor.

Anderson, MatthewW [Idaho National Laboratory (INL↗

Utilizing ensemble learning for performance and power modeling and improvement of parallel cancer deep learning CANDLE benchmarks

Abstract Machine learning (ML) continues to grow in importance across nearly all domains in modeling to learn from data. Often a tradeoff exists between a model's ability to minimize bias and variance. In this article, we utilize ensemble learning to combine linear, nonlinear, and tree‐/rule‐based ML methods to cope with the bias‐variance tradeoff and result in more accurate models. We use the datasets collected for two parallel cancer deep learning CANDLE benchmarks, NT3 and P1B2, to build performance and power models based on hardware performance counters using single‐object and multiple‐objects ensemble learning to identify the most important counters for improvement on the Cray XC40 Theta at Argonne National Laboratory. Based on the insights from these models, we improve the performance and energy of P1B2 and NT3 by optimizing the deep learning environments TensorFlow, Keras, Horovod, and Python under the huge page size of 8 MB. Experimental results show that ensemble learning not only produces more accurate models but also provides more robust performance counter ranking. We achieve up to 61.15% performance improvement and up to 62.58% energy saving for P1B2 and up to 55.81% performance improvement and up to 52.60% energy saving for NT3 on up to 24,576 cores.

Wu, Xingfu↗

Rapid Commissioning of Large Machine Tools Using Finite Element-Based Correction of Geometric Errors

Large computer numerical control (CNC) machine tools derive their stiffness from monolithic cast iron bases or weldments that are sometimes integral to machine motion systems like box ways or guideways. However, the sheer size of castings and even floor flatness deviations result in dimensional errors in these systems, which manifest as machine motion errors. Typical geometric alignment processes rely on an iterative approach, where measurements are taken to assess alignment (straightness, squareness, and parallelism), followed by adjustment of the machine supports (fixators or leveling pads), which can take weeks even for an experienced operator. Conversely, a novel method is proposed to shorten the correction time by eliminating the trial-and-error process in favor of a more deterministic approach guided by a finite element (FE) method. A feasibility study is conducted on a CNC polymer hybrid machine, with a steel weldment frame, supported by six leveling pads. An FE model of the frame is utilized to obtain recommended leveling pad adjustments, based on measurement of machine errors taken using a laser tracker. After a single adjustment cycle, measurements reveal that geometric errors of the machine tool are reduced from 2.22 mm of flatness deviation to 0.32 mm, achieving an 85.6% reduction. Furthermore, the entire process including measurement, adjustment, and assessment is completed in just 6 h by two operators who are not professional service engineers. In conclusion, this methodology demonstrates feasibility for scaling up, especially to large, high-precision CNC machine tools with bases mounted by fixators, offering the capability for bidirectional adjustment.

42 ENGINEERING↗

Laser Powder Bed Fusion Additive Manufacture Nb1Zr Development

Next generation fission and fusion nuclear reactors require materials that can withstand operating temperatures greater than 500 °C, neutron irradiation doses of up to 200 displacements per atom (dpa), and potentially corrosive coolants such as the alkali liquid metals sodium, lithium, and NaK (Na33K eutectic alloy). Refractory alloys, such as Nb1Zr (Nb-1wt%Zr) and Molybdenum alloy TZM (Mo-0.5wt%Ti-0.08wt%Zr) have been traditionally considered viable candidates for advanced fission and fusion reactor concepts. However, it is relatively difficult to generate complex geometries of interest from these alloys using traditional manufacturing methods. In addition, there needs to be a concentrated effort to address refractory metal challenges at elevated temperature operation. In order to generate complex geometries of interest, modern manufacturing techniques are considered to increase the technological readiness level (TRL), cost-effectiveness, and schedule savings. This work focused on the continued development of laser powder bed fusion (L-PBF) additive manufacturing (AM) to improve both design flexibility, evaluate microstructure and properties, and ultimately accelerate the TRL and qualification of these processes and alloys for components to potentially be put into service. Niobium alloy Nb1Zr was identified through a down-selection process outlined in previous reports as a candidate to develop in L-PBF AM. Historically, Nb1Zr had been explored for high temperature fast spectrum fission reactors for both terrestrial and space applications. Molybdenum alloy TZM has also been considered for these reactor concepts due to exceptional high-temperature strength, creep resistance, and stability under irradiation. L-PBF AM of TZM has previously been investigated at LANL under the Microreactor program, NASA, ORNL, and in academia. However, due to the crack prone nature of TZM, L-PBF AM of TZM resulted in significant microcracking and additional development is required to pursue viable maturation. Other AM methods have been found to be more successful in printing TZM, and those alternatives approaches are discussed in this effort. The efforts detailed in this report focused on continued development of Nb1Zr through L-PBF and development of TZM via L-PBF and electron powder bed fusion (E-PBF). The objective of this work was to further the development of these AM techniques for the chosen refractory alloys, elucidating and addressing associated challenges through characterization of several demonstration builds. At LANL, Nb1Zr builds were completed using an EOS M290 and M400 machines, and a refractory alloy-dedicated L-PBF system, the Xact Metal XM200G, was installed. The XM200G primary purpose was to do the Nb1Zr parameter development process; however, due to difficulties associated with the machine installation and qualification process, it was decided to pivot development to the larger M400 and M290 machines. Although the supply of Nb1Zr powder was limited, it was sufficient to generate sub-scale metallographic specimens for the purpose of parameter development. This was first accomplished on the EOS M400 then the M290 due to machine schedule availability. Further development of TZM has been initiated at the University of Texas El Paso (UTEP) under contract with LANL to use both a heated build envelope L-PBF machine and E-PBF machine that have been found in the literature to mitigate microcracking. UTEP was provided with TZM powder and build plates to support parallel TZM parameter development across both machines. As part of the contract, UTEP will also be conducting microstructural characterization once optimized process parameters have been identified. The optimized process parameters for each machine will be used to generate a series of metallographic, mechanical, and surface finish specimens for subsequent characterization and testing. In the next section, we provide a detailed discussion of the methodology used for investigating the feasibility of leveraging these alloys for use in advanced reactor applications.

36 MATERIALS SCIENCE↗

Modeling Data Movement Performance on Heterogeneous Architectures

The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.

97 MATHEMATICS AND COMPUTING↗

ExaSGD: 2022 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-407 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY22 the main objective of the Kernel Thrust was (i) provide sparse optimization solver that runs efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations, (ii) strengthen the reliability and increase the performance of the mixed-dense sparse (MDS) solver of HiOp for deployment on the FY22 target architectures, Summit and Crusher, and (iii) increase performance by improving the mathematical algorithm and refining the parallel MPI-based implementation of the coarse-grain parallel solver HiOp-PriDec for capabilities deployment on the FY22 target architectures, Summit and Crusher. This document presents the developments and contributions done by the Kernels Thrust Team in FY22 toward completion of the above-mentioned objectives. These contributions progressed along four main development (sub)thrusts: (1) Design and implementation of a sparse optimization solver for use on hardware accelerators; (2) Improvement of the mathematical algorithm and of the parallel implementation of HiOp-PriDec to ensure readiness and efficient coarse-grain parallelism for FY23 target exascale machine; and (3) Support Software and Application Development Thrusts of the exaSGD project in their deployment of the project’s software stack on AMD- and NVIDIA-based architectures. The development of the sparse optimization solver (thrust 1 above) was new in FY22 and resulted in a new sparse solver in HiOp (available as of version 0.6). The second development thrust was a continuation of the efforts from FY21 and improved the mathematical algorithm and the communication strategy of the HiOp-PriDec solver. The last developement thrust is a large collaborative effort. Namely, the project’s teams from multiple labs (LLNL, PNNL, ORNL, and NREL) performed large-scale demonstration of the ExaSGD software stack, namely the optimization solvers of HiOp interfaced with the modeling front-end ExaGO and the stochastic sampler PowerScenarios. These demonstration efforts solved large-scale instances of the SC-ACOPF challenge problem of medium network sizes (10, 000-bus system) and large number of contingencies on Summit (NVIDIA accelerators) and Crusher (AMD accelerators) systems at ORNL.

97 MATHEMATICS AND COMPUTING↗