Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel machines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Efficient Parallelization of Irregular Applications on GPU Architectures

With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.

Wang, Qihan↗

Designing workflows for materials characterization

Experimental science is enabled by the combination of synthesis, imaging, and functional characterization organized into evolving discovery loop. Synthesis of new material is typically followed by a set of characterization steps aiming to provide feedback for optimization or discover fundamental mechanisms. However, the sequence of synthesis and characterization methods and their interpretation, or research workflow, has traditionally been driven by human intuition and is highly domain specific. Here, we explore concepts of scientific workflows that emerge at the interface between theory, characterization, and imaging. In this study, we discuss the criteria by which these workflows can be constructed for special cases of multiresolution structural imaging and functional characterization, as a part of more general material synthesis workflows. Some considerations for theory–experiment workflows are provided. We further pose that the emergence of user facilities and cloud labs disrupts the classical progression from ideation, orchestration, and execution stages of workflow development. To accelerate this transition, we propose the framework for workflow design, including universal hyperlanguages describing laboratory operation, ontological domain matching, reward functions and their integration between domains, and policy development for workflow optimization. These tools will enable knowledge-based workflow optimization; enable lateral instrumental networks, sequential and parallel orchestration of characterization between dissimilar facilities; and empower distributed research.

36 MATERIALS SCIENCE↗

Machine Learning Driven Contouring of High-Frequency Four-Dimensional Cardiac Ultrasound Data

Automatic boundary detection of 4D ultrasound (4DUS) cardiac data is a promising yet challenging application at the intersection of machine learning and medicine. Using recently developed murine 4DUS cardiac imaging data, we demonstrate here a set of three machine learning models that predict left ventricular wall kinematics along both the endo- and epi-cardial boundaries. Each model is fundamentally built on three key features: (1) the projection of raw US data to a lower dimensional subspace, (2) a smoothing spline basis across time, and (3) a strategic parameterization of the left ventricular boundaries. Model 1 is constructed such that boundary predictions are based on individual short-axis images, regardless of their relative position in the ventricle. Model 2 simultaneously incorporates parallel short-axis image data into their predictions. Model 3 builds on the multi-slice approach of model 2, but assists predictions with a single ground-truth position at end-diastole. To assess the performance of each model, Monte Carlo cross validation was used to assess the performance of each model on unseen data. For predicting the radial distance of the endocardium, models 1, 2, and 3 yielded average R2 values of 0.41, 0.49, and 0.71, respectively. Monte Carlo simulations of the endocardial wall showed significantly closer predictions when using model 2 versus model 1 at a rate of 48.67%, and using model 3 versus model 2 at a rate of 83.50%. These finding suggest that a machine learning approach where multi-slice data are simultaneously used as input and predictions are aided by a single user input yields the most robust performance. Subsequently, we explore the how metrics of cardiac kinematics compare between ground-truth contours and predicted boundaries. We observed negligible deviations from ground-truth when using predicted boundaries alone, except in the case of early diastolic strain rate, providing confidence for the use of such machine learning models for rapid and reliable assessments of murine cardiac function. To our knowledge, this is the first application of machine learning to murine left ventricular 4DUS data. Future work will be needed to strengthen both model performance and applicability to different cardiac disease models.

4D ultrasound↗

RLGBS: Reinforcement Learning-Guided Beam Search for process optimization in a paper machine dryer section

Paper drying is responsible for over two-thirds of energy consumption in the U.S. pulp and paper industry, presenting significant potential for energy savings through optimization of process parameters. Current approaches often assume fixed operating conditions, neglecting dynamic ambient and process variations that limit achievable savings and real-world applicability. To this end, we develop a physics-based simulation environment for a paper machine dryer section and propose a reinforcement learning (RL) framework to minimize overall energy consumption by optimizing drying process parameters under diverse operating conditions. To mitigate overdrying and numerical instabilities caused by suboptimal local RL actions, we introduce Reinforcement Learning-Guided Beam Search (RLGBS), which explores multiple action sequences in parallel using beam search. Instead of making step-by-step decisions, RLGBS prioritizes solutions based on cumulative probability, reducing the impact of individual suboptimal actions. Experiments demonstrate that RLGBS achieves consistent energy savings under unseen operating conditions not encountered during training, outperforming conventional RL methods. While validated in drying optimization, this framework is broadly applicable to other RL-based industrial process control problems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Alignment of Au nanorods along de novo designed protein nanofibers studied with automated image analysis

In this paper, we focus on exploring the directional assembly of anisotropic Au nanorods along de novo designed 1D protein nanofiber templates. Using machine learning and automated image processing, we analyze scanning electron microscopy (SEM) images to study how the attachment density and alignment fidelity are influenced by variables such as the aspect ratio of the Au nanorods, and the salt concentration of the solution. We find that the Au nanorods prefer to align parallel to the protein nanofibers. This preference decreases with increasing salt concentration, but is only weakly sensitive to the nanorod aspect ratio. While the overall specific Au nanorod attachment density to the protein fibers increases with increasing solution ionic strength, this increase is dominated primarily by non-specific binding to the substrate background, and we find that greater specific attachment (nanorods attached to the nanofiber template as compared to the substrates) occurs at the lower studied salt concentrations, with the maximum ratio of specific to non-specific binding occurring when the protein fiber solutions are prepared in 75 mM NaCl concentration.

36 MATERIALS SCIENCE↗

MBX: A many-body energy and force calculator for data-driven many-body simulations

Many-Body eXpansion (MBX) is a C++ library that implements many-body potential energy functions (PEFs) within the “many-body energy” (MB-nrg) formalism. MB-nrg PEFs integrate an underlying polarizable model with explicit machine-learned representations of many-body interactions to achieve chemical accuracy from the gas to the condensed phases. MBX can be employed either as a stand-alone package or as an energy/force engine that can be integrated with generic software for molecular dynamics and Monte Carlo simulations. MBX is parallelized internally using Open Multi-Processing and can utilize Message Passing Interface when available in interfaced molecular simulation software. In this study, MBX enables classical and quantum molecular simulations with MB-nrg PEFs, as well as hybrid simulations that combine conventional force fields and MB-nrg PEFs, for diverse systems ranging from small gas-phase clusters to aqueous solutions and molecular fluids to biomolecular systems and metal-organic frameworks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

GLUE Code: A framework handling communication and interfaces between scales

Many scientific applications are inherently multiscale in nature. Such complex physical phenomena often require simultaneous execution and coordination of simulations spanning multiple time and length scales. This is possible by combining expensive small-scale simulations (such as molecular dynamics simulations) with larger scale simulations (such continuum limit/hydro solvers) to allow for considerably larger systems using task and data parallelism. However, the granularity of the tasks can be very large and often leads to load imbalance. Traditionally, we use approximations to streamline the computation of the more costly interactions and this introduces trade-offs between simulation cost and accuracy. In recent years, the available computational power and the advances in machine learning have made computing these scale-bridging interactions and multiscale simulations more feasible. One driving application has been plasma modeling in inertial confinement fusion (ICF), which is fundamentally multiscale in nature. This requires deep understanding of how to extrapolate microscopic information into macroscopically relevant scales. For example, in ICF one needs an accurate understanding of the connection between experimental observables and the underlying microphysics. The properties of the larger scales are often affected by the microscale behavior incorporated usually into the equations of state and ionic and electronic transport coefficients (Liboff, 1959; Rinderknecht et al., 2014; Rosenberg et al., 2015; Ross et al., 2017). Instead of incorporating this information using reliable molecular dynamics (MD) simulations, one often needs to use theoretical models, due to the inability of MD to reach engineering scales (Glosli et al., 2007; Marinak et al., 1998). One approach to resolve this issue is by coupling two MD simulations of different scales via force interpolation, e.g., the AdResS method (Krekeler et al., 2018; Nagarajan et al., 2013). Another approach, which we will pursue in the scope of this work, is by enabling scale bridging between MD simulations and meso/macro-scale models through the development and support of application programming interfaces that these different applications can interact through.

54 ENVIRONMENTAL SCIENCES↗

Scaling the SciDAC QuantOm Workflow

As part of the Scientific Discovery through Advanced Computing (SciDAC) program, the Quantum Chromodynamics Nuclear Tomography (QuantOM) project aims to analyze data from Deep Inelastic Scattering (DIS) experiments conducted at Jefferson Lab and the upcoming Electron Ion Collider. The DIS data analysis is performed on an event-level by combining the input from theoretical and experimental nuclear physics into a single, composable workflow. The optimization itself (I.e. fitting the experimental data with theoretical predictions) is carried out by a machine / deep learning algorithm. The size of the acquired DIS data as well as the complexity of the workflow itself require that the analysis is performed across multiple GPUs on high performance computing systems, such as Polaris at Argonne National Laboratory. This presentation discusses the novelties and challenges that came along with parallelizing this workflow. Recent results are compared to common distributed training techniques.

Lersch, Daniel↗

Performance Portability Evaluation of Fluid-Structure Interaction Simulations on Heterogeneous Platforms

The rapid proliferation of heterogeneous programming languages and multi-vendor hardware has underscored the critical need to evaluate the performance portability of scientific applications. In this work, we present the systematic porting and optimization of a massively parallel fluid-structure interaction code across multiple heterogeneous programming frameworks for deployment on leadership-class supercomputers from major vendors. Our analysis focuses on at-scale performance for simulations involving hundreds of millions of deformable cells, executed on a combination of CPUs and GPUs spanning thousands of nodes on exascale machines. We benchmark the performance of each implementation, highlighting the trade-offs inherent in adopting diverse programming models. Key insights regarding the portability of CUDA on multi-vendor platforms, the superior multi-core CPU performance from SYCL, and architectural considerations on performance optimization are distilled from our experience, offering guidance to other users of high performance computing based on our findings.

Martin, Aristotle [Duke University]↗

Human Factors for Advanced Reactors

Existing light water reactors in the U.S. are primarily large baseload electricity generating facilities. The concept of operations for these plants remains largely unchanged since the advent of commercial nuclear power—the main control room serves as the hub of plant activities and is staffed with multiple licensed operators who work in tandem under the shift supervisor, and staff such as field workers support the control room remotely. While newer plants have brought the advent of digital human-machine interfaces to replace earlier analog and mechanical instrumentation and controls, much of the control process remains unchanged and manual. It is simply a newer version of legacy concepts. Advanced reactors potentially bring considerable changes to the size, fuel type, automation, and staffing of nuclear power plants, necessitating a fundamental shift not just from analog to digital, but further from human to automation, from onsite to remote, from control to monitoring, and from many to few operators. Despite this multitude of parallel evolutions in reactor designs, many of the vendors developing the next generation of reactors represent smaller research and development enterprises. It is therefore not feasible to address all aspects of plant design at the same time. In particular, the competing design aspects of new reactors present a significant challenge to the development of robust and human factored systems at the plant. As vendors develop new reactor designs, much of the early focus is naturally on the fuel and reactor system technology. Looming behind these early advances is the daunting prospect of first-of-a-kind control concepts that have not yet been developed or validated. A failure to address the human element of reactor design early will lead to missed opportunities. The quickest development process is the replication of existing concepts of operations at legacy plants, even when such systems were long ago surpassed by better human-machine technologies outside the nuclear industry. Conversely, attempting to undertake novel concepts of operations late in the design life cycle of a plant could result in protracted development efforts and delays in licensing and deployment. This does not have to happen, and it is imperative that human factors be considered now, early in the design of new reactors.

99 GENERAL AND MISCELLANEOUS↗

Simulation of the Impact of Point Defects and Edge Dislocations on X-Ray Diffraction in Hexagonal (Ni,Co) 1+2 x Ti 1– x O 3 Thin Films

In this work, a computer code for simulating high-resolution X-ray diffraction (XRD) data from disordered crystals with arbitrary spatial composition and local lattice parameters is developed. Simulated patterns are compared with the experimental data collected on a single phase, highly crystalline (Ni 0.42 Co 0.58 ) 2.22 Ti 0.39 O 3 solid-solution thin film exhibiting a large number of subsidiary minima. As a case study, the edge dislocations in hexagonal (Ni,Co) 3 O 3 thin films with Burgers vector parallel and perpendicular to the a-axis and c-axis, respectively, are modeled. No peak profiles are assumed and thus issues related to profile fitting are avoided. Both macroscopic features, such as film thickness, and atomic-scale structure, such as dislocations, are simultaneously modeled. Simulations are run on a desktop machine. Commonly applied database-based phase identification routines can result in wrong phase identification, unless intensities are properly modeled. Modeling tools of diffractometer manufacturers are compared with the present approach. An example of how simulated reciprocal lattice patterns can be used to choose relevant measurement geometries in defected thin films is given.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Real-space solution to the electronic structure problem for nearly a million electrons

We report a Kohn–Sham density functional theory calculation of a system with more than 200 000 atoms and 800 000 electrons using a real-space high-order finite-difference method to investigate the electronic structure of large spherical silicon nanoclusters. Our system of choice was a 20 nm large spherical nanocluster with 202 617 silicon atoms and 13 836 hydrogen atoms used to passivate the dangling surface bonds. To speed up the convergence of the eigenspace, we utilized Chebyshev-filtered subspace iteration, and for sparse matrix–vector multiplications, we used blockwise Hilbert space-filling curves, implemented in the PARSEC code. For this calculation, we also replaced our orthonormalization + Rayleigh–Ritz step with a generalized eigenvalue problem step. We utilized all of the 8192 nodes (458 752 processors) on the Frontera machine at the Texas Advanced Computing Center. We achieved two Chebyshev-filtered subspace iterations, yielding a good approximation of the electronic density of states. Our work pushes the limits on the capabilities of the current electronic structure solvers to nearly 106 electrons and demonstrates the potential of the real-space approach to efficiently parallelize large calculations on modern high-performance computing platforms.

Chemistry↗

Development of Transformative Preparation Methods to Push up High Q&G Performance of FRIB Spare HWR Cryomodule Cavities

The FRIB accelerator project construction, a top priority of US nuclear science, was completed in January 2022, and is now moving to user operation. The stable and reliable operation of the accelerating cryomodules is essential in achieving/fulfilling DOE and user expectations. So far, FRIB cryomodules meet all FRIB specifications for cavity performance. However, during the lifetime of machine operation, degradation of cryomodule performance is possible, as reported in similar operating facilities (CEBAF, SNS). If cryomodule degradation is observed at FRIB, the under-performing cryomodule will require replacement/maintenance. In effort to manage operational reliability, FRIB plans to construct a 0.53 half-wave cryomodule to serve as an active spare. In a parallel effort, FRIB will also work toward increasing operational Q and gradient of spare cryomodule cavities to gain an overall performance margin to support future operational reliability. The current FRIB cavity designs have a potential to operate at gradients higher than 8 MV/m, but are currently limited by field emission (FE) and/or high field Q slope (HFQS); known issue in buffered chemical polished (BCP) treated cavities. The proposal looks to develop transformative surface preparation treatments to improve the operational gradient of spare cryomodules higher than 10 MV/m while maintaining high Q. Thus, increasing operational margin by 30 - 50%. With the goal to improve operational reliability set, the proposal will investigate multiple objectives as possible paths forward to achieve an overall increase in cavity performance and gain a better understanding of SRF limiting mechanisms. The proposal will study the application of different chemical surface treatments to 0.53 half-wave cavities, with the addition of low temperature bakes (LTB), and measure their effects on accelerating performance. Proposed chemical treatments to be explored in this proposal include conventional EP acid mixtures, as well as innovated EP and BCP acid mixtures designed to simplify processing paths in migrating FE and HFQS. The proposed transformative treatment wet N-doping also has the potential to replicate recent advancements in SRF technology relating to nitrogen doping and high Q operation without the requirement for an ultra-high vacuum annealing furnace; currently being developed at FNAL and JLAB. In parallel, high Q performance relating to flux trapping will be investigated with the installation of a second layer of magnetic shielding in the vertical test Dewar. The research objectives presented in the proposal, and their corresponding effects on cavity performance, will provide essential knowledge and future guidance to the SRF community and provide possible paths for future SRF based projects and applications.

43 PARTICLE ACCELERATORS↗

ExaCA: A performance portable exascale cellular automata application for alloy solidification modeling

Modeling the as-solidified grain structures that form during alloy processing is a critical component in understanding process-property relationships, particularly for additive manufacturing (AM) where grain structure is very sensitive to processing conditions. While cellular automata (CA)-based models have proven able to predict aspects of microstructure for several alloys and AM process conditions, long run times and large resource sets required limit the utility and the problem size to which existing CA models can be applied. As part of the ExaAM project, an initiative within the Exascale Computing Project (ECP) to develop, test, and optimize an exascale-capable coupled and self-consistent model of AM parts, we developed ExaCA (https://github.com/LLNL/ExaCA) for the liquid–solid phase transformation in the wake of AM melt pools. The CA-based code is parallelized using MPI and the Kokkos programming model, the latter enabling simulation on both CPUs and GPUs within a single-source implementation. Here, we detail the steps taken to transform a baseline, MPI-based CA code into one that is performant on CPUs and GPUs. Performance testing of ExaCA on Summit (a pre-exascale machine at Oak Ridge National Laboratory) was used to quantify CPU–GPU speedup comparing with equal numbers of nodes. Testing showed comparable CPU performance to the MPI-only CA code and a 5-20x speedup when running AM-based test problems using GPUs. The improved performance of CA through GPU utilization and the performance portable nature of ExaCA will enable accurate part-scale modeling by harnessing the power of current and future generations of high performance computing resources. Future work will include improving the strong scaling of ExaCA on GPUs by reducing load imbalance associated with the locality of the problem, and continuing performance optimization across exascale hardware.

36 MATERIALS SCIENCE↗

Learning plasma dynamics and robust rampdown trajectories with predict-first experiments at TCV

The rampdown phase of a tokamak pulse is difficult to simulate and often exacerbates multiple plasma instabilities. To reduce the risk of disrupting operations, we leverage advances in Scientific Machine Learning (SciML) to combine physics with data-driven models, developing a neural state-space model (NSSM) that predicts plasma dynamics during Tokamak à Configuration Variable (TCV) rampdowns. The NSSM efficiently learns dynamics from a modest dataset of 311 pulses with only five pulses in a reactor-relevant high-performance regime. The NSSM is parallelized across uncertainties, and reinforcement learning (RL) is applied to design trajectories that avoid instability limits. High-performance experiments at TCV show statistically significant improvements in relevant metrics. A predict-first experiment, increasing plasma current by 20% from baseline, demonstrates the NSSM’s ability to make small extrapolations. The developed approach paves the way for designing tokamak controls with robustness to considerable uncertainty and demonstrates the relevance of SciML for fusion experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Toward Large-Scale Image Segmentation on Summit

Semantic segmentation of images is an important computer vision task that emerges in a variety of application domains such as medical imaging, robotic vision and autonomous vehicles to name a few. While these domain-specific image analysis tasks involve relatively small image sizes (~ 10 2 × 10 2 ), there are many applications that need to train machine learning models on image data with extents that are orders of magnitude larger (~10 4 × 10 4 ). Training deep neural network (DNN) models on large extent images is extremely memory-intensive and often exceeds the memory limitations of a single graphical processing unit, a hardware accelerator of choice for computer vision workloads. Here, an efficient, sample parallel approach to train U-Net models on large extent image data sets is presented. Its advantages and limitations are analyzed and near-linear strong-scaling speedup demonstrated on 256 nodes (1536 GPUs) of the Summit supercomputer. Using a single node of the Summit supercomputer, an early evaluation of a recently released model parallel framework called GPipe is demonstrated to deliver ~ 2X speedup in executing a U-Net model with an order of magnitude larger number of trainable parameters than reported before. Performance bottlenecks for pipelined training of U-Net models are identified and mitigation strategies to improve the speedups are discussed. Together, these results open up the possibility of combining both approaches into a unified scalable pipelined and data parallel algorithm to efficiently train U-Net models with very large receptive fields on data sets of ultra-large extent images.

Seal, Sudip↗

Practical applications of machine-learned flows on gauge fields

Normalizing flows are machine-learned maps between different lattice theories which can be used as components in exact sampling and inference schemes. Ongoing work yields increasingly expressive flows on gauge fields, but it remains an open question how flows can improve lattice QCD at state-of-the-art scales. We discuss and demonstrate two applications of flows in replica exchange (parallel tempering) sampling, aimed at improving topological mixing, which are viable with iterative improvements upon presently available flows.

Abbott, Ryan↗

EdgeAI: Machine learning via direct attached accelerator for streaming data processing at high shot rate x-ray free-electron lasers

We present a case for low batch-size inference with the potential for adaptive training of a lean encoder model. We do so in the context of a paradigmatic example of machine learning as applied in data acquisition at high data velocity scientific user facilities such as the Linac Coherent Light Source-II x-ray Free-Electron Laser. We discuss how a low-latency inference model operating at the data acquisition edge can capitalize on the naturally stochastic nature of such sources. We simulate the method of attosecond angular streaking to produce representative results whereby simulated input data reproduce high-resolution ground truth probability distributions. By minimizing the mean-squared error between the decoded output of the latent representation and the ground truth distributions, we ensure that the encoding layers and resulting latent representation maintains full fidelity for any downstream task, be it classification or regression. We present throughput results for data-parallel inference of various batch sizes, some with throughput exceeding 100 k images per second. We also show in situ training below 10 s per epoch for the full encoder–decoder model as would be relevant for streaming and adaptive real-time data production at our nation’s scientific light sources.

97 MATHEMATICS AND COMPUTING↗