Neural Network Solver for Coherent Synchrotron Radiation Wakefield Calculations in Accelerator-Based Charged Particle Beams
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Machine learning is becoming prevalent in high energy physics, with numerous applications in physics analyses and event reconstruction showing great improvements compared to traditional computing methods. This thesis studies three projects which each propose new avenues for machine learning applications within the high energy physics CMS experiment located at CERN. In the first project, a search for a dark matter signal called “emerging jets” is performed, using graph neural networks to greatly increase sensitivity to the signal’s signature within the data. The result of this dark matter search sets the most stringent exclusion limits to date on theoretical emerging jet models. Motivated by inefficiencies encountered when processing the emerging jet graph neural network at Fermi National Accelerator Laboratory’s computing centers, the second project re-optimizes the computing centers for machine learning inference. This re-optimization uses NVIDIA Triton Inference Servers to process users’ analysis code heterogeneously, therefore achieving high processing throughput and decreasing user time-to-insight. The last project focuses on an upgrade to the CMS experiment’s real-time event selection system which improves physics object reconstruction under harsh processing conditions. A boosted decision tree is used to quickly and efficiently quantify a reconstructed particle’s “track quality” in order to remove particle tracks reconstructed erroneously. In summary, this thesis will not only present examples of how high energy physics can greatly benefit by leveraging machine learning techniques for physics analysis and reconstruction, but will also provide guidance on how the field can prepare for the inevitable increase in machine learning applications.
Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.
Abstract Field‐scale properties of fractured rocks play a crucial role in many subsurface applications, yet methodologies for identification of the statistical parameters of a discrete fracture network (DFN) are scarce. We present an inversion technique to infer two such parameters, fracture density and fractal dimension, from cross‐borehole thermal experiments data. It is based on a particle‐based heat‐transfer model, whose evaluation is accelerated with a deep neural network (DNN) surrogate that is integrated into a grid search. The DNN is trained on a small number of the heat‐transfer model runs and predicts the cumulative density function of the thermal field. The latter is used to compute fine posterior distributions of the (to be estimated) parameters. Our synthetic experiments reveal that fracture density is well constrained by data, while fractal dimension is harder to determine. Adding nonuniform prior information related to the DFN connectivity improves the inference of this parameter.
We develop a GPU-accelerated machine learning generative adversarial network model that can be used with observational data for the purpose of constructing causal inferences. The theoretical basis of our machine learning model is novel and is conceptualized to be operable and scalable for high performance computing platforms. Our GPU-accelerated code enables large-scale parallelization of the computation within a common and accessible computing environment. This will expand the reach of our model and empower research in new substantive domains while maintaining the underlying theoretical properties.
The principal objective of this research effort was to demonstrate the extraordinarily cost effective acceleration of finite element structural analysis problems using a transputer-based parallel processing network. This objective was accomplished in the form of a commercially viable parallel processing workstation. The workstation is a desktop size, low-maintenance computing unit capable of supercomputer performance yet costs two orders of magnitude less. To achieve the principal research objective, a transputer based structural analysis workstation termed XPFEM was implemented with linear static structural analysis capabilities resembling commercially available NASTRAN. Finite element model files, generated using the on-line preprocessing module or external preprocessing packages, are downloaded to a network of 32 transputers for accelerated solution. The system currently executes at about one third Cray X-MP24 speed but additional acceleration appears likely. For the NASA selected demonstration problem of a Space Shuttle main engine turbine blade model with about 1500 nodes and 4500 independent degrees of freedom, the Cray X-MP24 required 23.9 seconds to obtain a solution while the transputer network, operated from an IBM PC-AT compatible host computer, required 71.7 seconds. Consequently, the $80,000 transputer network demonstrated a cost-performance ratio about 60 times better than the $15,000,000 Cray X-MP24 system.
The proposed system is written in the hardware description language (HDL) verilog targeting an FPGA board. It is intended as a testbed to explore architecture tradeoffs in multi-tiled heterogeneous architectures. Although we target FPGAs, the system can be implemented as a monolithic SoC or a package comprised of many chiplets that are interconnected in the same package using a NoC. The proposed NoC is lightweight and follows an axi-lite interface. The endpoints of the NoC are a heterogeneous mix of "tiles" as endpoints that are general purpose processors, fixed function accelerators, and programmable accelerators. We assume that the network interfaces for the NoC endpoints are all addressable in a global name-space in that they represent an address range (for memory addresses) or a range of unique identifiers that are associated with each individual tile. This makes the functionality abstract from the standpoint of the NoC design details. Message queues offer a direct inter-processor interface between peer general purpose cores and diverse accelerators that comprise an SoC. Although they share the same NoC infrastructure for inter-tile communication within an SoC or SiP, the hardware message queues bypass the memory hierarchy and thus do not pollute the memory state or invoke the cache coherence mechanism.
We explore using neural operators, or neural network representations of nonlinear maps between function spaces, to accelerate infinite-dimensional Bayesian inverse problems (BIPs) with models governed by nonlinear parametric partial differential equations (PDEs). Neural operators have gained significant attention in recent years for their ability to approximate the parameter-to-solution maps defined by PDEs using as training data solutions of PDEs at a limited number of parameter samples. The computational cost of BIPs can be drastically reduced if the large number of PDE solves required for posterior characterization are replaced with evaluations of trained neural operators. However, reducing error in the resulting BIP solutions via reducing the approximation error of the neural operators in training can be challenging and unreliable. We provide an a priori error bound result that implies certain BIPs can be ill-conditioned to the approximation error of neural operators, thus leading to inaccessible accuracy requirements in training. To reliably deploy neural operators in BIPs, we consider a strategy for enhancing the performance of neural operators: correcting the prediction of a trained neural operator by solving a linear variational problem based on the PDE residual. We show that a trained neural operator with error correction can achieve a quadratic reduction of its approximation error, all while retaining substantial computational speedups of posterior sampling when models are governed by highly nonlinear PDEs. The strategy is applied to two numerical examples of BIPs based on a nonlinear reaction–diffusion problem and deformation of hyperelastic materials. We demonstrate that posterior representations of the two BIPs produced using trained neural operators are greatly and consistently enhanced by error correction.
Molecular Dynamics (MD) simulations are used to understand the effects of corrosion on metallic materials in salt brine. Reactive force fields in classical MD enable accurate modeling of bond formation and breakage in the aqueous medium and at the metal-electrolyte interface, while also facilitating dynamic partial charge equilibration. However, MD simulations are computationally intensive and unsuitable for modeling the long time scales characteristic of corrosive phenomena. To address this, we develop reduced-order machine learning models that provide accurate and efficient predictions of charge density in corrosive environments. Specifically, we use Long Short-Term Memory (LSTM) networks to forecast charge density evolution based on atomic environments represented by Smooth Overlap of Atomic Positions (SOAP) descriptors. A physics-informed loss function enforces charge neutrality and electronegativity equivalence. The atomic charges predicted by the deep learning model trained on this work were obtained two orders of magnitude faster than those from molecular dynamics (MD) simulations, with an error of less than 3% compared to the MD-obtained charges, even in extrapolative scenarios, while adhering to physical constraints. This demonstrates the excellent accuracy, computational efficiency, and validity of the developed model. Lastly, even though developed for corrosion, these protocols are formulated in a phenomenon-agnostic manner, allowing application to various variable-charge interatomic potentials and related fields.
Explore the source record for details and available documents.
Synaptic devices with linear high-speed switching can accelerate learning in artificial neural networks (ANNs) embodied in hardware. Conventional resistive memories however suffer from high write noise and asymmetric conductance tuning, preventing parallel programming of ANN arrays. Electrochemical random-access memories (ECRAMs), where resistive switching occurs by ion insertion into a redox-active channel, aim to address these challenges due to their linear switching and low noise. ECRAMs using 2D materials and metal oxides however suffer from slow ion kinetics, whereas organic ECRAMs enable high-speed operation but face challenges toward on-chip integration due to poor temperature stability of polymers. Here, ECRAMs using 2D titanium carbide (Ti 3 C 2 T x ) MXene that combine the high speed of organics and the integration compatibility of inorganic materials in a single high-performance device are demonstrated. These ECRAMs combine the speed, linearity, write noise, switching energy, and endurance metrics essential for parallel acceleration of ANNs, and importantly, they are stable after heat treatment needed for back-end-of-line integration with Si electronics. The high speed and performance of these ECRAMs introduces MXenes, a large family of 2D carbides and nitrides with more than 30 stoichiometric compositions synthesized to date, as promising candidates for devices operating at the nexus of electrochemistry and electronics.
Computational analysis of countercurrent flows in packed absorption columns, often used in solvent-based post-combustion carbon capture systems (CCSs), is challenging. Typically, computational fluid dynamics (CFD) approaches are used to simulate the interactions between a solvent, gas, and column's packing geometry while accounting for the thermodynamics, kinetics, heat, and mass transfer effects of the absorption process. These simulations can then be used explain a column's hydrodynamic characteristics and evaluate its CO 2 -capture efficiency. However, these approaches are computationally expensive, making it difficult to evaluate numerous designs and operating conditions to improve efficiency at industrial scales. In this work, we comprehensively explore the application of statistical ML methods, convolutional neural networks (CNNs), and graph neural networks (GNNs) to aid and accelerate the scale-up and design optimization of solvent-based post-combustion CCSs. We apply these methods to CFD datasets of countercurrent flows in absorption columns with structured packings characterized by several geometric parameters. We train models to use these parameters, inlet velocity conditions, and other model-specific representations of the column to estimate key determinants of CO 2 -capture efficiency without having to simulate additional CFD datasets. We also evaluate the impact of different input types on the accuracy and generalizability of each model. We discuss the strengths and limitations of each approach to further elucidate the role of CNNs, GNNs, and other machine learning approaches for CO 2 -capture property prediction and design optimization.
SmartNICs, accelerator devices integrated with a network, have conventionally been utilized to offload low-level networking functionality. However, newer SmartNIC variants, which incorporate a system-on-chip (SoC) with traditional designs, are challenging this precedent. Leveraging significantly augmented resources, these new devices offer increased versatility and the potential to more effectively complement a given architecture’s CPU. Naturally, there is performance overhead associated with offloading tasks from a host CPU to a SmartNIC. As SmartNICs become more versatile and are used for a wider range of computational tasks, it is critical to understand how well the SmartNICs perform those respective tasks in order to determine whether the cost of offloading is worthwhile. In this work, we lay out a path to understand the important question of when and when not to offload to SmartNIC devices by way of a series of microbenchmarks.
Large-scale, low -cost hydrogen production can enable an economically competitive, secure, and environmentally beneficial future energy system across multiple sectors. Furthermore, clean hydrogen can address specific sectors that are hard to decarbonize (e.g., heavy-duty trucking, load-following electricity, iron, steel, and cement) and can help the U.S. meet the net zero carbon goal by 2050. To achieve this goal, tens of millions of metric tons of clean, reliable, and affordable hydrogen will be needed annually1. In 2021, the Hydrogen Energy Earthshot was launched, and its goal is to reduce the cost of clean hydrogen to $1 per $1 kilogram in 1 decade (1 1 1) 2. One very promising pathway for large-scale hydrogen production is water splitting. Water splitting technologies range from commercial technologies such as electrolyzers to approaches that are at a much earlier stage of development, such as photoelectrochemical (PEC) and thermochemical (TCH) processes. All these water splitting pathways offer diverse benefits in energy storage, grid services, and cross-sector emissions reductions while taking advantage of the diverse domestic resources. However, critical materials-, component- and system-level challenges must be addressed to improve efficiency and durability and reduce cost. To address these barriers and move these promising and high impact technologies forward, the HydroGEN Advanced Water Splitting Materials (AWSM) and the H2 from the Next-generation of Electrolyzers of Water (H2NEW) consortia were formed and supported by the Department of Energy (DOE) EERE Hydrogen and Fuel Cell Technologies Office (HFTO). HydroGEN (https://www.energy.gov/eere/h2awsm/) consortium, established in 2016, is an Energy Materials Network (EMN) that aims to accelerate the materials R&D of low technology readiness level (TRL) advanced water splitting (AWS) technologies. The consortium comprises five core national laboratories and focuses on four early-stage AWS pathways: alkaline exchange membrane (AEM) electrolysis, proton conducting solid oxide electrolysis (p-SOEC), photoelectrochemical, and thermochemical water splitting. Liquid alkaline and PEM electrolyzers are already commercial and significant advancements in oxygen conducting solid oxide electrolysis cells (o-SOECs) have been realized. Yet, these systems are still too expensive and not sufficiently durable for wide-scale commercialization. To enable high-volume manufacturing of affordable, durable, efficient electrolyzers, H2NEW (https://h2new.energy.gov/), another multi-lab consortium, was established in 2020. This comprehensive, concerted effort is focused on overcoming barriers related to components and materials integration and scale-up to achieve performance, durability, with an initial focus to achieve $2/kg H2 by 2026.
Globally distributed computing infrastructures, such as clouds and supercomputers, are currently used to manage data that is generated with an unprecedented speed from a variety of resources. Coping with this trend, the volume of data exchanged across distant sites increases substantially. To accelerate data transfer, high-speed networks are provided to connect remote sites. Most existing data movement solutions are optimized for moving large files. However, it is still challenging to transfer a large number of small files across networks. This disadvantage not only lowers data transfer performance, but also decreases overall system utilization. Here, we identify that moving small files is mainly constrained by degraded file system throughput, not just network performance as might be suspected. We have built a data transfer pipeline model to analyze the impact of small network I/O and storage I/O on data movement. Extending one of the widely used open source data movement solutions, GridFTP, we demonstrate several appropriate engineering approaches that mitigate the bottleneck and increase data transfer efficiency. We show optimizations that improve data transfer performance more than 5 times. In comparison to existing solutions, our approaches can save a significant amount of system resources for moving lots of small files.
X-ray computed tomography (XCT) is an important tool for high-resolution non-destructive characterization of additively-manufactured metal components. XCT reconstructions of metal components may have beam hardening artifacts such as cupping and streaking which makes reliable detection of flaws and defects challenging. Furthermore, traditional workflows based on using analytic reconstruction algorithms require a large number of projections for accurate characterization - leading to longer measurement times and hindering the adoption of XCT for in-line inspections. In this paper, we introduce a new workflow based on the use of two neural networks to obtain high-quality accelerated reconstructions from sparse-view XCT scans of single material metal parts. The first network, implemented using fully-connected layers, helps reduce the impact of BH in the projection data without the need of any calibration or knowledge of the component material. The second network, a convolutional neural network, maps a low-quality analytic 3D reconstruction to a high-quality reconstruction. Using experimental data, we demonstrate that our method robustly generalizes across several alloys, and for a range of sparsity levels without any need for retraining the networks thereby enabling accurate and fast industrial XCT inspections.