Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Building a Trusted Roaming Hub [Slides]

The Trusted Roaming Hub is a U.S. Department of Energy-backed initiative led by the National Laboratory of the Rockies (NLR) to address one of the most persistent challenges in electric vehicle (EV) charging: fragmented roaming, inconsistent interoperability, and insufficient digital trust across charging networks. As EV adoption accelerates and charging infrastructure scales nationwide, today's many-to-many integration model between eMobility Service Providers (eMSPs) and Charge Point Operators (CPOs) has become increasingly brittle, costly, and difficult to secure. The Trusted Roaming Hub introduces a neutral, cybersecurity-forward "switchboard" architecture that enables standardized, secure, and scalable roaming interactions across the EV charging ecosystem. Rather than replacing existing networks or commercial relationships, the hub acts as a trusted intermediary that enforces consistent identity, authentication, authorization, and routing across participants improving reliability for drivers, lowering integration burden for industry, and creating a foundation for future grid-interactive charging services. This read-ahead provides an overview of the problem the hub is designed to solve, the core functional and security concepts behind the architecture, the value proposition to key stakeholders, and the near-term trajectory of the work.

33 ADVANCED PROPULSION SYSTEMS↗

An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog Memory

In this work, we demonstrate SONOS (silicon-oxide-nitrideoxide- silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a >10× gain in energy efficiency over state-of-the-art digital and analog inference accelerators.

97 MATHEMATICS AND COMPUTING↗

Voucher Opportunity 5-15: Independent Assessment of Monitoring, Reporting, and Verification (MRV) Technologies and Practices for Enhanced Rock Weathering (CRADA 718) Abstract

Development of robust, transparent, and precise monitoring, reporting, and verification (MRV) technologies and practices is critical for carbon dioxide removal (CDR) project developers to comply with regulatory and permitting requirements, voluntary carbon market (VCM) protocols, and to ensure safety while reducing environmental impacts. Enhanced rock weathering (ERW)-based CDR technologies focus on removing atmospheric carbon through conversion into thermodynamically stable solid or aqueous carbonate forms for permanent storage (i.e., mineralization). This highly durable form of CDR enhances naturally occurring silicate rock weathering cycles by optimizing application of finely-ground silicate rock particles (i.e., from basalt) on terrestrial agricultural lands to accelerate natural silicate rock weathering and mineralization. Enhanced rock weathering may also provide improved crop yields and enhance soil health. A critical aspect for commercialization of these technologies is the development of MRV to quantify the net removal and durable storage of atmospheric CO 2 . For ERW systems, it is essential to accurately characterize the mineral feedstock selected for application to establish the baseline geochemical composition, mineral dissolution rates, and carbon removal potential of the feedstocks to estimate overall net removal. Given the difficulty with conducting MRV for ERW in diverse soil/environment types, over large application areas, and due to complex chemical reaction networks, this project will accelerate understanding towards consensus on best practices for MRV. The overall objectives of the proposed voucher project are to: 1) Characterize and analyze feedstock(s) intended for ERW field application by Lithos Carbon (“Voucher Recipient”/ “CRADA Participant”) to determine overall mineralization potential; 2) Facilitate knowledge transfer and documentation of experimental protocols, instrumentation, and other relevant best practices; and 3) Support the Voucher Recipient’s broader technology commercialization and ERW Research Facility development plans. This work will align with the Voucher Recipient’s MRV plans for field sites and build upon complementary efforts conducted by PNNL on mineralization MRV.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Digital twin framework for PIP-II linac: AI-driven multi-scale modeling from ion source to 800 MeV

The PIP-II superconducting linac at Fermilab is designed to deliver multi-megawatt proton beams for neutrino physics and other high-intensity applications. To expedite commissioning and enhance operational reliability, we have developed an EPICS-based data flow framework that seamlessly integrates digital twins (DT) with physical twins (PT). These digital twins comprise high-fidelity beam dynamics models or data-driven surrogate models connected to their physical counterparts through real-time diagnostics and advanced machine-learning algorithms.Central to this framework is Linac_Gen, an accelerated simulation tool that incorporates convolutional neural networks, random forests, and genetic algorithms to provide up to a tenfold speedup in optimizing the accelerator geometry model. An EPICS translator layer ensures interoperability by efficiently mapping lattice parameters across diverse simulation platforms.Our EPICS-based framework supports multiple operational modes—monitoring, passive learning, closed-loop control, and online learning—covering the entire machine lifecycle. By leveraging HPC resources and multi-objective optimization techniques, the digital twin enables adaptive trajectory correction, real-time fault detection, and predictive modeling of beam stability. This comprehensive approach paves the way for robust, high-intensity operation and data-driven accelerator R&D at Fermilab.

Pathak, Abhishek [Fermilab]↗

A neural network for beam background decomposition in Belle II at SuperKEKB

Here, we describe a neural network for predicting the background hit rate in the Belle II detector produced by the SuperKEKB electron-positron collider. The neural network, BGNet, learns to predict the individual contributions of different physical background sources, such as beam-gas scattering or continuous top-up injections into the collider, to Belle II sub-detector rates. The samples for learning are archived 1 Hz time series of diagnostic variables from the SuperKEKB collider subsystems and measured hit rates of Belle II used as regression targets. We test the learned model by predicting detector hit rates on archived data from different run periods not used during training. We show that a feature attribution method can help interpret the source of changes in the background level over time.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Bayesian calibration of bubble size dynamics applied to CO 2 gas fermenters

To accelerate the scale-up of gaseous CO 2 fermentation reactors, computational models need to predict gas-to-liquid mass transfer which requires capturing the bubble size dynamics, i.e. bubble breakup and coalescence. However, the applicability of existing models beyond air–water mixtures remains to be established. Here, an inverse modeling approach, accelerated with a neural network surrogate, calibrates the breakup and coalescence closure models, that are used in class methods for population balance modeling (PBM). The calibration is performed based on experimental results obtained in a CO 2 -air–water-coflowing bubble column reactor. Bayesian inference is used to account for noise in the experimental dataset and bias in the simulation results. To accurately capture gas holdup and interphase mass transfer, the results show that the breakage rate needs to be increased by one order of magnitude. In conclusion, the inferred model parameters are then used on a separate configuration and shown to also improve bubble size distribution predictions.

09 BIOMASS FUELS↗

Artificial intelligence inferred microstructural properties from voltage–capacity curves

Abstract The quantification of microstructural properties to optimize battery design and performance, to maintain product quality, or to track the degradation of LIBs remains expensive and slow when performed through currently used characterization approaches. In this paper, a convolution neural network-based deep learning approach (CNN) is reported to infer electrode microstructural properties from the inexpensive, easy to measure cell voltage versus capacity data. The developed framework combines two CNN models to balance the bias and variance of the overall predictions. As an example application, the method was demonstrated against porous electrode theory-generated voltage versus capacity plots. For the graphite|LiMn $$_2$$ 2 O $$_4$$ 4 chemistry, each voltage curve was parameterized as a function of the cathode microstructure tortuosity and area density, delivering CNN predictions of Bruggeman’s exponent and shape factor with 0.97 $$R^2$$ R 2 score within 2 s each, enabling to distinguish between different types of particle morphologies, anisotropies, and particle alignments. The developed neural network model can readily accelerate the processing-properties-performance and degradation characteristics of the existing and emerging LIB chemistries.

25 ENERGY STORAGE↗

XploreNAS : Explore Adversarially Robust and Hardware-efficient Neural Architectures for Non-ideal Xbars

Compute In-Memory platforms such as memristive crossbars are gaining focus as they facilitate acceleration of Deep Neural Networks (DNNs) with high area and compute efficiencies. However, the intrinsic non-idealities associated with the analog nature of computing in crossbars limits the performance of the deployed DNNs. Furthermore, DNNs are shown to be vulnerable to adversarial attacks leading to severe security threats in their large-scale deployment. Thus, finding adversarially robust DNN architectures for non-ideal crossbars is critical to the safe and secure deployment of DNNs on the edge. This work proposes a two-phase algorithm-hardware co-optimization approach called XploreNAS that searches for hardware efficient and adversarially robust neural architectures for non-ideal crossbar platforms. We use the one-shot Neural Architecture Search approach to train a large Supernet with crossbar-awareness and sample adversarially robust Subnets therefrom, maintaining competitive hardware efficiency. Our experiments on crossbars with benchmark datasets (SVHN, CIFAR10, CIFAR100) show up to ~8–16% improvement in the adversarial robustness of the searched Subnets against a baseline ResNet-18 model subjected to crossbar-aware adversarial training. We benchmark our robust Subnets for Energy-Delay-Area-Products (EDAPs) using the Neurosim tool and find that with additional hardware efficiency–driven optimizations, the Subnets attain ~1.5–1.6× lower EDAPs than ResNet-18 baseline.

97 MATHEMATICS AND COMPUTING↗

Emerging Jets Search, Triton Server Deployment, and Track Quality Development: Machine Learning Applications in High Energy Physics

Machine learning is becoming prevalent in high energy physics, with numerous applications in physics analyses and event reconstruction showing great improvements compared to traditional computing methods. This thesis studies three projects which each propose new avenues for machine learning applications within the high energy physics CMS experiment located at CERN. In the first project, a search for a dark matter signal called “emerging jets” is performed, using graph neural networks to greatly increase sensitivity to the signal’s signature within the data. The result of this dark matter search sets the most stringent exclusion limits to date on theoretical emerging jet models. Motivated by inefficiencies encountered when processing the emerging jet graph neural network at Fermi National Accelerator Laboratory’s computing centers, the second project re-optimizes the computing centers for machine learning inference. This re-optimization uses NVIDIA Triton Inference Servers to process users’ analysis code heterogeneously, therefore achieving high processing throughput and decreasing user time-to-insight. The last project focuses on an upgrade to the CMS experiment’s real-time event selection system which improves physics object reconstruction under harsh processing conditions. A boosted decision tree is used to quickly and efficiently quantify a reconstructed particle’s “track quality” in order to remove particle tracks reconstructed erroneously. In summary, this thesis will not only present examples of how high energy physics can greatly benefit by leveraging machine learning techniques for physics analysis and reconstruction, but will also provide guidance on how the field can prepare for the inevitable increase in machine learning applications.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum

Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.

97 MATHEMATICS AND COMPUTING↗

Thermal Experiments for Fractured Rock Characterization: Theoretical Analysis and Inverse Modeling

Abstract Field‐scale properties of fractured rocks play a crucial role in many subsurface applications, yet methodologies for identification of the statistical parameters of a discrete fracture network (DFN) are scarce. We present an inversion technique to infer two such parameters, fracture density and fractal dimension, from cross‐borehole thermal experiments data. It is based on a particle‐based heat‐transfer model, whose evaluation is accelerated with a deep neural network (DNN) surrogate that is integrated into a grid search. The DNN is trained on a small number of the heat‐transfer model runs and predicts the cumulative density function of the thermal field. The latter is used to compute fine posterior distributions of the (to be estimated) parameters. Our synthetic experiments reveal that fracture density is well constrained by data, while fractal dimension is harder to determine. Adding nonuniform prior information related to the DFN connectivity improves the inference of this parameter.

Zhou, Zitong↗

A GPU-Accelerated Population Generation, Sorting, and Mutation Kernel for an Optimization-Based Causal Inference Model

We develop a GPU-accelerated machine learning generative adversarial network model that can be used with observational data for the purpose of constructing causal inferences. The theoretical basis of our machine learning model is novel and is conceptualized to be operable and scalable for high performance computing platforms. Our GPU-accelerated code enables large-scale parallelization of the computation within a common and accessible computing environment. This will expand the reach of our model and empower research in new substantive domains while maintaining the underlying theoretical properties.

Cho, Wendy K. Tam↗

P38 heterogeneous multi-tiled system with support for message queues (MoSAIC) v0.1

The proposed system is written in the hardware description language (HDL) verilog targeting an FPGA board. It is intended as a testbed to explore architecture tradeoffs in multi-tiled heterogeneous architectures. Although we target FPGAs, the system can be implemented as a monolithic SoC or a package comprised of many chiplets that are interconnected in the same package using a NoC. The proposed NoC is lightweight and follows an axi-lite interface. The endpoints of the NoC are a heterogeneous mix of "tiles" as endpoints that are general purpose processors, fixed function accelerators, and programmable accelerators. We assume that the network interfaces for the NoC endpoints are all addressable in a global name-space in that they represent an address range (for memory addresses) or a range of unique identifiers that are associated with each individual tile. This makes the functionality abstract from the standpoint of the NoC design details. Message queues offer a direct inter-processor interface between peer general purpose cores and diverse accelerators that comprise an SoC. Although they share the same NoC infrastructure for inter-tile communication within an SoC or SiP, the hardware message queues bypass the memory hierarchy and thus do not pollute the memory state or invoke the cache coherence mechanism.

Gonzalez, LouisaPatricia↗

Residual-based error correction for neural operator accelerated infinite-dimensional Bayesian inverse problems

We explore using neural operators, or neural network representations of nonlinear maps between function spaces, to accelerate infinite-dimensional Bayesian inverse problems (BIPs) with models governed by nonlinear parametric partial differential equations (PDEs). Neural operators have gained significant attention in recent years for their ability to approximate the parameter-to-solution maps defined by PDEs using as training data solutions of PDEs at a limited number of parameter samples. The computational cost of BIPs can be drastically reduced if the large number of PDE solves required for posterior characterization are replaced with evaluations of trained neural operators. However, reducing error in the resulting BIP solutions via reducing the approximation error of the neural operators in training can be challenging and unreliable. We provide an a priori error bound result that implies certain BIPs can be ill-conditioned to the approximation error of neural operators, thus leading to inaccessible accuracy requirements in training. To reliably deploy neural operators in BIPs, we consider a strategy for enhancing the performance of neural operators: correcting the prediction of a trained neural operator by solving a linear variational problem based on the PDE residual. We show that a trained neural operator with error correction can achieve a quadratic reduction of its approximation error, all while retaining substantial computational speedups of posterior sampling when models are governed by highly nonlinear PDEs. The strategy is applied to two numerical examples of BIPs based on a nonlinear reaction–diffusion problem and deformation of hyperelastic materials. We demonstrate that posterior representations of the two BIPs produced using trained neural operators are greatly and consistently enhanced by error correction.

97 MATHEMATICS AND COMPUTING↗

Accelerating charge estimation in molecular dynamics simulations using physics-informed neural networks: corrosion applications

Molecular Dynamics (MD) simulations are used to understand the effects of corrosion on metallic materials in salt brine. Reactive force fields in classical MD enable accurate modeling of bond formation and breakage in the aqueous medium and at the metal-electrolyte interface, while also facilitating dynamic partial charge equilibration. However, MD simulations are computationally intensive and unsuitable for modeling the long time scales characteristic of corrosive phenomena. To address this, we develop reduced-order machine learning models that provide accurate and efficient predictions of charge density in corrosive environments. Specifically, we use Long Short-Term Memory (LSTM) networks to forecast charge density evolution based on atomic environments represented by Smooth Overlap of Atomic Positions (SOAP) descriptors. A physics-informed loss function enforces charge neutrality and electronegativity equivalence. The atomic charges predicted by the deep learning model trained on this work were obtained two orders of magnitude faster than those from molecular dynamics (MD) simulations, with an error of less than 3% compared to the MD-obtained charges, even in extrapolative scenarios, while adhering to physical constraints. This demonstrates the excellent accuracy, computational efficiency, and validity of the developed model. Lastly, even though developed for corrosion, these protocols are formulated in a phenomenon-agnostic manner, allowing application to various variable-charge interatomic potentials and related fields.

Atomistic models↗

High-Speed Ionic Synaptic Memory Based on 2D Titanium Carbide MXene

Synaptic devices with linear high-speed switching can accelerate learning in artificial neural networks (ANNs) embodied in hardware. Conventional resistive memories however suffer from high write noise and asymmetric conductance tuning, preventing parallel programming of ANN arrays. Electrochemical random-access memories (ECRAMs), where resistive switching occurs by ion insertion into a redox-active channel, aim to address these challenges due to their linear switching and low noise. ECRAMs using 2D materials and metal oxides however suffer from slow ion kinetics, whereas organic ECRAMs enable high-speed operation but face challenges toward on-chip integration due to poor temperature stability of polymers. Here, ECRAMs using 2D titanium carbide (Ti 3 C 2 T x ) MXene that combine the high speed of organics and the integration compatibility of inorganic materials in a single high-performance device are demonstrated. These ECRAMs combine the speed, linearity, write noise, switching energy, and endurance metrics essential for parallel acceleration of ANNs, and importantly, they are stable after heat treatment needed for back-end-of-line integration with Si electronics. The high speed and performance of these ECRAMs introduces MXenes, a large family of 2D carbides and nitrides with more than 30 stoichiometric compositions synthesized to date, as promising candidates for devices operating at the nexus of electrochemistry and electronics.

2D materials↗