Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fast machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Machine learning of 27Al NMR electric field gradient tensors for crystalline structures from DFT

NMR crystallography has emerged as a promising technique for the determination and refinement of atomic coordinates in crystal structures. The crystal structure of compounds containing quadrupolar nuclei, such as 27Al, can be improved by directly comparing solid-state NMR measurements to DFT computations of the electric field gradient (EFG) tensor. The non-negligible computational cost of these first-principles calculations limits the applicability of this method to all but the most well-defined structures. We developed a fast, low-cost machine learning model to predict EFG parameters based on local structural motifs and elemental parameters. We computed 8081 EFG tensors from 1681 27Al crystalline solids using DFT and benchmarked them against 105 experimentally measured 27Al sites. Surprisingly, simple local geometric features dominate the predictive performance of the resulting random-forest model, yielding an R2 value of 0.98 and an RMSE of 0.61 MHz for CQ, the quadrupolar coupling constant. This model accuracy should enable pre-refining future structural assignments before finally validating with first-principles calculations. Such a catalogue of 27Al NMR tensors can serve as a tool for researchers assigning complex NMR spectra influenced by the nuclear electric quadrupole interaction.

Sun, He↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗

Fast Assessment of Metal Performance through Dislocation Physics and Machine Learning

The microstructure of metals is key to their mechanical properties. The types, density, composition and morphology of crystal defects all have pronounced impact on the properties. Changes to the microstructure occurring during processing and use can be very striking. The emerging technology additive manufacturing (AM) has the potential to improve performance by allowing optimized designs, but the process and environments can lead to unusual microscale features whose properties must be understood and characterized to enable higher technological readiness levels and application. Experimentally, an extensive evaluation of mechanical properties of 3D printed metals is a challenge, and anomalous effects related to the AM process add complexity. We present a new machine learning (ML) model predicting mechanical response based on dislocation mediated plasticity simulations. A large set of 3D discrete dislocation dynamics simulations with wide ranges of loading conditions is transformed to preprocessed data ready for training with the ML model. The trained model can predict the mechanical response of Mo30W for a given microstructure evolution, providing key information essential for optimization of AM processing.

Jaehyun Cho↗

Monte Carlo Dropout Uncertainty Quantification of Long Short-Term Memory Autoencoder Anomaly Detection in a Liquid Sodium Cold Trap

Advanced high-temperature fluid reactors, such as sodium-cooled fast reactors (SFRs) and molten salt–cooled reactors (MSCRs), require coolant purification systems to prevent fluid contamination and local freezing that can lead to plugging. Liquid sodium purification can be achieved with a cold trap, where the sodium temperature is reduced to a near-freezing point to precipitate out impurities. Automation of monitoring of the cold trap performance with machine learning algorithms can aid in early detection of incipient anomalies. An efficient approach to loss-of-coolant–type anomaly detection in a cold trap monitored with more than two dozen thermal-hydraulic sensors consists of a long short-term memory (LSTM) autoencoder. This work develops the uncertainty quantification of the LSTM autoencoder performance for cold trap anomaly detection using the Monte Carlo (MC) dropout method. The MC dropout methodology creates a distribution of sister distributions that all slightly differ from each other because of random neurons being turned off for testing. The variances of the sister network distributions are used to make an uncertainty interval. Our analysis shows that the uncertainty in the autoencoder performance is largest near the peak of the anomaly signal. Using the MC dropout method, we investigate the uncertainty in the anomaly detection with missing sensor inputs. This capability allows the reactor operator to evaluate resilience of the anomaly detection system and to make informed decisions about continuity of operation in the event of sensor failure.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

MHD mode tracking using high-speed cameras and deep learning

Abstract We present a new algorithm to track the amplitude and phase of rotating magnetohydrodynamic (MHD) modes in tokamak plasmas using high speed imaging cameras and deep learning. This algorithm uses a convolutional neural network (CNN) to predict the amplitudes of the n = 1 sine and cosine mode components using solely optical measurements from one or more cameras. The model was trained and tested on an experimental dataset consisting of camera frame images and magnetic-based mode measurements from the High Beta Tokamak - Extended Pulse (HBT-EP) device, and it outperformed other, more conventional, algorithms using identical image inputs. The effect of different input data streams on the accuracy of the model’s predictions is also explored, including using a temporal frame stack or images from two cameras viewing different toroidal regions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Development of lean, efficient, and fast physics-framed deep-learning-based proxy models for subsurface carbon storage

In this work, we present deep-learning-based surrogate models for CCUS developed with four different algorithms and a physics-framed two-phase flow problem involving displacement of water by CO 2 . The deep-learning models were trained using 3D datasets describing the pressure plume, CO 2 saturation plume, and water extraction rate generated by numerical simulation. The hyperparameters defining the architecture of the neural networks were optimized to determine the slimmest network size and training parameters that give the most efficient performance at the least training cost. To develop a robust model that closely mimics the governing physical laws, the discretized form of the two-phase fluid transport equation was used to formulate the supervised deep-learning task. The algorithms investigated in this study predicted the data to above 95% accuracy, with the multi-layer perceptron model demonstrating the best performance by balancing training speed, prediction time, and prediction accuracy with lean network capacity. Furthermore, the surrogate models simultaneously predict reservoir pressure and CO 2 saturation in every grid block, including the surface well extraction rate and bottomhole pressure, at all simulation times for a given static model realization in just a few seconds on a standard desktop computer. A key outcome of this study is that limits can be placed on network design parameters to avoid over designing neural networks, with associated efficiencies in training and prediction times. This is very useful because large volumes of data may be generated in CCUS projects and over-design of neural network architectures imposes penalties that are antithetical to the goal of near-real time forecasting.

58 GEOSCIENCES↗

A hybrid neural architecture: Online attosecond x-ray characterization

The emergence of high-repetition-rate x-ray free-electron lasers (XFELs), such as SLAC’s LCLS-II, serves as our canonical example for autonomous controls that necessitate high-throughput diagnostics paired with streaming computational pipelines capable of single-shot analysis with extremely low latency. We present the deterministic characterization with an integrated parallelizable hybrid resolver architecture, a hybrid machine learning framework designed for fast, accurate analysis of XFEL diagnostics using angular streaking-based sinogram images. This architecture integrates convolutional neural networks and bidirectional long short-term memory models to denoise input, identify x-ray sub-spike features, and extract sub-spike relative delays with sub-30 attosecond temporal resolution. Deployed on low-latency hardware, it achieves over 10 kHz throughput with 168.3 μs inference latency, indicating scalability to 14 kHz with field-programmable gate array integration. By transforming regression tasks into classification problems and leveraging optimized error encoding, we achieve high precision with low-latency performance that is critical for real-time streaming event selection and experimental control feedback signals. This represents a key development in real-time control pipelines for next-generation autonomous science, generally, and high repetition-rate x-ray experiments in particular.

Accelerator Physics (physics.acc-ph)↗

Fast and Accurate Pixel Calibration of Tof Neutron Diffractometers with Machine Learning

At a spallation neutron source, neutron pulses of varying energies are generated, and the detection of neutrons by instrument detectors is recorded as time-of-flight from the emission of the neutron pulse to its arrival at specific detector pixels with high time resolution. The flight path of neutrons from the moderator to the sample and then to the detector must be precisely calibrated at the detector-pixel level using standard powders, so the neutron events from all pixels can be time-focused to produce high-resolution diffraction patterns. Modern time-of-flight neutron diffractometers at spallation neutron sources are equipped with two-dimensional detectors with millimeter-scale pixelations. The number of pixels in a diffraction instrument can reach millions, which makes a single-pixel-level calibration process time-consuming or even impossible with conventional refinement or fitting approaches. Here we present a machine-learning-aided calibration process using a train-and-predict approach, in which machine learning models are trained on the relationship between an individual pixel time-of-flight diffraction pattern and its diffraction constant. These models use a portion of the available pixels for training, and a good model then predicts the diffraction constants precisely and rapidly for large sets of pixel diffraction patterns.

detector pixel calibration↗

Neural Network Reflectance Prediction Model for Both Open Ocean and Coastal Waters

Remote sensing of global ocean color is a valuable tool for understanding the ecology and biogeochemistry of the worlds oceans, and provides critical input to our knowledge of the global carbon cycle and the impacts of climate change. Ocean polarized reflectance contains information about the constituents of the upper ocean euphotic zone, such as colored dissolved organic matter (CDOM), sediments, phytoplankton, and pollutants. In order to retrieve the information on these constituents, remote sensing algorithms typically rely on radiative transfer models to interpret water color or remote-sensing reflectance; however, this can be resource-prohibitive for operational use due to the extensive CPU time involved in radiative transfer solutions. In this work, we report a fast model based on machine learning techniques, called Neural Network Reflectance Prediction Model (NNRPM), which can be used to predict ocean bidirectional polarized reflectance given inherent optical properties of ocean waters. This supervised model is trained using a large volume of data derived from radiative transfer simulations for coupled atmosphere and ocean systems using the successive order of scattering technique (SOS-CAOS). The performance of the model is validated against another large independent test dataset generated from SOS-CAOS. The model is able to predict both polarized and unpolarized reflectances with an absolute error (AE) less than 0.004 for 99% of test cases. We have also shown that the degree of linear polarization (DoLP) for unpolarized incident light can be predicted with an AE less than 0.002 for 99% of test cases. In general, the simulation time of SOS-CAOS depends on optical depth, and required accuracy. When comparing the average speeds of the NNRPM against the SOS-CAOS model for the same parameters, we see that the NNRPM is able to predict the Ocean BRDF 6000 times faster than SOS-CAOS. Both ultraviolet and visible wavelengths are included in the model to help differentiate between dissolved organic material and chlorophyll in the study of the open ocean and the coastal zone. The incorporation of this model into the retrieval algorithm will make the retrieval process more efficient, and thus applicable for operational use with global satellite observations.

radiative transfer↗

Data reduction through optimized scalar quantization for more compact neural networks

Raw data generation for several existing and planned large physics experiments now exceeds TB/s rates, generating untenable data sets in very little time. Those data often demonstrate high dimensionality while containing limited information. Meanwhile, Machine Learning algorithms are now becoming an essential part of data processing and data analysis. Those algorithms can be used offline for post processing and post data analysis, or they can be used online for real time processing providing ultra low latency experiment monitoring. Both use cases would benefit from data throughput reduction while preserving relevant information: one by reducing the offline storage requirements by several orders of magnitude and the other by allowing ultra fast online inferencing with low complexity Machine Learning models. Moreover, reducing the data source throughput also reduces material cost, power and data management requirements. In this work we demonstrate optimized nonuniform scalar quantization for data source reduction. This data reduction allows lower dimensional representations while preserving the relevant information of the data, thus enabling high accuracy Tiny Machine Learning classifier models for online fast inferences. We demonstrate this approach with an initial proof of concept targeting the CookieBox, an array of electron spectrometers used for angular streaking, that was developed for LCLS-II as an online beam diagnostic tool. We used the Lloyd-Max algorithm with the CookieBox dataset to design an optimized nonuniform scalar quantizer. Optimized quantization lets us reduce input data volume by 69% with no significant impact on inference accuracy. When we tolerate a 2% loss on inference accuracy, we achieved 81% of input data reduction. Finally, the change from a 7-bit to a 3-bit input data quantization reduces our neural network size by 38%.

97 MATHEMATICS AND COMPUTING↗

Machine learning without a processor: Emergent learning in a nonlinear analog network

Standard deep learning algorithms require differentiating large nonlinear networks, a process that is slow and power-hungry. Electronic contrastive local learning networks (CLLNs) offer potentially fast, efficient, and fault-tolerant hardware for analog machine learning, but existing implementations are linear, severely limiting their capabilities. These systems differ significantly from artificial neural networks as well as the brain, so the feasibility and utility of incorporating nonlinear elements have not been explored. Here, we introduce a nonlinear CLLN—an analog electronic network made of self-adjusting nonlinear resistive elements based on transistors. We demonstrate that the system learns tasks unachievable in linear systems, including XOR (exclusive or) and nonlinear regression, without a computer. We find our decentralized system reduces modes of training error in order (mean, slope, curvature), similar to spectral bias in artificial neural networks. The circuitry is robust to damage, retrainable in seconds, and performs learned tasks in microseconds while dissipating only picojoules of energy across each transistor. This suggests enormous potential for fast, low-power computing in edge systems like sensors, robotic controllers, and medical devices, as well as manufacturability at scale for performing and studying emergent learning.

Science & Technology - Other Topics↗

Probing degradation at solid-state battery interfaces using machine-learning interatomic potential

Solid-state batteries featuring fast ion-conducting solid electrolytes are promising next-generation energy storage technologies, yet challenges remain for practical deployment due to electro-chemo-mechanical instabilities at solid-solid interfaces. These interfaces, which include homogeneous/internal interfaces such as grain boundaries (GBs) and heterogeneous/external interfaces between solid-electrolyte and electrode materials, can impede Li-ion transport, deteriorate performance, and eventually lead to cell failure. Here, in this study, we leverage large-scale molecular simulations, enabled by validated machine-learning interatomic potentials, to directly probe the onset of interfacial degradation at the garnet Li 7 La 3 Zr 2 O 12 (LLZO) solid-electrolyte/LiCoO 2 (LCO) cathode interface. By surveying different interfacial geometries and compositions, it is found that Li-deficient interfaces can lead to severe interfacial disordering with cation mixing and Co interdiffusion from LCO into LLZO. By contrast, Li-sufficient interfaces are less disordered, although elemental segregation with local ordering is observed. As a consequence of Co interdiffusion, Co-rich regions are formed at the GBs of LLZO due to cation segregation and trapping effects. This behavior is independent of the GB tilting axis, degree of disorder at the GBs, and Co concentration, which implies Co clustering at GBs is a general phenomenon in polycrystalline LLZO and can dictate its overall transport and mechanical properties. Our findings elucidate the underlying fundamental mechanisms that give rise to experimentally observed physicochemical properties and provide guidelines for interface design that can mitigate interfacial degradation and improve cycling performance.

25 ENERGY STORAGE↗

Fast and flexible analysis of direct dark matter search data with machine learning

We present the results from combining machine learning with the profile likelihood fit procedure, using data from the Large Underground Xenon (LUX) dark matter experiment. This approach demonstrates reduction in computation time by a factor of 30 when compared with the previous approach, without loss of performance on real data. We establish its flexibility to capture non-linear correlations between variables (such as smearing in light and charge signals due to position variation) by achieving equal performance using pulse areas with and without position-corrections applied. Its efficiency and scalability furthermore enables searching for dark matter using additional variables without significant computational burden. We demonstrate this by including a light signal pulse shape variable alongside more traditional inputs such as light and charge signal strengths. Furthermore, this technique can be exploited by future dark matter experiments to make use of additional information, reduce computational resources needed for signal searches and simulations, and make inclusion of physical nuisance parameters in fits tractable.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hls4ml Synthesis Testing

HLS4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where HLS4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗

hls4ml

hls4ml (high level synthesis for machine learning) Is a Python package used to translate commonly used open-source machine learning models into HLS. This is useful in machine learning applications on FPGAs. Machine learning algorithms are only as fast as the hardware that they are used on, and some applications require high speed without sacrificing accuracy. In these situations, an FPGA is a good choice since it is faster than a CPU or a GPU, but programming an FPGA is difficult. This is where hls4ml can be used to simplify the process, as a well-known learning model can be converted to HLS and more easily deployed onto an FPGA. There are many use cases for a machine learning algorithm running on an FPGA. For example, detectors in a particle accelerator cannot keep every event that they detect, and so a computer must decide which events to keep and which to discard. Using an FPGA with a machine learning algorithm would be a good way to keep as many events as possible.

Swanson, Caiden↗