Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “empirical deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Comparison of DeePMD, MTP, GAP, ACE and MACE Machine‐Learned Potentials for Radiation‐Damage Simulations: A User Perspective

Accurate and efficient interatomic potentials are essential for molecular dynamics (MD) simulations of radiation damage, gas diffusion, and phase stability in complex ceramics such as LiAlO 2 , especially under extreme conditions relevant to tritium production. Here, we evaluate the performance of six machine-learned interatomic potentials (MLIPs), moment tensor potential (MTP), Gaussian approximation potential, deep potential (DeePMD), atomic cluster expansion (ACE), message-passing ACE (multilayer atomic cluster expansion (MACE) pretrained) and MACE (trained from-scratch), all trained on the same density functional theory dataset with inclusion of tritium. The MLIPs are benchmarked against traditional Buckingham and ReaxFF potentials in terms of energy accuracy, density predictions, thermal equilibration behavior, threshold displacement energy (E d ), tritium diffusivity, and computational cost. Among the models, MTP shows the best overall balance between efficiency and accuracy, with low force and energy errors and realistic E d values for Li and Al. The ACE and MACE (pretrained and trained from scratch) models exhibit high E d (>200 eV) and unphysical pair interactions. DeePMD underestimates Ed due to overly repulsive behavior even at equilibrium distances. All models over-estimate tritium diffusion but the pretrained MACE model behaves well during tritium-diffusion simulations up to 500 K, maintaining diffusivities in the physically consistent 10 −11 m 2 /s range. Finally, we quantify the computational cost of each potential in large-scale atomic/molecular massively parallel simulator, finding that only MTP is more efficient than traditional empirical potentials, while others are significantly more expensive. These findings explain the trade-offs between accuracy and computational cost in MLIP development and provide essential guidance for use in high-throughput radiation damage and gas diffusion simulations in nuclear ceramics.

74 ATOMIC AND MOLECULAR PHYSICS↗

On the Stochastic Stability of Deep Markov Models

Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. DMMs using deep neural networks to parametrize the transition of Markov probability distributions have recently been shown to provide more expressiveness in modeling sequential data and dynamical system responses. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a rigorous analytical method to prove the necessary and sufficient conditions of DMM's stochastic stability. This task is achieved by spectral analysis of the efficiently computed Jacobians of probabilistic maps modeled by deep neural networks. We make theoretical connections between the eigenvalues of neural network's weights and the different activation function types used on the stability and overall dynamic behavior of DMMs with Gaussian distributions. We empirically substantiate our theoretical results on stochastic stability and eigenvalue spectra via several numerical experiments. Formal stability guarantees of DMMs can substantially improve their robustness and trustworthiness, necessary for reliable use in safety-critical real-world applications.

Drgona, Jan↗

Cultural Shifts in High Energy Physics Collaboration from the Cold War to the Present: A Historical and Philosophical Perspective

Here, this article employs empirical history and the philosophy of science to study cultural convergences and divergences in international collaborations in high energy physics. We examine two cases: (1) E-36, an experiment on small angle proton-proton scattering conducted during the Cold War at the National Accelerator Laboratory (NAL) in the USA by Soviet and US scientists and (2) an ongoing collaborative experiment, NICA, at the Joint Institute for Nuclear Research (JINR, Dubna), which is a project devoted to heavy-ion physics. The JINR, particularly its Laboratory of High Energy Physics (formerly the “Laboratory of High Energies”) is the main mediating actor between these two cases (i.e., E-36 and NICA), as the majority of Soviet participants in E-36 were representatives of the Institute. Using empirical data collected through archival searches, field observations conducted at JINR in 2018–2019, and in-depth interviews, we tell a story of cultural differences in high energy physics by applying the concepts of ‘trading zones’ (P. Galison) and the translation of interests in actor-networks (B. Latour, M. Callon and others). We analyze three types of cultural diversity (specialization, nationality, and generational) in light of the implications of temporal context and the dichotomy between East and West, showing the roles cultural diversity plays in scientific collaboration (which is an integral part of as well as obstacle to scientific research that can nevertheless provide learning opportunities). Our study aims to demonstrate how disunity and diversity may function in scientific research and how high energy physics collaborations can remain productive despite sometimes deep divergences, including those between East and West.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Spike-and-Slab Shrinkage Priors for Structurally Sparse Bayesian Neural Networks

Network complexity and computational efficiency have become increasingly significant aspects of deep learning. Sparse deep learning addresses these challenges by recovering a sparse representation of the underlying target function by reducing heavily overparameterized deep neural networks. Specifically, deep neural architectures compressed via structured sparsity (e.g., node sparsity) provide low-latency inference, higher data throughput, and reduced energy consumption. In this article, we explore two well-established shrinkage techniques, Lasso and Horseshoe, for model compression in Bayesian neural networks (BNNs). To this end, we propose structurally sparse BNNs, which systematically prune excessive nodes with the following: 1) spike-and-slab group Lasso (SS-GL) and 2) SS group Horseshoe (SS-GHS) priors, and develop computationally tractable variational inference, including continuous relaxation of Bernoulli variables. We establish the contraction rates of the variational posterior of our proposed models as a function of the network topology, layerwise node cardinalities, and bounds on the network weights. Furthermore, we empirically demonstrate the competitive performance of our models compared with the baseline models in prediction accuracy, model compression, and inference latency.

97 MATHEMATICS AND COMPUTING↗

Improving Single-Stage Object Detectors for Nighttime Pedestrian Detection

We report Improving the reliability of nighttime pedestrian detection is a crucial challenge towards the design of robust autonomous systems. Not surprisingly, most pedestrian fatalities occur in low-illumination settings, thus emphasizing the need for new algorithmic advances. This work presents a novel pedestrian detection approach that makes a number of crucial modifications to the state-of-the-art YOLOV5-PANet architecture, in order to improve the reliability of features extracted from nighttime images. More specifically, the proposed architecture systematically incorporates powerful shuffle attention mechanisms and a transformer module to improve the feature learning pipeline. Instead of advocating the use of other sensing modalities that are better suited for nighttime detection, our approach relies only on conventional RGB cameras and is hence broadly applicable. Our empirical studies with nighttime pedestrian detection benchmarks show that with only minimal increase in model complexity, our approach provides significant improvements in detection efficacy over existing solutions. Finally, we explore the impact of post-hoc network pruning on the speed-accuracy trade-off of our approach and demonstrate that it is well suited for reduced memory/compute requirements.

97 MATHEMATICS AND COMPUTING↗

Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees

We present differentiable predictive control (DPC), a method for offline learning of constrained neural control policies for nonlinear dynamical systems with performance guarantees. We show that the sensitivities of the parametric optimal control problem can be used to obtain direct policy gradients. Specifically, we employ automatic differentiation (AD) to efficiently compute the sensitivities of the model predictive control (MPC) objective function and constraints penalties. To guarantee safety upon deployment, we derive probabilistic guarantees on closed-loop stability and constraint satisfaction based on indicator functions and Hoeffding’s inequality. We empirically demonstrate that the proposed method can learn neural control policies for various parametric optimal control tasks. In particular, we show that the proposed DPC method can stabilize systems with unstable dynamics, track time-varying references, and satisfy nonlinear state and input constraints. Our DPC method has practical time savings compared to alternative approaches for fast and memory-efficient controller design. Specifically, DPC does not depend on a supervisory controller as opposed to approximate MPC based on imitation learning. We demonstrate that, without losing performance, DPC is scalable with greatly reduced demands on memory and computation compared to implicit and explicit MPC while being more sample efficient than model-free reinforcement learning (RL) algorithms.

97 MATHEMATICS AND COMPUTING↗

Assessing the potential of deep learning for protein–ligand docking

The effects of ligand binding on protein structures and their in vivo functions carry numerous implications for modern biomedical research and biotechnology development efforts such as drug discovery. Although several deep learning (DL) methods and benchmarks designed for protein–ligand docking have recently been introduced, so far no previous works have systematically studied the behaviour of the latest docking and structure prediction methods within the broadly applicable context of: (1) using predicted (apo) protein structures for docking (for example, for applicability to new proteins); (2) binding multiple (cofactor) ligands concurrently to a given target protein (for example, for enzyme design); and (3) having no previous knowledge of binding pockets (for example, for generalization to unknown pockets). To enable a deeper understanding of the real-world utility of docking methods, we introduce PoseBench, a comprehensive benchmark for broadly applicable protein–ligand docking. PoseBench enables researchers to rigorously and systematically evaluate DL methods for apo-to-holo protein–ligand docking and protein–ligand structure prediction using both primary ligand and multiligand benchmark datasets, the latter of which we introduce to the DL community. Empirically, using PoseBench, we find that: (1) DL cofolding methods generally outperform comparable conventional and DL docking baseline algorithms, but popular methods such as AlphaFold 3 are still challenged by prediction targets with new protein–ligand binding poses; (2) certain DL cofolding methods are highly sensitive to their input multiple sequence alignments, whereas others are not; and (3) DL methods struggle to strike a balance between structural accuracy and chemical specificity when predicting new or multiligand protein targets.

Morehead, Alex [Lawrence Berkeley National Laborat↗

Rapid wavefield forecasting for earthquake early warning via deep sequence to sequence learning

We propose a deep learning model, WaveCastNet, to forecast high-dimensional wavefields. WaveCastNet integrates a convolutional long expressive memory architecture into a sequence-to-sequence forecasting framework, enabling it to model long-term dependencies and multiscale patterns in both space and time. By sharing weights across spatial and temporal dimensions, WaveCastNet requires significantly fewer parameters than more resource-intensive models such as transformers, resulting in faster inference times. Crucially, WaveCastNet also generalizes better than transformers to rare and critical seismic scenarios, such as high-magnitude earthquakes. Here, we show the ability of the model to predict the intensity and timing of destructive ground motions in real time, using simulated data from the San Francisco Bay Area. Furthermore, we demonstrate its zero-shot capabilities by evaluating WaveCastNet on real earthquake data. Our approach does not require estimating earthquake magnitudes and epicenters, steps that are prone to error in conventional methods, nor does it rely on empirical ground-motion models, which often fail to capture strongly heterogeneous wave propagation effects.

Geophysics↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Evaluating Physics-Informed Neural Network Performance for Seismic Discrimination between Earthquakes and Explosions

In this article, we evaluate adding a weak physics constraint, that is, a physics‐based empirical relationship, to the loss function with a physics‐informed manner in local distance explosion discrimination in the hope of improving the generalization capability of the machine learning (ML) model. We compare the proposed model with the two‐branch model we previously developed, as well as with a pure data‐driven model. Unexpectedly, the proposed model did not consistently outperform the pure data‐driven model. By varying the level of inconsistency in the training data, we find this approach is modulated by the strength of the physics relationship. In conclusion, this result has important implications for how to best incorporate physical constraints in ML models.

58 GEOSCIENCES↗

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder↗

Using Data-Driven Prediction of Downstream 1D River Flow to Overcome the Challenges of Hydrologic River Modeling

Methods for downstream river flow prediction can be categorized into physics-based and empirical approaches. Although based on well-studied physical relationships, physics-based models rely on numerous hydrologic variables characteristic of the specific river system that can be costly to acquire. Moreover, simulation is often computationally intensive. Conversely, empirical models require less information about the system being modeled and can capture a system’s interactions based on a smaller set of observed data. This article introduces two empirical methods to predict downstream hydraulic variables based on observed stream data: a linear programming (LP) model, and a convolutional neural network (CNN). We apply both empirical models within the Colorado River system to a site located on the Green River, downstream of the Yampa River confluence and Flaming Gorge Dam, and compare it to the physics-based model Streamflow Synthesis and Reservoir Regulation (SSARR) currently used by federal agencies. Results show that both proposed models significantly outperform the SSARR model. Moreover, the CNN model outperforms the LP model for hourly predictions whereas both perform similarly for daily predictions. Although less accurate than the CNN model at finer temporal resolution, the LP model is ideal for linear water scheduling tools.

13 HYDRO ENERGY↗

Graph neural networks predict energetic and mechanical properties for models of solid solution metal alloy phases

Here, we developed a PyTorch-based architecture called HydraGNN that implements graph convolutional neural networks (GCNNs) to predict the formation energy and the bulk modulus for models of solid solution alloys for various atomic crystal structures and relaxed volumes. We trained the GCNN surrogate model on a dataset for nickel–niobium (NiNb) generated by the embedded atom model (EAM) empirical interatomic potential for demonstration purposes. The dataset was generated by calculating the formation energy and the bulk modulus as a prototypical elastic property for optimized geometries starting from initial body-centered cubic (BCC), face-centered cubic (FCC), and hexagonal compact packed (HCP) crystal structures, with configurations spanning the possible compositional range for each of the three types of initial crystal structures. Numerical results show that the GCNN model effectively predicts both the formation energy and the bulk modulus as function of the optimized crystal structure, relaxed volume, and configurational entropy of the model structures for solid solution alloys.

36 MATERIALS SCIENCE↗

Neural-Network-Enhanced COTSIM: Advancing Predictive Capabilities for Fast DIII-D Simulations

Sustaining fusion reactions in tokamaks requires heating plasma to thermonuclear temperatures while maintaining confinement and stability. Neutral beam injection (NBI) provides heating, current drive, torque, and fueling, while electron cyclotron (EC) waves are widely used for heating and current drive; together, these actuators shape the plasma current, temperature, and density profiles. The control-oriented tokamak simulator (COTSIM), a predictive, control-oriented code, has been enhanced with neural-network surrogates for transport and sources. Turbulent transport is predicted by MMMnet—a neural-network version of the updated multimode model (MMM 9.0.10)—with significantly reduced computation time relative to MMM; neoclassical transport follows the Chang–Hinton model. NUBEAMnet, a surrogate of the Monte Carlo NUBEAM module, predicts beam-driven heating, current, and torque. EC heating and current drive use a control-oriented, empirically scaled source model; plasma resistivity follows the Spitzer formulation; bootstrap current uses the Sauter model. Equilibrium is computed using both prescribed and fixed-boundary solvers (FBSs), and the pedestal structure is modeled with an empirical pedestal model. For a representative DIII-D discharge, COTSIM predicts electron and ion temperature and safety-factor profiles in close agreement with TRANSP predictive and interpretive simulations while extending predictions through the pedestal region to the plasma edge (versus 80% of the minor radius in TRANSP). Furthermore, the equivalent COTSIM simulation runs in under 3 min compared to about 2 h for TRANSP, enabling rapid scenario planning, optimization of tokamak operation, and between-pulse control design.

Control-oriented tokamak simulator (COTSIM)↗

Predicting Solar Energetic Particles Using SDO/HMI Vector Magnetic Data Products and a Bidirectional LSTM Network

Solar energetic particles (SEPs) are an essential source of space radiation, and are hazardous for humans in space, spacecraft, and technology in general. In this paper, we propose a deep-learning method, specifically a bidirectional long short-term memory (biLSTM) network, to predict if an active region (AR) would produce an SEP event given that (i) the AR will produce an M- or X-class flare and a coronal mass ejection (CME) associated with the flare, or (ii) the AR will produce an M- or X-class flare regardless of whether or not the flare is associated with a CME. The data samples used in this study are collected from the Geostationary Operational Environmental Satellite's X-ray flare catalogs provided by the National Centers for Environmental Information. We select M- and X-class flares with identified ARs in the catalogs for the period between 2010 and 2021, and find the associations of flares, CMEs, and SEPs in the Space Weather Database of Notifications, Knowledge, Information during the same period. Each data sample contains physical parameters collected from the Helioseismic and Magnetic Imager on board the Solar Dynamics Observatory. Experimental results based on different performance metrics demonstrate that the proposed biLSTM network is better than related machine-learning algorithms for the two SEP prediction tasks studied here. We also discuss extensions of our approach for probabilistic forecasting and calibration with empirical evaluation

79 ASTRONOMY AND ASTROPHYSICS↗

Enhancing generative molecular design via uncertainty-guided fine-tuning of variational autoencoders

In recent years, deep generative models have been successfully applied to various molecular design tasks, particularly in the life and materials sciences. One critical challenge for pre-trained generative molecular design (GMD) models is to fine-tune them to be better suited for downstream design tasks that aim at optimizing specific molecular properties. However, redesigning and training an existing effective generative model from scratch for each new design task are impractical. Furthermore, the black-box nature of typical downstream tasks that involve property prediction makes it nontrivial to optimize the generative model in a task-specific manner. In this work, we propose an uncertainty-guided fine-tuning strategy that can effectively enhance a pre-trained variational autoencoder (VAE) for GMD through performance feedback in an active learning setting. The strategy begins by quantifying the model uncertainty of the generative model using an efficient active subspace-based UQ (uncertainty quantification) scheme. Next, the decoder diversity within the characterized model uncertainty class is explored to expand the viable space of molecular generation. The low-dimensionality of the active subspace makes this exploration tractable using a black-box optimization scheme, which in turn enables us to identify and leverage a diverse set of high-performing models to generate enhanced molecules. Empirical results across six target molecular properties using multiple VAE-based generative models demonstrate that our uncertainty-guided fine-tuning strategy consistently leads to improved models that outperform the original pre-trained models.

97 MATHEMATICS AND COMPUTING↗

Data-Driven Discovery of Linear Molecular Probes with Optimal Selective Affinity for PFAS in Water

Approaches to tackle the wide and growing variety of highly persistent per- and polyfluoroalkyl substances (PFAS) are of pressing global need because of their detrimental human health effects, such as cancer, birth defects, and hormone imbalance. Sensitive, selective, and easy-to-use real-time sensors to monitor and detect PFAS and sorbents to extract them are critical to meeting government-mandated environmental concentrations. In this work, we combine all-atom molecular dynamics simulations, enhanced sampling, deep representational learning, and Bayesian optimization to perform high-throughput virtual screening for highly sensitive and selective molecular probes. Our molecular design space consists of 3850 linear hydrocarbon chains with varying degrees of halogenation with and without amine- and phosphine-based headgroups. By employing a data-driven search process, we efficiently explore the molecular design space to optimize the sensitivity to perfluorooctanesulfonic acid (PFOS) as a prototypical PFAS analyte and selectivity relative to a sodium dodecyl sulfate (SDS) interferent. We calculate 504 Gibbs free energies of probe-analyte and probe-interferent interactions and identify probes with PFOS association free energies of up to (-ΔG PFOS ) = 9.8 ± 0.2 kJ/mol and selectivities relative to SDS of (-ΔΔG PFOS–SDS ) = 3.1 ± 1.5 kJ/mol. A C 11 Br 23 P(CH 3 ) 2 probe containing 11 backbone brominated carbons and a tertiary phosphine headgroup possesses the most sensitive binding constant to PFOS within the defined search space of K b PFOS = 177.4 ± 12.7, and a semibrominated probe C 5 H 11 C 7 Br 14 N(CH 3 ) 2 containing 12 backbone carbons and a tertiary amine headgroup possesses the highest selectivity relative to SDS of K b PFOS /K b SDS = 4.6 ± 1.7. A retrospective analysis of our data to extract interpretable design rules reveals that the sensitivity of linear hydrogenated probes increases by approximately 1 kJ/mol per C–C bond. The addition or removal of halogen atoms and amine or phosphine headgroups produces nonmonotonic changes in both sensitivity and selectivity with changes to the sensitivity of up to 2.5 kJ/mol. Finally, this work places empirical limitations on the performance of a wide range of linear probes for PFOS detection and offers a generic strategy for high-throughput computational screening to promote selective and sensitive binding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Short-term Electricity Price Forecasting with Constrained Regressors

The volatility of electricity price presents a challenge to market participants as their decision-making process are highly depend on the accuracy of price forecasts. However, there is growing empirical evidence of increasing price volatility and price spikes in electricity markets as a result of variable renewable energy generation, extreme weather events, and other factors. The distribution shift caused by spikes in electricity price data differentiates the forecasting tasks from other renewable energy sources. Moreover, the observations may be compromised by cyberattacks and thus not available in the testing phase. To this end, we propose a Similarity-Enhanced Electricity Decomposition Forecasting model (SEED-Forecaster) to address the missing response problem and spikes capturing in short-term electricity price forecasting. The effectiveness of the proposed framework is tested on real-world electricity price data from California Independent System Operator (CAISO). Numerical results of case studies show that the proposed SEED-Forecsater can enhance forecasting performance, particularly in capturing electricity spikes, even under conditions without regressors during testing stage.

24 POWER TRANSMISSION AND DISTRIBUTION↗