Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Learning theory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗

Predicting Metabolic Reaction Networks with Perturbation-Theory Machine Learning (PTML) Models

Background: Checking the connectivity (structure) of complex Metabolic Reaction Networks(MRNs) models proposed for new microorganisms with promising properties is an importantgoal for chemical biology. Objective: In principle, we can perform a hand-on checking (Manual Curation). However, this is achallenging task due to the high number of combinations of pairs of nodes (possible metabolic reactions). Results: The CPTML linear model obtained using the LDA algorithm is able to discriminate nodes(metabolites) with the correct assignation of reactions from incorrect nodes with values of accuracy,specificity, and sensitivity in the range of 85-100% in both training and external validation dataseries. Methods: In this work, we used Combinatorial Perturbation Theory and Machine Learning techniquesto seek a CPTML model for MRNs >40 organisms compiled by Barabasis’ group. First, wequantified the local structure of a very large set of nodes in each MRN using a new class of node indexcalled Markov linear indices fk. Next, we calculated CPT operators for 150000 combinationsof query and reference nodes of MRNs. Last, we used these CPT operators as inputs of differentML algorithms. Conclusion: Meanwhile, PTML models based on Bayesian network, J48-Decision Tree and RandomForest algorithms were identified as the three best non-linear models with accuracy greaterthan 97.5%. The present work opens the door to the study of MRNs of multiple organisms usingPTML models.

Pharmacology & Pharmacy↗

Predicting band gaps and band-edge positions of oxide perovskites using density functional theory and machine learning

Density functional theory (DFT) within the local or semilocal density approximations, i.e., the local density approximation (LDA) or generalized gradient approximation (GGA), has become a workhorse in the electronic structure theory of solids, being extremely fast and reliable for energetics and structural properties, yet remaining highly inaccurate for predicting band gaps of semiconductors and insulators. The accurate prediction of band gaps using first-principles methods is time consuming, requiring hybrid functionals, quasiparticle GW, or quantum Monte Carlo methods. Efficiently correcting DFT-LDA/GGA band gaps and unveiling the main chemical and structural factors involved in this correction is desirable for discovering novel materials in high-throughput calculations. In this direction, we, in this study, use DFT and machine learning techniques to correct band gaps and band-edge positions of a representative subset of ABO 3 perovskite oxides. Relying on the results of HSE06 hybrid functional calculations as target values of band gaps, we find a systematic band-gap correction of ~1.5 eV for this class of materials, where ~1 eV comes from downward shifting the valence band and ~0.5 eV from uplifting the conduction band. The main chemical and structural factors determining the band-gap correction are determined through a feature selection procedure.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Machine Learned Hückel Theory: Interfacing Physics and Deep Neural Networks

The Hückel Hamiltonian is an incredibly simple tight-binding model known for its ability to capture qualitative physics phenomena arising from electron interactions in molecules and materials. Part of its simplicity arises from using only two types of empirically fit physics-motivated parameters: the first describes the orbital energies on each atom and the second describes electronic interactions and bonding between atoms. By replacing these empirical parameters with machine-learned dynamic values, we vastly increase the accuracy of the extended Hückel model. The dynamic values are generated with a deep neural network, which is trained to reproduce orbital energies and densities derived from density functional theory. The resulting model retains interpretability, while the deep neural network parameterization is smooth and accurate and reproduces insightful features of the original empirical parameterization. Altogether, this work shows the promise of utilizing machine learning to formulate simple, accurate, and dynamically parameterized physics models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Observable optimization for precision theory: machine learning energy correlators

The practice of collider physics typically involves the marginalization of multi-dimensional collider data to uni-dimensional observables relevant for some physics task. In many cases, such as classification or anomaly detection, the observable can be arbitrarily complicated, such as the output of a neural network. However, for precision measurements, the observable must correspond to something computable systematically beyond the level of current simulation tools. In this work, we demonstrate that precision-theory-compatible observable space exploration can be systematized by using neural simulation-based inference techniques from machine learning. We illustrate this approach by exploring the space of marginalizations of the energy 3-point correlator to optimize sensitivity to the top quark mass. We first learn the energy-weighted probability density from simulation, then search in the space of marginalizations for an optimal triangle shape. Although simulations and machine learning are used in the process of observable optimization, the output is an observable definition which can be then computed to high precision and compared directly to data without any memory of the computations which produced it. We find that the optimal marginalization is isosceles triangles on the sphere with a side ratio approximately $1 : 1 : \sqrt{2}$ (i.e. right triangles) within the set of marginalizations we consider.

Jets and Jet Substructure↗

Reformulation of the No-Free-Lunch Theorem for Entangled Datasets

The No-Free-Lunch (NFL) theorem is a celebrated result in learning theory that limits one’s ability to learn a function with a training data set. With the recent rise of quantum machine learning, it is natural to ask whether there is a quantum analog of the NFL theorem, which would restrict a quantum computer’s ability to learn a unitary process with quantum training data. However, in the quantum setting, the training data can possess entanglement, a strong correlation with no classical analog. In this work, we show that entangled data sets lead to an apparent violation of the (classical) NFL theorem. This motivates a reformulation that accounts for the degree of entanglement in the training set. As our main result, we prove a quantum NFL theorem whereby the fundamental limit on the learnability of a unitary is reduced by entanglement. We employ Rigetti's quantum computer to test both the classical and quantum NFL theorems. In conclusion, our work establishes that entanglement is a commodity in quantum machine learning.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Statistical Complexity of Quantum Learning

Abstract Learning problems involve settings in which an algorithm has to make decisions based on data, and possibly side information such as expert knowledge. This study has two main goals. First, it reviews and generalizes different results on the data and model complexity of quantum learning, where the data and/or the algorithm can be quantum, focusing on information‐theoretic techniques. Second, it introduces the notion of copy complexity, which quantifies the number of copies of a quantum state required to achieve a target accuracy level. Copy complexity arises from the destructive nature of quantum measurements, which irreversibly alter the state to be processed, limiting the information that can be extracted about quantum data. As a result, empirical risk minimization is generally inapplicable. The paper presents novel results on the copy complexity for both training and testing. To make the paper self‐contained and approachable by different research communities, an extensive background material is provided on classical results from statistical learning theory, as well as on the distinguishability of quantum states. Throughout, the differences between quantum and classical learning are highlighted by addressing both supervised and unsupervised learning, and extensive pointers are provided to the literature.

97 MATHEMATICS AND COMPUTING↗

Universal Compiling and (No-)Free-Lunch Theorems for Continuous-Variable Quantum Learning

Quantum compiling, where a parameterized quantum circuit is trained to learn a target unitary, is an important primitive for quantum computing that can be used as a subroutine to obtain optimal circuits or as a tomographic tool to study the dynamics of an experimental system. While much attention has been paid to quantum compiling on discrete-variable hardware, less has been paid to compiling in the continuous-variable paradigm. Here we motivate several, closely related, short-depth continuous-variable algorithms for quantum compilation. We analyze the trainability of our proposed cost functions and numerically demonstrate our algorithms by learning arbitrary Gaussian operations and Kerr nonlinearities. We further make connections between this framework and quantum learning theory in the continuous-variable setting by deriving no-free-lunch theorems. These generalization bounds demonstrate a linear resource reduction for learning Gaussian unitaries using entangled coherent-Fock states and an exponential resource reduction for learning arbitrary unitaries using two-mode-squeezed states.

97 MATHEMATICS AND COMPUTING↗

Seeding picoscale solutions for social macro goals: complex thinking in projects for vulnerable communities

This chapter presents lessons learned from two experiences of implementing sociopoliticalsustainability strategies to achieve environmental and economic sustainability.The projects were carried out in rural communities in Guinea-Bissau andJamaica, which are both vulnerable to climate change. Each experience, developedfrom a systemic perspective and under two frameworks (the water–food–energynexus and ‘Appropriate Technology’), is described in four phases (problemcharacterization, solution design, solution pre-evaluation, and implementation),with the significant findings made by the team in terms of challenges, rewards, andessential conditions for development. Both deal with environment-friendly solutionsthat solve the problems of water and energy provision. Based on complex thinking,the experiences are analyzed from the social inclusion, participation, empowerment,and learning theory viewpoints. Finally, the insights gained by the team arehighlighted: communication requirements, the level of dependency of industrializedsolutions, budget constraints, training, testing, and planning recommendations, thediversity of interaction with the communities, and partners’ skills. As a result, futureactions can exploit these experiences and contribute to the success of subsequentprojects.

Pereira Pinto, Joao↗

A HPC Theory-Guided Machine Learning Cyberinfrastructure for Communicating Hydrometeorological Data Across Scales

High-resolution predictions of hydrometeorological variables are critical for supporting hydropower generation decisions and flood control at hydroelectric power plants. Traditional climate and hydrologic models rely on the numerical simulation of detailed physical processes. Therefore, running these simulations is time-, labor-, and computation-intensive. Improving the spatial and temporal resolution in these modeling outputs could lead to cubic increases in both the simulation time and computational demands, rendering high-resolution hydrometeorological predictions expensive and impractical. Many past studies apply the super resolution (SR) technique to downscale climate models using deep learners. However, deep learners are deemed “black-boxes,” as their derivation processes from low-resolution outputs to high-resolution outputs are often hidden. Their results are difficult for domain scientists to interpret and validate. Thus, there is a need for an exploratory machine learning approach that can partially integrate domain-specific theory and knowledge into the data-driven mapping process between simulation outputs of different spatial scales. The domain-specific theory and knowledge can be incorporated into the data model through an inductive approach in which process-related environmental variables are used and analyzed as key drivers (i.e., environmental surrogates) to reflect the complex physical processes. Many of these variables, such as land use land cover, soil types, topography, digital elevation, air temperature, and various watershed characteristics, can be directly measured through sensors or remote sensing techniques. Additionally, SR applications that can downscale hydrological and hydrodynamics models to efficiently produce high-resolution (1 m) flood depth grids are still rare. Since the flood depth grid can be used to support critical decisions for flood control operation at hydroelectric power plants, it is crucial to enable an SR-based capability for interpolating high-resolution flood inundation maps.

13 HYDRO ENERGY↗

Model-free estimation of completeness, uncertainties, and outliers in atomistic machine learning using information theory

Abstract An accurate description of information is relevant for a range of problems in atomistic machine learning (ML), such as crafting training sets, performing uncertainty quantification (UQ), or extracting physical insights from large datasets. However, atomistic ML often relies on unsupervised learning or model predictions to analyze information contents from simulation or training data. Here, we introduce a theoretical framework that provides a rigorous, model-free tool to quantify information contents in atomistic simulations. We demonstrate that the information entropy of a distribution of atom-centered environments explains known heuristics in ML potential developments, from training set sizes to dataset optimality. Using this tool, we propose a model-free UQ method that reliably predicts epistemic uncertainty and detects out-of-distribution samples, including rare events in systems such as nucleation. This method provides a general tool for data-driven atomistic modeling and combines efforts in ML, simulations, and physical explainability.

36 MATERIALS SCIENCE↗

Sparse Data Machine Learning Integration with Theory, Experiment and Uncertainty Quantification: Process-Structure-Property-Performance of Friction Deformation Processing

Computer vision and deep learning tools that advance the ability to establish processing-structure-property-performance (PSPP) relations are presented. The Bayesian binning method for image segmentation enables quantitative analysis of microstructural features in an automated way, while the analysis of shapes and relative orientation of these features reveals local deformation maps indicative of both, material flow and residual stresses due to materials processing. The deep learning method leads to the previous knowledge agnostic mapping of empirically observed microstructural zones in friction stir welding (FSW) process and synthetic microstructure generation capability that is statistically equivalent to experimentally collected data.

97 MATHEMATICS AND COMPUTING↗

Accelerating multiscale electronic stopping power predictions with time-dependent density functional theory and machine learning

Knowing the rate at which particle radiation releases energy in a material, the “stopping power,” is key to designing nuclear reactors, medical treatments, semiconductor and quantum materials, and many other technologies. While the nuclear contribution to stopping power, i.e., elastic scattering between atoms, is well understood in the literature, the route for gathering data on the electronic contribution has for decades remained costly and reliant on many simplifying assumptions, including that materials are isotropic. We establish a method that combines time-dependent density functional theory (TDDFT) and machine learning to reduce the time to assess new materials to hours on a supercomputer and provide valuable data on how atomic details influence electronic stopping. Our approach uses TDDFT to compute the electronic stopping from first principles in several directions and then machine learning to interpolate to other directions at a cost of 10 million times fewer core-hours. We demonstrate the combined approach in a study of proton irradiation in aluminum and employ it to predict how the depth of maximum energy deposition, the “Bragg Peak,” varies depending on the incident angle—a quantity otherwise inaccessible to modelers and far outside the scales of quantum mechanical simulations. The lack of any experimental information requirement makes our method applicable to most materials, and its speed makes it a prime candidate for enabling quantum-to-continuum models of radiation damage. The prospect of reusing valuable TDDFT data for training the model makes our approach appealing for applications in the age of materials data science.

36 MATERIALS SCIENCE↗

Extreme Risk Mitigation in Reinforcement Learning using Extreme Value Theory

Risk-sensitive reinforcement learning (RL) has garnered significant attention in recent years due to the growing interest in deploying RL agents in real-world scenarios. A critical aspect of risk awareness involves modelling highly rare risk events (rewards) that could potentially lead to catastrophic outcomes. These infrequent occurrences present a formidable challenge for data-driven methods aiming to capture such risky events accurately. While risk-aware RL techniques do exist, they suffer from high variance estimation due to the inherent data scarcity. Our work proposes to enhance the resilience of RL agents when faced with very rare and risky events by focusing on refining the predictions of the extreme values predicted by the state-action value distribution. To achieve this, we formulate the extreme values of the state-action value function distribution as parameterized distributions, drawing inspiration from the principles of extreme value theory (EVT). We propose an extreme value theory based actor-critic approach, namely, Extreme Valued Actor-Critic (EVAC) which effectively addresses the issue of infrequent occurrence by leveraging EVT-based parameterization. Importantly, we theoretically demonstrate the advantages of employing these parameterized distributions in contrast to other risk-averse algorithms. Our evaluations show that the proposed method outperforms other risk averse RL algorithms on a diverse range of benchmark tasks, each encompassing distinct risk scenarios.

Wang, Yu↗

Identifying Key Drivers of Wildfires in the Contiguous US Using Machine Learning and Game Theory Interpretation

Abstract Understanding the complex interrelationships between wildfire and its environmental and anthropogenic controls is crucial for wildfire modeling and management. Although machine learning (ML) models have yielded significant improvements in wildfire predictions, their limited interpretability has been an obstacle for their use in advancing understanding of wildfires. This study builds an ML model incorporating predictors of local meteorology, land‐surface characteristics, and socioeconomic variables to predict monthly burned area at grid cells of 0.25° × 0.25° resolution over the contiguous United States. Besides these predictors, we construct and include predictors representing the large‐scale circulation patterns conducive to wildfires, which largely improves the temporal correlations in several regions by 14%–44%. The Shapley additive explanation is introduced to quantify the contributions of the predictors to burned area. Results show a key role of longitude and latitude in delineating fire regimes with different temporal patterns of burned area. The model captures the physical relationship between burned area and vapor pressure deficit, relative humidity (RH), and energy release component (ERC), in agreement with the prior findings. Aggregating the contribution of predictor variables of all the grids by region, analyses show that ERC is the major contributor accounting for 14%–27% to large burned areas in the western US. In contrast, there is no leading factor contributing to large burned areas in the eastern US, although large‐scale circulation patterns featuring less active upper‐level ridge‐trough and low RH two months earlier in winter contribute relatively more to large burned areas in spring in the southeastern US.

54 ENVIRONMENTAL SCIENCES↗

Integrating Maximum Entropy Production Theory and Machine Learning to Improve Global Evapotranspiration Modeling

Accurate estimation of terrestrial evapotranspiration (ET) is vital for understanding global water and energy cycles. However, current global ET estimations are not well constrained. This study introduces an integrated framework combining the Maximum Entropy Production (MEP) theory with Random Forest (RF) model to improve global ET estimation. Specifically, in contrast to direct ET estimation by the RF model, the integrated framework (MEP‐RF) trains to predict error of MEP‐simulated ET. MEP‐RF outperforms RF in spatiotemporal extrapolation. Attribution analysis with in situ observations reveals that the inputs of MEP are the most critical variables for the ET process, including net radiation, vegetated area, soil moisture, and surface temperature. We further drive MEP‐RF with global reanalysis and satellite data sets of these four inputs, yielding a global mean terrestrial ET of 548 mm/year, with 77% attributed to transpiration. The global ET increased at a rate of 0.85 mm/year per year during 2003–2021, primarily due to vegetation greening rather than rising temperature, while decreasing soil moisture led to decreasing regional ET. The integrated framework provides a novel approach for the estimation of global ET without the need for hard‐to‐obtain and thus uncertain inputs, such as wind speed, surface roughness, aerodynamic and canopy stomatal resistance. Therefore, MEP‐RF offers an independent method on existing global ET products. It represents a promising physically based approach that can be incorporated into Earth System Models to enhance water and energy cycle simulations.

54 ENVIRONMENTAL SCIENCES↗