Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “empirical machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Machine Learning Predictions of Simulated Self-Diffusion Coefficients for Bulk and Confined Pure Liquids

Diffusion properties of bulk fluids have been predicted using empirical expressions and machine learning (ML) models, suggesting that predictions of diffusion also should be possible for fluids in confined environments. The ability to quickly and accurately predict diffusion in porous materials would enable new discoveries and spur development in relevant technologies such as separations, catalysis, batteries, and subsurface applications. Here in this work, we apply artificial neural network (ANN) models to predict the simulated self-diffusion coefficients of real liquids in both bulk and pore environments. The training data sets were generated from molecular dynamics (MD) simulations of Lennard-Jones particles representing a diverse set of 14 molecules ranging from ammonia to dodecane over a range of liquid pressures and temperatures. Planar, cylindrical, and hexagonal pore models consisted of walls composed of carbon atoms. Our simple model for these liquids was primarily used to generate ANN training data, but the simulated self-diffusion coefficients of bulk liquids show excellent agreement with experimental diffusion coefficients. ANN models based on simple descriptors accurately reproduced the MD diffusion data for both bulk and confined liquids, including the trend of increased mobility in large pores relative to the corresponding bulk liquid.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Myths and legends in learning classification rules

A discussion is presented of machine learning theory on empirically learning classification rules. Six myths are proposed in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, universal learning algorithms, and interactive learning. Some of the problems raised are also addressed from a Bayesian perspective. Questions are suggested that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

Myths and legends in learning classification rules

This paper is a discussion of machine learning theory on empirically learning classification rules. The paper proposes six myths in the machine learning community that address issues of bias, learning as search, computational learning theory, Occam's razor, 'universal' learning algorithms, and interactive learnings. Some of the problems raised are also addressed from a Bayesian perspective. The paper concludes by suggesting questions that machine learning researchers should be addressing both theoretically and experimentally.

Buntine, Wray↗

A Machine Learning Approach to Jet-Surface Interaction Noise Modeling

This paper investigates using machine learning to rapidly develop empirical models suitable for system-level aircraft noise studies. In particular, machine learning is used to train a neural network to predict the noise spectra produced by a round jet near a surface over a range of surface lengths, surface standoff distances, jet Mach numbers, and observer angles. These spectra include two sources, jet-mixing noise and jet-surface interaction (JSI) noise, with different scale factors as well as surface shielding and reflection effects to create a multi- dimensional problem. A second model is then trained using data from three rectangular nozzles to include nozzle aspect ratio in the spectral prediction. The training and validation data are from an extensive jet-surface interaction noise database acquired at the NASA Glenn Research Center's Aero-Acoustic Propulsion Laboratory. Although the number of training and validation points is small compared a typical machine learning application, the results of this investigation show that this approach is viable if the underlying data are well behaved.

Brown, Cliff↗

Effect of particle size and moisture on flow performance of loblolly pine anatomical fractions: Experimental findings and model predictions

The rising energy demand has highlighted biomass as a promising next-generation energy source. However, commercializing biomass-derived energy faces challenges, particularly in handling biomass feedstock. Factors like particle size, shape, moisture content, and surface roughness significantly impact biomass flowability. This study addresses a crucial knowledge gap by examining the effects of particle size and moisture content on the flow behavior and shear properties of different anatomical fractions of loblolly pine (Pinus taeda). The bulk shear behavior was examined using a Schulze ring shear tester, while flow performance was tested through gravity-driven flow experiments in a variable wedge-shape hopper. Results were incorporated into empirical and machine learning-based flow prediction models to evaluate their accuracy and limitations. The study found that samples with higher moisture content show higher unconfined yield strength. The critical arching distance increased with particle size, e.g., from approximately 13 and 33 mm for 2- and 6-mm whole chips, respectively at a 32-degree inclination angle. Conversely, the flow rate decreased for a given hopper opening as particle size increased. For instance, at a 60-mm hopper opening and a 32-degree inclination angle, the mass flow rates for 2- and 6-mm whole chips were 7.83 and 6.42 tonne/h, respectively. The empirical model consistently overpredicted the mass flow rate for all anatomical fractions, while the machine learning model more accurately predicted the central tendency of flow rate but was insensitive to varying tissue proportions. These novel findings provide comprehensive characterization of anatomical fractions, reveal significant combined effects of particle size and moisture content on biomass flow behavior, and demonstrate a better predictive accuracy of a machine learning model, all of which are useful for optimizing material handling strategies and biomass utilization technologies in the industry.

09 - BIOMASS FUELS↗

Synergy of semiempirical models and machine learning in computational chemistry

Catalyzed by enormous success in the industrial sector, many research programs have been exploring data-driven, machine learning approaches. Performance can be poor when the model is extrapolated to new regions of chemical space, e.g., new bonding types, new many-body interactions. Another important limitation is the spatial locality assumption in model architecture, and this limitation cannot be overcome with larger or more diverse datasets. The outlined challenges are primarily associated with the lack of electronic structure information in surrogate models such as interatomic potentials. Given the fast development of machine learning and computational chemistry methods, we expect some limitations of surrogate models to be addressed in the near future; nevertheless spatial locality assumption will likely remain a limiting factor for their transferability. Here, we suggest focusing on an equally important effort—design of physics-informed models that leverage the domain knowledge and employ machine learning only as a corrective tool. In the context of material science, we will focus on semi-empirical quantum mechanics, using machine learning to predict corrections to the reduced-order Hamiltonian model parameters. The resulting models are broadly applicable, retain the speed of semiempirical chemistry, and frequently achieve accuracy on par with much more expensive ab initio calculations. These early results indicate that future work, in which machine learning and quantum chemistry methods are developed jointly, may provide the best of all worlds for chemistry applications that demand both high accuracy and high numerical efficiency.

36 MATERIALS SCIENCE↗

Machine Learned Hückel Theory: Interfacing Physics and Deep Neural Networks

The Hückel Hamiltonian is an incredibly simple tight-binding model known for its ability to capture qualitative physics phenomena arising from electron interactions in molecules and materials. Part of its simplicity arises from using only two types of empirically fit physics-motivated parameters: the first describes the orbital energies on each atom and the second describes electronic interactions and bonding between atoms. By replacing these empirical parameters with machine-learned dynamic values, we vastly increase the accuracy of the extended Hückel model. The dynamic values are generated with a deep neural network, which is trained to reproduce orbital energies and densities derived from density functional theory. The resulting model retains interpretability, while the deep neural network parameterization is smooth and accurate and reproduces insightful features of the original empirical parameterization. Altogether, this work shows the promise of utilizing machine learning to formulate simple, accurate, and dynamically parameterized physics models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Inexact Newton-CG algorithms with complexity guarantees

Abstract We consider variants of a recently developed Newton-CG algorithm for nonconvex problems (Royer, C. W. & Wright, S. J. (2018) Complexity analysis of second-order line-search algorithms for smooth nonconvex optimization. SIAM J. Optim., 28, 1448–1477) in which inexact estimates of the gradient and the Hessian information are used for various steps. Under certain conditions on the inexactness measures, we derive iteration complexity bounds for achieving $\epsilon $-approximate second-order optimality that match best-known lower bounds. Our inexactness condition on the gradient is adaptive, allowing for crude accuracy in regions with large gradients. We describe two variants of our approach, one in which the step size along the computed search direction is chosen adaptively, and another in which the step size is pre-defined. To obtain second-order optimality, our algorithms will make use of a negative curvature direction on some steps. These directions can be obtained, with high probability, using the randomized Lanczos algorithm. In this sense, all of our results hold with high probability over the run of the algorithm. We evaluate the performance of our proposed algorithms empirically on several machine learning models. Our approach is a first attempt to introduce inexact Hessian and/or gradient information into the Newton-CG algorithm of Royer & Wright (2018, Complexity analysis of second-order line-search algorithms for smooth nonconvex optimization. SIAM J. Optim., 28, 1448–1477).

Mathematics↗

Future Building Archetypes for Los Angeles (2100 Projection)

This dataset (Data.zip) includes empirical and machine learning-generated building information for the Los Angeles urban region. The MAv1_LA.csv file provides the baseline 2015 building data while Final_IECC_LO_2100_GAN.csv represents generative adversarial network-projected urban morphologies for the year 2100. Building archetypes were created for both datasets (Basecase_LA_Archetype.csv and LA_Simulation_2100_GAN_Archetype.csv) using footprint area as the key aggregation variable. More details about the dataset are provided in the attached readme file (README_LA_Archetype_MAv1.txt)

AutoBEM↗

Review of Solar Energetic Particle Models

Solar Energetic Particle (SEP) events are interesting from a scientific perspective as they are the product of a broad set of physical processes from the corona out through the extent of the heliosphere, and provide insight into processes of particle acceleration and transport that are widely applicable in astrophysics. From the operations perspective, SEP events pose a radiation hazard for aviation, electronics in space, and human space exploration, in particular for missions outside of the Earth’s protective magnetosphere including to the Moon and Mars. Thus, it is critical to improve the scientific understanding of SEP events and use this understanding to develop and improve SEP forecasting capabilities to support operations. Many SEP models exist or are in development using a wide variety of approaches and with differing goals. These include computationally intensive physics-based models, fast and light empirical models, machine learning-based models, and mixed-model approaches. The aim of this paper is to summarize all of the SEP models currently developed in the scientific community, including a description of model approach, inputs and outputs, free parameters, and any published validations or comparisons with data.

Kathryn Whitman↗

Machine learning models for volumetric swelling in uranium nitride

Machine learning methods are applied to predict the volumetric swelling rate of the nuclear fuel uranium nitride (UN) over various temperatures, irradiation conditions, and power densities. Both kernel-based methods and symbolic regression models for UN swelling are developed and compared with multiple experimental datasets. We find that the UN pellet geometry and dimensions must be taken into account to accurately model swelling behavior. Strong agreement is observed between the developed machine learning models and the data. The predictive error generated by the machine learning models improves on empirical models taken from the literature. Sensitivity analysis is performed to determine which properties such as temperature, burnup, and power density, are most important in the swelling process. We find that machine learning can be used to quickly develop accurate swelling models for nuclear materials. In conclusion, the presented results illustrate the potential of machine learning to determine volumetric swelling in UN.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Empirical relationships between environmental factors and soil organic carbon produce comparable prediction accuracy as the Machine Learning

Accurate representation of environmental controllers of soil organic carbon (SOC) stocks in Earth System Model (ESM) land models could reduce uncertainties in future carbon-climate feedback projections. Using empirical relationships between environmental factors and SOC stocks to evaluate land models can help modelers understand prediction biases beyond what can be achieved with the observed SOC stocks alone. In this study, we used 31 observed environmental factors, field SOC observations (n = 6,213) from the continental US, and two Machine Learning approaches [Random Forest (RF) and Generalized Additive Modeling (GAM)] to (1) select important environmental predictors of SOC stocks, (2) derive empirical relationships between environmental factors and SOC stocks, and (3) use the derived relationships to predict SOC stocks and compare the prediction accuracy of simpler model developed with the machine learning predictions. Out of the 31 environmental factors we investigated, 12 were identified as important predictors of SOC stocks by the RF approach. In contrast, the GAM approach identified six (of those 12) environmental factors as important controllers of SOC stocks: potential evapotranspiration, normalized difference vegetation index, soil drainage condition, precipitation, elevation, and net primary productivity. The GAM approach showed minimal SOC predictive importance of the remaining six environmental factors identified by the RF approach. Our derived empirical relations produced comparable prediction accuracy as the GAM and RF approach using only a subset of environmental factors. The empirical relationships we derived using the GAM approach can serve as important benchmarks to evaluate environmental control representations of SOC stocks in ESMs, which could reduce uncertainty in predicting future carbon-climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗

Quantum circuit fidelity estimation using machine learning

The computational power of real-world quantum computers is limited by errors. When using quantum computers to perform algorithms which cannot be efficiently simulated classically, it is important to quantify the accuracy with which the computation has been performed. In this work, we introduce a machine learning-based technique to estimate the fidelity between the state produced by a noisy quantum circuit and the target state corresponding to ideal noise-free computation. Our machine learning model is trained in a supervised manner, using smaller or simpler circuits for which the fidelity can be estimated using other techniques like direct fidelity estimation and quantum state tomography. Here we demonstrate that, for simulated random quantum circuits with a realistic noise model, the trained model can predict the fidelities of more complicated circuits for which such methods are infeasible. In particular, we show that the trained model may make predictions for circuits with higher degrees of entanglement than were available in the training set and that the model may make predictions for non-Clifford circuits even when the training set included only Clifford-reducible circuits. This empirical demonstration suggests classical machine learning may be useful for making predictions about beyond-classical quantum circuits for some non-trivial problems.

97 MATHEMATICS AND COMPUTING↗

Predictability and empirical dynamics of fisheries time series in the North Pacific

Previous studies have documented a strong relationship between marine ecosystems and large-scale modes of sea surface height (SSH) and sea surface temperature (SST) variability in the North Pacific such as the Pacific Decadal Oscillation and the North Pacific Gyre Oscillation. In the central and western North Pacific along the Kuroshio-Oyashio Extension (KOE), the expression of these modes in SSH and SST is linked to the propagation of long oceanic Rossby waves, which extend the predictability of the climate system to ~3 years. Using a multivariate physical-biological linear inverse model (LIM) we explore the extent to which this physical predictability leads to multi-year prediction of dominant fishery indicators inferred from three datasets (i.e., estimated biomasses, landings, and catches). We find that despite the strong autocorrelation in the fish indicators, the LIM adds dynamical forecast skill beyond persistence up to 5-6 years. By performing a sensitivity analysis of the LIM forecast model, we find that two main factors are essential for extending the dynamical predictability of the fishery indicators beyond persistence. The first is the interaction of the fishery indicators with the SST/SSH of the North and tropical Pacific. The second is the empirical relationship among the fisheries time series. This latter component reflects stock-stock interactions as well as common technological and human socioeconomic factors that may influence multiple fisheries and are captured in the training of the LIM. These results suggest that empirical dynamical models and machine learning algorithms, such as the LIM, provide an alternative and promising approach for forecasting key ecological indicators beyond the skill of persistence.

60 APPLIED LIFE SCIENCES↗

tenzing

SAND2022-3576 O tenzing provides techniques for improving the performance of key applications and libraries. The program is specified as a directed acyclic graph of operations. It uses Monte-Carlo tree search to explore the design space of possible implementations and applies machine learning to the observed empirical performance to deliver design guidance to users. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

A Qualitative Strategy for Fusion of Physics into Empirical Models for Process Anomaly Detection

To facilitate the automated online monitoring of power plants, a systematic and qualitative strategy for anomaly detection is presented. This strategy is essential to provide credible reasoning on why and when an empirical versus hybrid (i.e., physics-supported) approach should be used and to determine the ideal mix of these two approaches for a defined anomaly detection scope. Empirical methods are usually based on pattern, statistical, and causal inference. Hybrid methods include the use of physics models to train and test data methods, reduce data dimensionality, reduce data-model complexity, augment data, and reduce empirical uncertainty; hybrid methods also include the use of data to tune physics models. The presented strategy is driven by key decision points related to data relevance, simple modeling feasibility, data inference, physics-modeling value, data dimensionality, physics knowledge, method of validation, performance, data availability, and suitability for training and testing, cause-effect, entropy inference, and model fitting. The strategy is demonstrated through a pilot use case for the application of anomaly detection to capture a valve packing leak at the high-pressure coolant injection system of a nuclear power plant.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine Learning for the Validation of Expert-Elicited Causal Risk Diagrams

Exposure to spaceflight poses risk to human health in complex ways. To help manage this risk, the Human Systems Risk Board (HSRB) at the National Aeronautics and Space Administration (NASA) maintains a set of causal diagrams that attempt to explain how spaceflight hazards generate health risks and lead to adverse outcomes both in-mission, immediately post-mission, and over the long term. These causal risk diagrams are formulated as directed acyclic graphs (DAGs) and can function as knowledge graphs of connected risks and outcomes. These DAGs have proven useful for communication, and, through network analysis, have allowed for the identification of structurally important factors in the risk network. However, the utility these DAGs provide is directly proportional to their verisimilitude, making assessment of this trait using empirical data – whether from actual human spaceflight or various spaceflight analogue exposures and model organisms – a high priority. In this research we explore the use of machine learning algorithms to learn DAG structure from empirical data as a means of evaluating human-elicited DAG structures. To do so, we test several different graph structure-learning algorithms on data concerning changes in the bones of rats and mice after exposure to either spaceflight or a spaceflight analogue. We explore potential methods for indexing the similarity between each algorithm’s output DAG with all the others and with that of the expert-elicited DAG. We discuss next steps in this ongoing line of research and open science initiatives underway to complete them.

directed acyclic graphs↗

The seventh blind test of crystal structure prediction: structure ranking methods

A seventh blind test of crystal structure prediction has been organized by the Cambridge Crystallographic Data Centre. The results are presented in two parts, with this second part focusing on methods for ranking crystal structures in order of stability. The exercise involved standardized sets of structures seeded from a range of structure generation methods. Participants from 22 groups applied several periodic DFT-D methods, machine learned potentials, force fields derived from empirical data or quantum chemical calculations, and various combinations of the above. In addition, one non-energy-based scoring function was used. Results showed that periodic DFT-D methods overall agreed with experimental data within expected error margins, while one machine learned model, applying system-specific AIMnet potentials, agreed with experiment in many cases demonstrating promise as an efficient alternative to DFT-based methods. For target XXXII, a consensus was reached across periodic DFT methods, with consistently high predicted energies of experimental forms relative to the global minimum (above 4 kJ mol −1 at both low and ambient temperatures) suggesting a more stable polymorph is likely not yet observed. The calculation of free energies at ambient temperatures offered improvement of predictions only in some cases (for targets XXVII and XXXI). Several avenues for future research have been suggested, highlighting the need for greater efficiency considering the vast amounts of resources utilized in many cases.

Chemistry↗