Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29

A machine learning model for predicting the minimum miscibility pressure of CO 2 and crude oil system based on a support vector machine algorithm approach

CO 2 enhanced oil recovery (EOR) is a potential way for carbon capture, utilization and storage (CCUS). Though, the effect of CO 2 injection is greatly influenced by the reservoir conditions. Typically, Minimum miscible pressure (MMP) is selected as one of the key parameters for the screening and evaluation of prospective CO 2 flooding. Conventional slim tube test is both accurate and widely accepted but it is inefficient. Existing empirical formulas for MMPs are easy to be used but have been proved inaccurate and unreliable. Machine learning-based methods have great advantages in predicting MMP. However, only predication accuracy is discussed for most models without the screening of the main control factors and further validation of the model reliability. In this paper, a new prediction model based on support vector machine (SVM) was developed for pure/impure CO 2 and crude oil system. This study was based on 147 sets of MMP data from the literature with full information on reservoir temperature, oil composition and gas composition. The main control factors were screened by several statistical methods. Unlike the conventional prediction models that verified by only prediction accuracy, learning curve and single factor control variable analysis are further validated to obtain the optimum model.

02 PETROLEUM↗

Bridging the Gap between Cosmological Simulations with Graph Neural Networks and Domain Adaptation

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross dataset tasks (up to 28% better relative error and up to almost an order of magnitude better χ 2 ). Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

97 MATHEMATICS AND COMPUTING↗

Domain Adaptive Graph Neural Networks for Constraining Cosmological Parameters Across Multiple Data Sets

Deep learning models have been shown to outperform methods that rely on summary statistics, like the power spectrum, in extracting information from complex cosmological data sets. However, due to differences in the subgrid physics implementation and numerical approximations across different simulation suites, models trained on data from one cosmological simulation show a drop in performance when tested on another. Similarly, models trained on any of the simulations would also likely experience a drop in performance when applied to observational data. Training on data from two different suites of the CAMELS hydrodynamic cosmological simulations, we examine the generalization capabilities of Domain Adaptive Graph Neural Networks (DA-GNNs). By utilizing GNNs, we capitalize on their capacity to capture structured scale-free cosmological information from galaxy distributions. Moreover, by including unsupervised domain adaptation via Maximum Mean Discrepancy (MMD), we enable our models to extract domain-invariant features. We demonstrate that DA-GNN achieves higher accuracy and robustness on cross-dataset tasks (up to $28\%$ better relative error and up to almost an order of magnitude better $\chi^2$). Using data visualizations, we show the effects of domain adaptation on proper latent space data alignment. This shows that DA-GNNs are a promising method for extracting domain-independent cosmological information, a vital step toward robust deep learning for real cosmic survey data.

79 ASTRONOMY AND ASTROPHYSICS↗

A semblance measure for model comparison

Algorithmic and computational advances have made it possible that geophysical survey and earth model design can be aided by many systematic trial inverse-modelling runs with synthetic data. Such may, for example, come up in machine-learning approaches. Automated image appraisal pertaining to such applications will involve common statistical tests for goodness-of-data fit as a primary evaluation method. However, solution non-uniqueness may render multiple images equivalent in terms of their data fit, requiring secondary categorizers. A logical choice for classifying synthetic-imaging results quantifies the goodness of model fit where a known reference model replaces the observational input. The task of model intercomparison in terms of measuring the resemblance to the reference model poses challenges to common distance-based metrics like root mean square error and mean absolute error. First, distance-based metrics can introduce spurious contributions when smooth models with fuzzy target contours are to be compared against a sharp reference. Second, large differences due to parameter-estimation overshoots can dominate distance metrics. Here, we propose a remedy that is referred to as semblance and is based on the idea of logistic functions, where a binary-dependent variable adds non-zero or zero accumulation terms for the, respectively, passing or failing of preset target thresholds. This classifying approach is amenable to an objective where model feature recognition is primary. Numerical comparisons to distance-based metrics provide evidence for the advantages of the semblance in view of this objective. Geophysical imaging in conjunction with machine-learning is seen as a benefitting upcoming application area.

58 GEOSCIENCES↗

Detecting Quantum Critical Points of Correlated Systems by Quantum Convolutional Neural Network Using Data from Variational Quantum Eigensolver

Machine learning has been applied to a wide variety of models, from classical statistical mechanics to quantum strongly correlated systems, for classifying phase transitions. The recently proposed quantum convolutional neural network (QCNN) provides a new framework for using quantum circuits instead of classical neural networks as the backbone of classification methods. We present the results from training the QCNN by the wavefunctions of the variational quantum eigensolver for the one-dimensional transverse field Ising model (TFIM). We demonstrate that the QCNN identifies wavefunctions corresponding to the paramagnetic and ferromagnetic phases of the TFIM with reasonable accuracy. The QCNN can be trained to predict the corresponding ‘phase’ of wavefunctions around the putative quantum critical point even though it is trained by wavefunctions far away. The paper provides a basis for exploiting the QCNN to identify the quantum critical point.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE↗

Guiding the Design of Heterogeneous Electrode Microstructures for Li-Ion Batteries: Microscopic Imaging, Predictive Modeling, and Machine Learning

Electrochemical and mechanical properties of lithium-ion battery materials are heavily dependent on their 3D microstructure characteristics. A quantitative understanding of the role played by stochastic microstructures is critical for the prediction of material properties and for guiding synthesis processes. Furthermore, tailoring microstructure morphology is also a viable way of achieving optimal electrochemical and mechanical performances of lithium-ion cells. To facilitate the establishment of microstructure-resolved modeling and design methods, a review covering spatially and temporally resolved imaging of microstructure and electrochemical phenomena, microstructure statistical characterization and stochastic reconstruction, microstructure-resolved modeling for property prediction, and machine learning for microstructure design is presented here. The perspectives on the unresolved challenges and opportunities in applying experimental data, modeling, and machine learning to improve the understanding of materials and identify paths toward enhanced performance of lithium-ion cells are presented.

25 ENERGY STORAGE↗

Physics-informed machine learning of the correlation functions in bulk fluids

The Ornstein–Zernike (OZ) equation is the fundamental equation for pair correlation function computations in the modern integral equation theory for liquids. In this work, machine learning models, notably physics-informed neural networks and physics-informed neural operator networks, are explored to solve the OZ equation. The physics-informed machine learning models demonstrate great accuracy and high efficiency in solving the forward and inverse OZ problems of various bulk fluids. The results highlight the significant potential of physics-informed machine learning for applications in thermodynamic state theory.

97 MATHEMATICS AND COMPUTING↗

Neural network representation and learning of mappings and their derivatives

Discussed here are recent theorems proving that artificial neural networks are capable of approximating an arbitrary mapping and its derivatives as accurately as desired. This fact forms the basis for further results establishing the learnability of the desired approximations, using results from non-parametric statistics. These results have potential applications in robotics, chaotic dynamics, control, and sensitivity analysis. An example involving learning the transfer function and its derivatives for a chaotic map is discussed.

White, Halbert↗

Machine-Learning Assisted Identification of Battery Life Models

Predictive battery life models are commonly utilized to extrapolate degradation trends observed during accelerated aging tests for simulation of degradation in real-world applications. Thus, fitting accelerated aging data as accurately as possible and with low uncertainty is crucial for making believable projections of battery lifetime, but it is challenging to identify algebraic expressions that accurately fit multivariate degradation trends. A review of models published in literature reveal some common expressions for fitting calendar aging data, which is only dependent on temperature and state-of-charge, but no consistency across many models for fitting cycle aging data, indicating the need for a statistically rigorous data driven approach for developing empirical models. This talk will describe a machine-learning assisted method for identification of predictive battery life models utilizing bilevel optimization and symbolic regression. Bilevel optimization with cross-validation is used to statistically determine cell- and stress-dependent model parameters, while symbolic regression identifies both linear and multiplicative candidate expressions to predict stress-dependent degradation rates by selecting low-order subsets of features from a generated feature library. Because model expressions are identified empirically, it is crucial to ensure resulting models behave according to physical expectations, so the stability of models for interpolation or extrapolation is interrogated qualitatively through simulation and quantitatively through cross-validation and uncertainty quantification via bootstrap resampling. This model identification approach substantially improves upon models identified purely using expert judgement in terms of both accuracy and uncertainty. Model simulation and validation is then conducted by deriving a state-equation form of the predictive model, enabling simulation of battery aging under dynamic stresses. This enables validation of the predictive battery model on lab-based tests with varying conditions or on drive-cycle or application-cycle testing protocols. Parameter uncertainty can be carried forward into model simulation, giving lifetime estimates and confidence windows for cell- or system-level lifetime. The financial impact of battery model uncertainty can be estimated by incorporating uncertainty into a technoeconomic model.

battery↗

Monitoring and prediction of porosity in laser powder bed fusion using physics-informed meltpool signatures and machine learning

In this work we accomplished the monitoring and prediction of porosity in laser powder bed fusion (LPBF) additive manufacturing process. This objective was realized by extracting physics-informed meltpool signatures from an in-situ dual-wavelength imaging pyrometer, and subsequently, analyzing these signatures via computationally tractable machine learning approaches. Porosity in LPBF occurs despite extensive optimization of processing conditions due to stochastic causes. Hence, it is essential to continually monitor the process with in-situ sensors for detecting and mitigating incipient pore formation. In this work a tall cuboid-shaped part (10 mm × 10 mm × 137 mm, material ATI 718Plus) was built with controlled porosity by varying laser power and scanning speed. This test caused various types of porosity, such as lack-of-fusion and keyhole formation, with varying degrees of severity in the part. The meltpool was continuously monitored using a dual-wavelength imaging pyrometer installed in the machine. Physically intuitive process signatures, such as meltpool length, temperature distribution, and ejecta (spatter) characteristics, were extracted from the meltpool images. Subsequently, relatively simple machine learning models, e.g., K-Nearest Neighbors, were trained to predict both the severity and type of porosity as a function of these physics-informed meltpool signatures. These models resulted in a prediction accuracy exceeding 95% (statistical F1-score). The same analysis was carried out with a complex, black-box deep learning convolutional neural network which directly used the meltpool images instead of physics-informed features. The convolutional neural network produced a comparable F1-score in the range of 89–97%. Finally, these results demonstrate that using pragmatic, physics-informed meltpool signatures within a simple machine learning model is as effective for flaw prediction in LPBF as using a complex and computationally demanding black-box deep learning model.

36 MATERIALS SCIENCE↗

Machine learning approaches for structural and thermodynamic properties of a Lennard-Jones fluid

Predicting the functional properties of many molecular systems relies on understanding how atomistic interactions give rise to macroscale observables. However, current attempts to develop predictive models for the structural and thermodynamic properties of condensed-phase systems often rely on extensive parameter fitting to empirically selected functional forms whose effectiveness is limited to a narrow range of physical conditions. Here, we illustrate how these traditional fitting paradigms can be superseded using machine learning. Specifically, we use the results of molecular dynamics simulations to train machine learning protocols that are able to produce the radial distribution function, pressure, and internal energy of a Lennard-Jones fluid with increased accuracy in comparison to previous theoretical methods. The radial distribution function is determined using a variant of the segmented linear regression with the multivariate function decomposition approach developed by Craven et al. [J. Phys. Chem. Lett. 11, 4372 (2020)]. The pressure and internal energy are determined using expressions containing the learned radial distribution function and also a kernel ridge regression process that is trained directly on thermodynamic properties measured in simulation. The presented results suggest that the structural and thermodynamic properties of fluids may be determined more accurately through machine learning than through human-guided functional forms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ESM data downscaling: a comparison of super-resolution deep learning models

Abstract Climate projections at fine spatial resolutions are required to conduct accurate risk assessment for critical infrastructure and design adaptation planning. Generating these projections using advanced Earth system models (ESM) requires significant computational resources. To address this issue, various statistical downscaling techniques have been introduced to generate fine-resolution data from coarse-resolution simulations. In this study, we evaluate and compare five deep learning-based downscaling techniques, namely, super-resolution convolutional neural networks, fast super-resolution convolutional neural network ESM, efficient sub-pixel convolutional neural network, enhanced deep residual network (EDRN), and super-resolution generative adversarial network (SRGAN). These techniques are applied to a dataset generated by the Energy Exascale Earth System Model (E3SM), focusing on key surface variables such as surface temperature, shortwave heat flux, and longwave heat flux. Models are trained and validated using paired fine-resolution (0.25 $$^{\circ }$$ ∘ ) and coarse-resolution (1 $$^{\circ }$$ ∘ ) monthly data obtained from a 9-year simulation. Next, blind testing is performed using monthly data obtained from two different years outside of the training and validation set. To evaluate the efficiency of each technique, different statistical metrics are used, including mean squared error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The results show that EDRN outperforms other algorithms in terms of PSNR, SSIM, and MSE, but struggles to capture fine-scale features in the data. In contrast, SRGAN, a generative model that uses perceptual loss, excels in capturing fine details at boundaries and internal structures, resulting in lower LPIPS than other methods.

Pawar, Nikhil M. (ORCID:0000000211613289)↗

Avalanches and the distribution of solar flares

The solar coronal magnetic field is proposed to be in a self-organized critical state, thus explaining the observed power-law dependence of solar-flare-occurrence rate on flare size which extends over more than five orders of magnitude in peak flux. The physical picture that arises is that solar flares are avalanches of many small reconnection events, analogous to avalanches of sand in the models published by Bak and colleagues in 1987 and 1988. Flares of all sizes are manifestations of the same physical processes, where the size of a given flare is determined by the number of elementary reconnection events. The relation between small-scale processes and the statistics of global-flare properties which follows from the self-organized magnetic-field configuration provides a way to learn about the physics of the unobservable small-scale reconnection processes. A simple lattice-reconnection model is presented which is consistent with the observed flare statistics. The implications for coronal heating are discussed and some observational tests of this picture are given.

Lu, Edward T.↗

Machine learning for single-ended event reconstruction in PROSPECT experiment

The Precision Reactor Oscillation and Spectrum Experiment, PROSPECT, was a segmented antineutrino detector that successfully operated at the High Flux Isotope Reactor in Oak Ridge, TN, during its 2018 run. Despite challenges with photomultiplier tube base failures affecting some segments, innovative machine learning approaches were employed to perform position and energy reconstruction, and particle classification. This work highlights the effectiveness of convolutional neural networks and graph convolutional networks in enhancing data analysis. By leveraging these techniques, a 3.3% increase in effective statistics was achieved compared to traditional methods, showcasing their potential to improve analysis performance. Furthermore, these machine learning methodologies offer promising applications for other segmented particle detectors, underscoring their versatility and impact.

47 OTHER INSTRUMENTATION↗

Effect of Solvent on the Local Structure, Dynamics, and Vibrational Density of States in Sn-BEA Zeolite

Lewis acid zeolites are attractive catalysts for epoxidation and biomass valorization, as they are highly active and selective in the liquid phase and can operate at or near ambient conditions. While a rich experimental literature exists on liquid-phase Lewis acid zeolite catalysis, our understanding of the molecular organization and solvent dynamics in the vicinity of Lewis acid sites with differing metal site speciation remains limited. In this work, we investigate the molecular coordination and diffusion of two common solvents (methanol and water) around the closed and open Sn-BEA zeolite active sites using molecular dynamics simulations with a machine-learned interatomic potential trained on ab initio molecular dynamics trajectories. Molecular dynamics simulations reveal that introducing active sites significantly enhances local order in the first and second solvation shells compared to the pure silica case. For methanol, both closed and open active sites are singly coordinated, while more than two water molecules coordinate the open site. In contrast to methanol, we observed that water molecules dissociate, leading to the formation of additional Sn-OH and silanol groups away from the active site. The diffusion coefficients of water and methanol are functions of the solvent population in the pore. Here, our work provides insights into how active site speciation in Lewis acid zeolites affects solvent coordination, diffusion, and vibrational signature. This information is foundational for catalyst design and optimization of liquid-phase catalytic processes in zeolites. It also demonstrates the suitability of machine-learned interatomic potentials for modeling reactive systems, enabling sufficiently long trajectories for appropriate statistical averaging.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Active Learning with Rationales for Identifying Operationally Significant Anomalies in Aviation

A major focus of the commercial aviation community is discovery of unknown safety events in flight operations data. Data-driven unsupervised anomaly detection methods are better at capturing unknown safety events compared to rule-based methods which only look for known violations. However, not all statistical anomalies that are discovered by these unsupervised anomaly detection methods are operationally significant (e.g., represent a safety concern). Subject Matter Experts (SMEs) have to spend significant time reviewing these statistical anomalies individually to identify a few operationally significant ones. In this paper we propose an active learning algorithm that incorporates SME feedback in the form of rationales to build a classifier that can distinguish between uninteresting and operationally significant anomalies. Experimental evaluation on real aviation data shows that our approach improves detection of operationally significant events by as much as 75% compared to the state-of-the-art. The learnt classifier also generalizes well to additional validation data sets.

anomaly detection↗