Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Analysis of leading edge protection application on wind turbine performance through energy and power decomposition approaches

Abstract Wind power production is driven by, and varies with, the stochastic yet uncontrollable wind and environmental inputs. To compare a wind turbine's performance, a direct comparison on power outputs is always confounded by the stochastic effect of weather inputs. It is therefore crucial to control for the weather and environmental influence. Toward that objective, our study proposes an energy decomposition approach. We start with comparing the change in the total energy production and refer to the change in total energy as delta energy. On this delta energy, we apply our decomposition method, which is to separate the portion of energy change due to weather effects from that due to the turbine itself. We derive a set of mathematical relationships allowing us to perform this decomposition and examine the credibility and robustness of the proposed decomposition approach through extensive cross‐validation and case studies. We then apply the decomposition approach to Supervisory Control and Data Acquisition data associated with several wind turbines to which leading‐edge protection was carried out. Our study shows that the leading‐edge protection applied on blades may cause a small decline to the power production efficiency in the short term, although we expect the leading‐edge protection to benefit the blade's reliability in the long term.

17 WIND ENERGY↗

Snowmass2021 theory frontier white paper: Astrophysical and cosmological probes of dark matter

While astrophysical and cosmological probes provide a remarkably precise and consistent picture of the quantity and general properties of dark matter, its fundamental nature remains one of the most significant open questions in physics. Obtaining a more comprehensive understanding of dark matter within the next decade will require overcoming a number of theoretical challenges: the groundwork for these strides is being laid now, yet much remains to be done. Chief among the upcoming challenges is establishing the theoretical foundation needed to harness the full potential of new observables in the astrophysical and cosmological domains, spanning the early Universe to the inner portions of galaxies and the stars therein. Identifying the nature of dark matter will also entail repurposing and implementing a wide range of theoretical techniques from outside the typical toolkit of astrophysics, ranging from effective field theory to the dramatically evolving world of machine learning and artificial-intelligence-based statistical inference. Through this work, the theory frontier will be at the heart of dark matter discoveries in the upcoming decade.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Detect and correct bias in multi-site neuroimaging datasets

The desire to train complex machine learning algorithms and to increase the statistical power in association studies drives neuroimaging research to use ever-larger datasets. The most obvious way to increase sample size is by pooling scans from independent studies. However, simple pooling is often ill-advised as selection, measurement, and confounding biases may creep in and yield spurious correlations. In this work, we combine 35,320 magnetic resonance images of the brain from 17 studies to examine bias in neuroimaging. In the first experiment, Name That Dataset, we provide empirical evidence for the presence of bias by showing that scans can be correctly assigned to their respective dataset with 71.5% accuracy. Given such evidence, we take a closer look at confounding bias, which is often viewed as the main shortcoming in observational studies. In practice, we neither know all potential confounders nor do we have data on them. Hence, we model confounders as unknown, latent variables. Kolmogorov complexity is then used to decide whether the confounded or the causal model provides the simplest factorization of the graphical model. Finally, we present methods for dataset harmonization and study their ability to remove bias in imaging features. In particular, we propose an extension of the recently introduced ComBat algorithm to control for global variation across image features, inspired by adjusting for unknown population stratification in genetics. Overall, our results demonstrate that harmonization can reduce dataset-specific information in image features. Further, confounding bias can be reduced and even turned into a causal relationship. However, harmonization also requires caution as it can easily remove relevant subject-specific information. Code is available at https://github.com/ai-med/Dataset-Bias.

42 ENGINEERING↗

Reconstruction and uncertainty quantification of lattice Hamiltonian model parameters from observations of microscopic degrees of freedom

The emergence of scanning probe and electron beam imaging techniques has allowed quantitative studies of atomic structure and minute details of electronic and vibrational structure on the level of individual atomic units. These microscopic descriptors, in turn, can be associated with local symmetry breaking phenomena, representing the stochastic manifestation of the underpinning generative physical model. In this work, we explore the reconstruction of exchange integrals in the Hamiltonian for a lattice model with two competing interactions from observations of microscopic degrees of freedom and establish the uncertainties and reliability of such analysis in a broad parameter-temperature space. In contrast to other approaches, we specifically specify a loss function inherent to thermodynamic systems and utilize it to estimate uncertainty in simulated realizations of different models. As an ancillary task, we develop a machine learning approach based on histogram clustering to predict phase diagrams efficiently using a reduced descriptor space. We further demonstrate that reconstruction is possible well above the phase transition and in the regions of parameter space when the macroscopic ground state of the system is poorly defined due to frustrated interactions. This suggests that this approach can be applied to the traditionally complex problems of condensed matter physics such as ferroelectric relaxors and morphotropic phase boundary systems, spin and cluster glasses, and quantum systems once the local descriptors linked to the relevant physical behaviors are known.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Interpretation of autoencoder-learned collective variables using Morse–Smale complex and sublevelset persistent homology: An application on molecular trajectories

Dimensionality reduction often serves as the first step toward a minimalist understanding of physical systems as well as the accelerated simulations of them. In particular, neural network-based nonlinear dimensionality reduction methods, such as autoencoders, have shown promising outcomes in uncovering collective variables (CVs). However, the physical meaning of these CVs remains largely elusive. In this work, we constructed a framework that (1) determines the optimal number of CVs needed to capture the essential molecular motions using an ensemble of hierarchical autoencoders and (2) provides topology-based interpretations to the autoencoder-learned CVs with Morse–Smale complex and sublevelset persistent homology. Furthermore, this approach was exemplified using a series of n-alkanes and can be regarded as a general, explainable nonlinear dimensionality reduction method.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

What drives the scatter of local star-forming galaxies in the BPT diagrams? A Machine Learning based analysis

ABSTRACT We investigate which physical properties are most predictive of the position of local star forming galaxies on the BPT diagrams, by means of different Machine Learning (ML) algorithms. Exploiting the large statistics from the Sloan Digital Sky Survey (SDSS), we define a framework in which the deviation of star-forming galaxies from their median sequence can be described in terms of the relative variations in a variety of observational parameters. We train artificial neural networks (ANN) and random forest (RF) trees to predict whether galaxies are offset above or below the sequence (via classification), and to estimate the exact magnitude of the offset itself (via regression). We find, with high significance, that parameters primarily associated to variations in the nitrogen-over-oxygen abundance ratio (N/O) are the most predictive for the [N ii]-BPT diagram, whereas properties related to star formation (like variations in SFR or EW(H α)) perform better in the [S ii]-BPT diagram. We interpret the former as a reflection of the N/O–O/H relationship for local galaxies, while the latter as primarily tracing the variation in the effective size of the S+ emitting region, which directly impacts the [S ii] emission lines. This analysis paves the way to assess to what extent the physics shaping local BPT diagrams is also responsible for the offsets seen in high redshift galaxies or, instead, whether a different framework or even different mechanisms need to be invoked.

79 ASTRONOMY AND ASTROPHYSICS↗

The entropy of galaxy spectra: how much information is encoded?

Abstract The inverse problem of extracting the stellar population content of galaxy spectra is analysed here from a basic standpoint based on information theory. By interpreting spectra as probability distribution functions, we find that galaxy spectra have high entropy, thus leading to a rather low effective information content. The highest variation in entropy is unsurprisingly found in regions that have been well studied for decades with the conventional approach. We target a set of six spectral regions that show the highest variation in entropy – the 4000 Å break being the most informative one. As a test case with real data, we measure the entropy of a set of high-quality spectra from the Sloan Digital Sky Survey, and contrast entropy-based results with the traditional method based on line strengths. The data are classified into star-forming (SF), quiescent (Q), and active galactic nucleus (AGN) galaxies, and show – independently of any physical model – that AGN spectra can be interpreted as a transition between SF and Q galaxies, with SF galaxies featuring a more diverse variation in entropy. The high level of entanglement complicates the determination of population parameters in a robust, unbiased way, and affects traditional methods that compare models with observations, as well as machine learning (especially deep learning) algorithms that rely on the statistical properties of the data to assess the variations among spectra. Entropy provides a new avenue to improve population synthesis models so that they give a more faithful representation of real galaxy spectra.

Ferreras, Ignacio (ORCID:0000000345843127)↗

Lithium-Ion Battery Life Model with Electrode Cracking and Early-Life Break-in Processes

This paper develops a physically justified reduced-order capacity fade model from accelerated calendar- and cycle-aging data for 32 lithium-ion (Li-ion) graphite/nickel-manganese-cobalt (NMC) cells. The large data set reveals temperature-, charge C-rate-, depth-of-discharge-, and state of charge (SOC)-dependent degradation patterns that would be unobserved in a smaller test matrix. Model structure is informed by incremental capacity analysis that shows loss of lithium inventory and cathode-material loss as the dominant capacity fade mechanisms. The model includes terms attributable to solid-electrolyte interface (SEI) growth, electrode cracking, cycling-driven acceleration of SEI growth, and "break-in" mechanisms that slightly decrease or increase available Li inventory early in life. The study explores what mathematical couplings of these mechanisms best describe calendar aging, cycle aging, and mixed calendar/cycle aging. Various approaches are discussed for extracting relevant stress factors from complex cycling profiles to predict lifetime during real-world battery loads using models trained on constant-current laboratory test results. The complexity of the present human-driven model identification process motivates future work in machine learning to more widely search and statistically discern the optimal model that correctly extrapolates capacity fade based on physical knowledge.

25 ENERGY STORAGE↗

An Open Benchmark of One Million High-Fidelity Cislunar Trajectories

Cislunar space spans from geosynchronous altitudes to beyond the Moon and will underpin future exploration, science, and security operations. We describe and release an open dataset of one million numerically propagated cislunar trajectories generated with the open-source Space Situational Awareness Python package (SSAPy). The model includes high-degree Earth/Moon gravity, solar gravity, and Earth/Sun radiation pressure; other planetary gravities are omitted by design for computational efficiency. Initial conditions uniformly sample commonly used osculating-element ranges, and each trajectory is propagated for up to six years under a single, fixed start epoch. The dataset is intended as a reusable benchmark for method development (e.g., space domain awareness, navigation, and machine-learning pipelines), a reference library for statistical studies of orbit families, and a starting point for community-driven extensions (e.g., alternative epochs). We report empirically observed stability trends (e.g., a band near ~5 GEO and persistence of some co-orbital classes including L4/L5 librators) as dataset descriptors rather than new dynamical results. The chief contribution is the scale, fidelity, organization (CSV/HDF5 with full state time series and metadata), and open availability, which together lower the barrier to comparative and data-driven studies in the cislunar regime.

79 ASTRONOMY AND ASTROPHYSICS↗

Systems and methods for binary code analysis

Human-readable (HR) code may be derived from a binary. The HR code may be configured to have statistical properties suitable for machine-learned (ML) translation. The HR code may comprise source code, intermediate code, assembly code, or the like. A machine-learned translator may be configured to translate the HR code into labels comprising semantic information pertaining to respective functions of the binary, such as a function name, role, or the like. Execution of the binary may be blocked in response to translating the HR code to a label associated with malware, such as cryptocurrency mining malware or the like. Conversely, the binary may be permitted to proceed to execution in response to determining that the translation is free from labels indicative of malware.

Anderson, Matthew W.↗

Predicting battery capacity from impedance at varying temperature and state of charge using machine learning

Prediction of battery health from electrochemical impedance spectroscopy (EIS) data can enable rapid measurement of battery state in real-world applications without using additional sensors or time-consuming performance measurements. However, deconvoluting the effect of capacity, state of charge, and temperature on EIS response is complicated analytically. Here, various machine-learning models, such as linear, Gaussian process, random forest, and artificial neural network regression, are utilized to predict capacity from EIS using hundreds of capacity, direct current (DC) resistance, and EIS measurements recorded under varying conditions of health, temperature, and state of charge (SOC). Several feature extraction and selection methods from traditional electrochemical analysis and statistical modeling are explored using machine-learning pipelines. EIS data from just two frequencies can accurately predict capacity, and interrogation shows that the optimal set of frequencies is not usually intuitive. Best results are achieved with an ensemble model, which predicts battery capacity with a mean absolute error of 1.9% on data from unobserved cells.

25 ENERGY STORAGE↗

Simulated wildfire burned area over the CONUS during 2001-2020

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM). A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

Liu, Ye↗

Improving North American Wildfire Prediction by Integrating a Machine-Learning Fire Model in a Land Surface Model

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

54 ENVIRONMENTAL SCIENCES↗

Learning process mapping heuristics under stochastic sampling overheads

A statistical method was developed previously for improving process mapping heuristics. The method systematically explores the space of possible heuristics under a specified time constraint. Its goal is to get the best possible heuristics while trading between the solution quality of the process mapping heuristics and their execution time. The statistical selection method is extended to take into consideration the variations in the amount of time used to evaluate heuristics on a problem instance. The improvement in performance is presented using the more realistic assumption along with some methods that alleviate the additional complexity.

Ieumwananonthachai, Arthur↗

Accelerating high-strain continuum-scale brittle fracture simulations with machine learning

Failure in brittle materials under dynamic loading conditions is a result of the propagation and coalescence of microcracks. Simulating this discrete crack evolution at the continuum level is computationally expensive or, in some cases, intractable, resulting in the need to make broad assumptions or neglect key physics. In this work, we have developed an approach using machine learning that overcomes the current inability to represent meso-scale physics at the macro-scale. Our approach leverages damage and stress data from a computationally expensive high-fidelity model that explicitly resolves microcrack behavior to build an inexpensive machine learning emulator. Once trained, the machine learning emulator is used to predict the evolution of crack length statistics, which then informs a continuum-scale constitutive model. This results in a significant speed-up of the workflow by four orders of magnitude. Both the machine learning emulator and the continuum-scale model are validated against the high-fidelity model and experimental data, respectively, showing excellent agreement. There are two key findings. The first is that we can reduce the dimensionality of the problem, establishing that the machine learning emulator only needs the length of the longest crack and one of the maximum stress components to capture the necessary physics. Another compelling finding is that the emulator can be trained in one experimental setting and transferred successfully to predict behavior in a different setting.

36 MATERIALS SCIENCE↗

Ambient-temperature liquid jet targets for high-repetition-rate HED discovery science

High-power lasers can generate energetic particle beams and astrophysically relevant pressure and temperature states in the high-energy-density (HED) regime. Recently-commissioned high-repetition-rate (HRR) laser drivers are capable of producing these conditions at rates exceeding 1 Hz. However, experimental output from these systems is often limited by the difficulty of designing targets that match these repetition rates. To overcome this challenge, we have developed tungsten microfluidic nozzles, which produce a continuously replenishing jet that operates at flow speeds of approximately 10 m/s and can sustain shot frequencies up to 1 kHz. The ambient-temperature planar liquid jets produced by these nozzles can have thicknesses ranging from hundreds of nanometers to tens of micrometers. In this work, we illustrate the operational principle of the microfluidic nozzle and describe its implementation in a vacuum environment. Further, we provide evidence of successful laser-driven ion acceleration using this target and discuss the prospect of optimizing the ion acceleration performance through an in situ jet thickness scan. Future applications for the jet throughout HED science include shock compression and studies of strongly heated nonequilibrium plasmas. When fielded in concert with HRR-compatible laser, diagnostic, and active feedback technology, this target will facilitate advanced automated studies in HRR HED science, including machine learning-based optimization and high-dimensional statistical analysis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗