Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Intensity of sample processing methods impacts wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes

Wastewater SARS-CoV-2 surveillance has been deployed since the beginning of the COVID-19 pandemic to monitor the dynamics in virus burden in local communities. Genomic surveillance of SARS-CoV-2 in wastewater, particularly efforts aimed at whole genome sequencing for variant tracking and identification, are still challenging due to low target concentration, complex microbial and chemical background, and lack of robust nucleic acid recovery experimental procedures. The intrinsic sample limitations are inherent to wastewater and are thus unavoidable. Here, we use a statistical approach that couples correlation analyses to a random forest-based machine learning algorithm to evaluate potentially important factors associated with wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes, with a specific focus on the breadth of genome coverage. We collected 182 composite and grab wastewater samples from the Chicago area between November 2020 to October 2021. Samples were processed using a mixture of processing methods reflecting different homogenization intensities (HA + Zymo beads, HA + glass beads, and Nanotrap), and were sequenced using one of the two library preparation kits (the Illumina COVIDseq kit and the QIAseq DIRECT kit). Technical factors evaluated using statistical and machine learning approaches include sample types, certain sample intrinsic features, and processing and sequencing methods. The results suggested that sample processing methods could be a predominant factor affecting sequencing outcomes, and library preparation kits was considered a minor factor. Finally, a synthetic SARS-CoV-2 RNA spike-in experiment was performed to validate the impact from processing methods and suggested that the intensity of the processing methods could lead to different RNA fragmentation

60 APPLIED LIFE SCIENCES↗

Robust Carbon Dioxide Plume Imaging Using Joint Tomographic Inversion of Seismic Onset Time and Distributed Pressure and Temperature Measurements (Final Report)

We develop and demonstrate rapid and cost-effective methodologies for spatiotemporal tracking of CO2 plumes during geologic sequestration using joint inversion of seismic data and distributed pressure and temperature measurements. Key elements of our methodology are: (a) a computationally efficient approach to pressure and temperature propagation, (b) analysis of time lapse seismic data using a novel ‘seismic onset time’ approach to detect fluid front propagation, and (c) data assimilation and uncertainty assessment via joint inversion of pressure, temperature and time lapse seismic data, and (d) validating the numerical tomographic inversion using a CO2 injection demonstration projects, specifically data collected from the from the Petra Nova Parish Holdings CCUS project in the West Ranch Field, Texas and the Chester-16 reef CO2 injection site in Northern Michigan which is part of the DOE Midwestern Carbon Sequestration Project. The research team is led by Texas A&M University and includes Battelle as a subcontractor with support from Shell, Anadarko, Chevron and JX Nippon. A carbon dioxide (CO2) water-alternating-gas (WAG) pilot was conducted to gain insights into tertiary oil recovery potential via CO2 flood in the West Ranch Field as part of the Petra Nova project, the world’s largest post-combustion CO2 capture and utilization initiative. With a fluvial formation geology and large contrasts in permeability, this is a challenging and novel application of CO2 enhanced oil recovery (EOR). We build a predictive dynamic model of the subsurface that incorporates the multiphase and compositional data acquired during the pilot operation. The calibrated model is used for the carbon dioxide plume imaging. The study began with an initialization of the pilot sector model extracted from a calibrated full-field model. The pilot model calibration follows a two-step hierarchical workflow. First, we performed a large-scale update of the permeability distribution by integrating available bottomhole pressure and multiphase production data. In the second step, local permeability field is fine-tuned using a streamline-based method to match CO2 breakthrough times at the producers. The predictive capability of the calibrated model was verified through two blind validation tests: (1) the model showed good agreement with saturation logs acquired at two observation wells; and (2) the model reproduced the CO2 recovery as a fraction of the injected CO2. The use of seismic onset times has shown great promise for integrating near-continuous seismic surveys for updating geologic models. In this study, we analyze the impact of seismic survey frequency on the onset time approach aiming to extend the application of onset time to infrequent seismic surveys. In addition, we quantitatively examine the nonlinearity of the onset time method and compare it to the commonly used amplitude inversion method. We carry out a sensitivity analysis of seismic survey frequency based on the complete seismic survey data (over 175 surveys) of steam injection in a heavy oil reservoir (Peace River Unit) in Canada. Our results show that an adequate onset time map can be obtained from the infrequent seismic surveys by interpolation between seismic surveys as long as there is no change in the dominant underlying physics between the successive surveys. The study also shows that nonlinearity of the onset time method can be -smaller than that of the amplitude inversion method by several orders of magnitude. Application to the Brugge benchmark case shows that the onset time method obtains comparable permeability update as the traditional seismic amplitude inversion method with faster computation and improved convergence characteristics. We extend the streamline-based data integration approach to incorporate distributed temperature sensor (DTS) data using the concept of thermal tracer travel time. Then, a hierarchical workflow composed of evolutionary and streamline methods is employed to jointly history match the DTS and pressure data. Finally, CO2 saturation and streamline maps are used to visualize the CO2 plume movement during the sequestration process. The hierarchical workflow is applied to a carbon sequestration project in a carbonate reef reservoir within the Northern Niagaran Pinnacle Reef Trend in Michigan, USA. The monitoring data set consists of distributed temperature sensing (DTS) data acquired at the injection well and a monitoring well, flowing bottom-hole pressure data at the injection well, and time-lapse pressure measurements at several locations along the monitoring well. The history matching results indicate that the CO2 movement is mostly restricted to the intended zones of injection which is consistent with an independent warm-back analysis of the temperature data. In addition to employing simulation models and inverse methods for CO2 plume imaging, we also initialized a data-driven technology for detecting inter-well connectivity based on production and pressure data. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO2 EOR projects utilizing the water-alternating-gas (WAG) process. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. Texas A&M University, the lead organization in the project, was primarily responsible for the development of tomographic approaches for CO2 plume mapping in conjunction with distributed pressure, temperature and seismic onset time data. Battelle, as a subcontractor, was primarily responsible for the development of analytical and empirical methods for analyzing transient injection rate and pressure data from point/line sources such as injection and monitoring wells. An additional area of emphasis for Battelle was the use of machine learning for such tasks as inferring reservoir connectivity information from injection-production data, and identifying variable importance for machine learning-based proxy models developed from full-physics simulations. The two organizations also collaborated on the application of the tomographic inversion methodology for a field data set.

02 PETROLEUM↗

Unstructured clinical notes within the 24 hours since admission predict short, mid & long-term mortality in adult ICU patients

Mortality prediction for intensive care unit (ICU) patients is crucial for improving outcomes and efficient utilization of resources. Accessibility of electronic health records (EHR) has enabled data-driven predictive modeling using machine learning. However, very few studies rely solely on unstructured clinical notes from the EHR for mortality prediction. In this work, we propose a framework to predict short, mid, and long-term mortality in adult ICU patients using unstructured clinical notes from the MIMIC III database, natural language processing (NLP), and machine learning (ML) models. Depending on the statistical description of the patients’ length of stay, we define the short-term as 48-hour and 4-day period, the mid-term as 7-day and 10-day period, and the long-term as 15-day and 30-day period after admission. We found that by only using clinical notes within the 24 hours of admission, our framework can achieve a high area under the receiver operating characteristics (AU-ROC) score for short, mid and long-term mortality prediction tasks. The test AU-ROC scores are 0.87, 0.83, 0.83, 0.82, 0.82, and 0.82 for 48-hour, 4-day, 7-day, 10-day, 15-day, and 30-day period mortality prediction, respectively. We also provide a comparative study among three types of feature extraction techniques from NLP: frequency-based technique, fixed embedding-based technique, and dynamic embedding-based technique. Lastly, we provide an interpretation of the NLP-based predictive models using feature-importance scores.

60 APPLIED LIFE SCIENCES↗

Analysis of leading edge protection application on wind turbine performance through energy and power decomposition approaches

Abstract Wind power production is driven by, and varies with, the stochastic yet uncontrollable wind and environmental inputs. To compare a wind turbine's performance, a direct comparison on power outputs is always confounded by the stochastic effect of weather inputs. It is therefore crucial to control for the weather and environmental influence. Toward that objective, our study proposes an energy decomposition approach. We start with comparing the change in the total energy production and refer to the change in total energy as delta energy. On this delta energy, we apply our decomposition method, which is to separate the portion of energy change due to weather effects from that due to the turbine itself. We derive a set of mathematical relationships allowing us to perform this decomposition and examine the credibility and robustness of the proposed decomposition approach through extensive cross‐validation and case studies. We then apply the decomposition approach to Supervisory Control and Data Acquisition data associated with several wind turbines to which leading‐edge protection was carried out. Our study shows that the leading‐edge protection applied on blades may cause a small decline to the power production efficiency in the short term, although we expect the leading‐edge protection to benefit the blade's reliability in the long term.

17 WIND ENERGY↗

Snowmass2021 theory frontier white paper: Astrophysical and cosmological probes of dark matter

While astrophysical and cosmological probes provide a remarkably precise and consistent picture of the quantity and general properties of dark matter, its fundamental nature remains one of the most significant open questions in physics. Obtaining a more comprehensive understanding of dark matter within the next decade will require overcoming a number of theoretical challenges: the groundwork for these strides is being laid now, yet much remains to be done. Chief among the upcoming challenges is establishing the theoretical foundation needed to harness the full potential of new observables in the astrophysical and cosmological domains, spanning the early Universe to the inner portions of galaxies and the stars therein. Identifying the nature of dark matter will also entail repurposing and implementing a wide range of theoretical techniques from outside the typical toolkit of astrophysics, ranging from effective field theory to the dramatically evolving world of machine learning and artificial-intelligence-based statistical inference. Through this work, the theory frontier will be at the heart of dark matter discoveries in the upcoming decade.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Detect and correct bias in multi-site neuroimaging datasets

The desire to train complex machine learning algorithms and to increase the statistical power in association studies drives neuroimaging research to use ever-larger datasets. The most obvious way to increase sample size is by pooling scans from independent studies. However, simple pooling is often ill-advised as selection, measurement, and confounding biases may creep in and yield spurious correlations. In this work, we combine 35,320 magnetic resonance images of the brain from 17 studies to examine bias in neuroimaging. In the first experiment, Name That Dataset, we provide empirical evidence for the presence of bias by showing that scans can be correctly assigned to their respective dataset with 71.5% accuracy. Given such evidence, we take a closer look at confounding bias, which is often viewed as the main shortcoming in observational studies. In practice, we neither know all potential confounders nor do we have data on them. Hence, we model confounders as unknown, latent variables. Kolmogorov complexity is then used to decide whether the confounded or the causal model provides the simplest factorization of the graphical model. Finally, we present methods for dataset harmonization and study their ability to remove bias in imaging features. In particular, we propose an extension of the recently introduced ComBat algorithm to control for global variation across image features, inspired by adjusting for unknown population stratification in genetics. Overall, our results demonstrate that harmonization can reduce dataset-specific information in image features. Further, confounding bias can be reduced and even turned into a causal relationship. However, harmonization also requires caution as it can easily remove relevant subject-specific information. Code is available at https://github.com/ai-med/Dataset-Bias.

42 ENGINEERING↗

Reconstruction and uncertainty quantification of lattice Hamiltonian model parameters from observations of microscopic degrees of freedom

The emergence of scanning probe and electron beam imaging techniques has allowed quantitative studies of atomic structure and minute details of electronic and vibrational structure on the level of individual atomic units. These microscopic descriptors, in turn, can be associated with local symmetry breaking phenomena, representing the stochastic manifestation of the underpinning generative physical model. In this work, we explore the reconstruction of exchange integrals in the Hamiltonian for a lattice model with two competing interactions from observations of microscopic degrees of freedom and establish the uncertainties and reliability of such analysis in a broad parameter-temperature space. In contrast to other approaches, we specifically specify a loss function inherent to thermodynamic systems and utilize it to estimate uncertainty in simulated realizations of different models. As an ancillary task, we develop a machine learning approach based on histogram clustering to predict phase diagrams efficiently using a reduced descriptor space. We further demonstrate that reconstruction is possible well above the phase transition and in the regions of parameter space when the macroscopic ground state of the system is poorly defined due to frustrated interactions. This suggests that this approach can be applied to the traditionally complex problems of condensed matter physics such as ferroelectric relaxors and morphotropic phase boundary systems, spin and cluster glasses, and quantum systems once the local descriptors linked to the relevant physical behaviors are known.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Interpretation of autoencoder-learned collective variables using Morse–Smale complex and sublevelset persistent homology: An application on molecular trajectories

Dimensionality reduction often serves as the first step toward a minimalist understanding of physical systems as well as the accelerated simulations of them. In particular, neural network-based nonlinear dimensionality reduction methods, such as autoencoders, have shown promising outcomes in uncovering collective variables (CVs). However, the physical meaning of these CVs remains largely elusive. In this work, we constructed a framework that (1) determines the optimal number of CVs needed to capture the essential molecular motions using an ensemble of hierarchical autoencoders and (2) provides topology-based interpretations to the autoencoder-learned CVs with Morse–Smale complex and sublevelset persistent homology. Furthermore, this approach was exemplified using a series of n-alkanes and can be regarded as a general, explainable nonlinear dimensionality reduction method.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

What drives the scatter of local star-forming galaxies in the BPT diagrams? A Machine Learning based analysis

ABSTRACT We investigate which physical properties are most predictive of the position of local star forming galaxies on the BPT diagrams, by means of different Machine Learning (ML) algorithms. Exploiting the large statistics from the Sloan Digital Sky Survey (SDSS), we define a framework in which the deviation of star-forming galaxies from their median sequence can be described in terms of the relative variations in a variety of observational parameters. We train artificial neural networks (ANN) and random forest (RF) trees to predict whether galaxies are offset above or below the sequence (via classification), and to estimate the exact magnitude of the offset itself (via regression). We find, with high significance, that parameters primarily associated to variations in the nitrogen-over-oxygen abundance ratio (N/O) are the most predictive for the [N ii]-BPT diagram, whereas properties related to star formation (like variations in SFR or EW(H α)) perform better in the [S ii]-BPT diagram. We interpret the former as a reflection of the N/O–O/H relationship for local galaxies, while the latter as primarily tracing the variation in the effective size of the S+ emitting region, which directly impacts the [S ii] emission lines. This analysis paves the way to assess to what extent the physics shaping local BPT diagrams is also responsible for the offsets seen in high redshift galaxies or, instead, whether a different framework or even different mechanisms need to be invoked.

79 ASTRONOMY AND ASTROPHYSICS↗

The entropy of galaxy spectra: how much information is encoded?

Abstract The inverse problem of extracting the stellar population content of galaxy spectra is analysed here from a basic standpoint based on information theory. By interpreting spectra as probability distribution functions, we find that galaxy spectra have high entropy, thus leading to a rather low effective information content. The highest variation in entropy is unsurprisingly found in regions that have been well studied for decades with the conventional approach. We target a set of six spectral regions that show the highest variation in entropy – the 4000 Å break being the most informative one. As a test case with real data, we measure the entropy of a set of high-quality spectra from the Sloan Digital Sky Survey, and contrast entropy-based results with the traditional method based on line strengths. The data are classified into star-forming (SF), quiescent (Q), and active galactic nucleus (AGN) galaxies, and show – independently of any physical model – that AGN spectra can be interpreted as a transition between SF and Q galaxies, with SF galaxies featuring a more diverse variation in entropy. The high level of entanglement complicates the determination of population parameters in a robust, unbiased way, and affects traditional methods that compare models with observations, as well as machine learning (especially deep learning) algorithms that rely on the statistical properties of the data to assess the variations among spectra. Entropy provides a new avenue to improve population synthesis models so that they give a more faithful representation of real galaxy spectra.

Ferreras, Ignacio (ORCID:0000000345843127)↗

Lithium-Ion Battery Life Model with Electrode Cracking and Early-Life Break-in Processes

This paper develops a physically justified reduced-order capacity fade model from accelerated calendar- and cycle-aging data for 32 lithium-ion (Li-ion) graphite/nickel-manganese-cobalt (NMC) cells. The large data set reveals temperature-, charge C-rate-, depth-of-discharge-, and state of charge (SOC)-dependent degradation patterns that would be unobserved in a smaller test matrix. Model structure is informed by incremental capacity analysis that shows loss of lithium inventory and cathode-material loss as the dominant capacity fade mechanisms. The model includes terms attributable to solid-electrolyte interface (SEI) growth, electrode cracking, cycling-driven acceleration of SEI growth, and "break-in" mechanisms that slightly decrease or increase available Li inventory early in life. The study explores what mathematical couplings of these mechanisms best describe calendar aging, cycle aging, and mixed calendar/cycle aging. Various approaches are discussed for extracting relevant stress factors from complex cycling profiles to predict lifetime during real-world battery loads using models trained on constant-current laboratory test results. The complexity of the present human-driven model identification process motivates future work in machine learning to more widely search and statistically discern the optimal model that correctly extrapolates capacity fade based on physical knowledge.

25 ENERGY STORAGE↗

An Open Benchmark of One Million High-Fidelity Cislunar Trajectories

Cislunar space spans from geosynchronous altitudes to beyond the Moon and will underpin future exploration, science, and security operations. We describe and release an open dataset of one million numerically propagated cislunar trajectories generated with the open-source Space Situational Awareness Python package (SSAPy). The model includes high-degree Earth/Moon gravity, solar gravity, and Earth/Sun radiation pressure; other planetary gravities are omitted by design for computational efficiency. Initial conditions uniformly sample commonly used osculating-element ranges, and each trajectory is propagated for up to six years under a single, fixed start epoch. The dataset is intended as a reusable benchmark for method development (e.g., space domain awareness, navigation, and machine-learning pipelines), a reference library for statistical studies of orbit families, and a starting point for community-driven extensions (e.g., alternative epochs). We report empirically observed stability trends (e.g., a band near ~5 GEO and persistence of some co-orbital classes including L4/L5 librators) as dataset descriptors rather than new dynamical results. The chief contribution is the scale, fidelity, organization (CSV/HDF5 with full state time series and metadata), and open availability, which together lower the barrier to comparative and data-driven studies in the cislunar regime.

79 ASTRONOMY AND ASTROPHYSICS↗

Systems and methods for binary code analysis

Human-readable (HR) code may be derived from a binary. The HR code may be configured to have statistical properties suitable for machine-learned (ML) translation. The HR code may comprise source code, intermediate code, assembly code, or the like. A machine-learned translator may be configured to translate the HR code into labels comprising semantic information pertaining to respective functions of the binary, such as a function name, role, or the like. Execution of the binary may be blocked in response to translating the HR code to a label associated with malware, such as cryptocurrency mining malware or the like. Conversely, the binary may be permitted to proceed to execution in response to determining that the translation is free from labels indicative of malware.

Anderson, Matthew W.↗

Predicting battery capacity from impedance at varying temperature and state of charge using machine learning

Prediction of battery health from electrochemical impedance spectroscopy (EIS) data can enable rapid measurement of battery state in real-world applications without using additional sensors or time-consuming performance measurements. However, deconvoluting the effect of capacity, state of charge, and temperature on EIS response is complicated analytically. Here, various machine-learning models, such as linear, Gaussian process, random forest, and artificial neural network regression, are utilized to predict capacity from EIS using hundreds of capacity, direct current (DC) resistance, and EIS measurements recorded under varying conditions of health, temperature, and state of charge (SOC). Several feature extraction and selection methods from traditional electrochemical analysis and statistical modeling are explored using machine-learning pipelines. EIS data from just two frequencies can accurately predict capacity, and interrogation shows that the optimal set of frequencies is not usually intuitive. Best results are achieved with an ensemble model, which predicts battery capacity with a mean absolute error of 1.9% on data from unobserved cells.

25 ENERGY STORAGE↗

Simulated wildfire burned area over the CONUS during 2001-2020

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM). A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

Liu, Ye↗

Improving North American Wildfire Prediction by Integrating a Machine-Learning Fire Model in a Land Surface Model

Wildfires have shown increasing trends in both frequency and severity across the Contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth System Models (ESMs). Alternatively, machine learning (ML) based fire models, which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ML4Fire-XGB) that integrates a pretrained eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran-C-Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001-2020, the ML4Fire-XGB model outperforms process-based fire models in terms of spatial distribution and seasonal variations. Sensitivity analysis confirms that the ML4Fire-XGB well captures the responses of the burned area to rising temperatures. The ML4Fire-XGB model has proved to be a new tool for studying vegetation-fire interactions, and more importantly, enables seamless exploration of climate-fire feedback, working as an active component in E3SM.

54 ENVIRONMENTAL SCIENCES↗

Accelerating high-strain continuum-scale brittle fracture simulations with machine learning

Failure in brittle materials under dynamic loading conditions is a result of the propagation and coalescence of microcracks. Simulating this discrete crack evolution at the continuum level is computationally expensive or, in some cases, intractable, resulting in the need to make broad assumptions or neglect key physics. In this work, we have developed an approach using machine learning that overcomes the current inability to represent meso-scale physics at the macro-scale. Our approach leverages damage and stress data from a computationally expensive high-fidelity model that explicitly resolves microcrack behavior to build an inexpensive machine learning emulator. Once trained, the machine learning emulator is used to predict the evolution of crack length statistics, which then informs a continuum-scale constitutive model. This results in a significant speed-up of the workflow by four orders of magnitude. Both the machine learning emulator and the continuum-scale model are validated against the high-fidelity model and experimental data, respectively, showing excellent agreement. There are two key findings. The first is that we can reduce the dimensionality of the problem, establishing that the machine learning emulator only needs the length of the longest crack and one of the maximum stress components to capture the necessary physics. Another compelling finding is that the emulator can be trained in one experimental setting and transferred successfully to predict behavior in a different setting.

36 MATERIALS SCIENCE↗