A Model Ensemble Approach Enables Data-Driven Property Prediction for Chemically Deconstructable Thermosets in the Low-Data Regime
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Explore the source record for details and available documents.
Abstract We present an innovative approach called boosting Barlow Twins reduced order modeling (BBT‐ROM) to enhance the reliability of machine learning surrogate models for multiphase flow problems. BBT‐ROM builds upon Barlow Twins reduced order modeling that leverages self‐supervised learning to effectively handle linear and nonlinear manifolds by constructing well‐structured latent spaces of input parameters and output quantities. To address the challenge of high contrast data in multiphase flow problems due to injection wells and faults, we employ a boosting algorithm within BBT‐ROM. This algorithm sequentially trains a set of weak models (i.e., inaccurate models), improving prediction accuracy through ensemble learning. To evaluate the performance of BBT‐ROM, we conduct three three‐dimensional multiphase flow problems, including waterflooding and geologic carbon storage (GCS), with varying numbers of input parameter cases and model domain features. The results demonstrate that BBT‐ROM excels at predicting non‐wetting phase saturation (e.g., oil or saturation) and fluid pressure, with average relative errors ranging from 0.5% to 3%. Importantly, BBT‐ROM showcases robustness when faced with limited input parameter space during GCS testing.
Observational analysis shows that the Atlantic multidecadal variability (AMV) is associated with climate variability in the Northern Hemisphere through a zonal atmospheric teleconnection extending from the North Atlantic Ocean and propagating eastward around the Northern Hemisphere. Here, we studied the fidelity of model simulations in reproducing the observed summer AMV and the associated impacts on the mid-latitude climate by analysing simulations using the National Centre for Atmospheric Research Community Earth System Model Version 1 (CESM1), including CESM1 North Atlantic idealized and pacemaker simulations, CESM1 large ensemble twentieth century uninitialized simulations and large ensemble initialized CESM1 decadal predictions. To further compare the fidelity of CESM1, we also analysed large ensemble simulations from three other models. Our results suggest that the uninitialized large ensemble simulations from all models can produce an AMV time evolution and its regional climate impacts similar to the observations to certain degree. By initializing the observed oceanic condition in decadal prediction simulations, the simulated AMV and its regional impacts are closer to the observed ones than those in uninitialized ensemble simulations. In addition, the pacemaker simulations that nudged the time-evolving observed North Atlantic sea surface temperature anomalies produce spatiotemporal characteristics of the AMV and AMV climate impacts closer to the observed ones than the uninitialized simulations. We conclude that although coupled models can produce AMV and its regional impacts similar to observed, proper initialization and bias correction of the sea surface temperature spatial and temporal structure can improve this capability.
Prior to the launch of STS-119 NASA had completed a study of an issue in the flow control valve (FCV) in the Main Propulsion System of the Space Shuttle using an adaptive learning method known as Virtual Sensors. Virtual Sensors are a class of algorithms that estimate the value of a time series given other potentially nonlinearly correlated sensor readings. In the case presented here, the Virtual Sensors algorithm is based on an ensemble learning approach and takes sensor readings and control signals as input to estimate the pressure in a subsystem of the Main Propulsion System. Our results indicate that this method can detect faults in the FCV at the time when they occur. We use the standard deviation of the predictions of the ensemble as a measure of uncertainty in the estimate. This uncertainty estimate was crucial to understanding the nature and magnitude of transient characteristics during startup of the engine. This paper overviews the Virtual Sensors algorithm and discusses results on a comprehensive set of Shuttle missions and also discusses the architecture necessary for deploying such algorithms in a real-time, closed-loop system or a human-in-the-loop monitoring system. These results were presented at a Flight Readiness Review of the Space Shuttle in early 2009.
Numerical simulations in cosmology require trade-offs between volume, resolution and run-time that limit the volume of the Universe that can be simulated, leading to sample variance in predictions of ensemble-average quantities such as the power spectrum or correlation function(s). Sample variance is particularly acute at large scales, which is also where analytic techniques can be highly reliable. This provides an opportunity to combine analytic and numerical techniques in a principled way to improve the dynamic range and reliability of predictions for clustering statistics. In this paper we extend the technique of Zel'dovich control variates, previously demonstrated for 2-point functions in real space, to reduce the sample variance in measurements of 2-point statistics of biased tracers in redshift space. We demonstrate that with this technique, we can reduce the sample variance of these statistics down to their shot-noise limit out to k ~ 0.2 h Mpc -1 . This allows a better matching with perturbative models and improved predictions for the clustering of e.g. quasars, galaxies and neutral Hydrogen measured in spectroscopic redshift surveys at very modest computational expense. We discuss the implementation of ZCV, give some examples and provide forecasts for the efficacy of the method under various conditions.
Secondary forest regrowth shapes community succession and biogeochemistry for decades, including in the Upper Great Lakes region. Vegetation models encapsulate our understanding of forest function, and whether models can reproduce multi‐decadal succession patterns is an indication of our ability to predict forest responses to future change. We test the ability of a vegetation model to simulate C cycling and community composition during 100 years of forest regrowth following stand‐replacing disturbance, asking (a) Which processes and parameters are most important to accurately model Upper Midwest forest succession? (b) What is the relative importance of model structure versus parameter values to these predictions? We ran ensembles of the Ecosystem Demography model v2.2 with different representations of processes important to competition for light. We compared the magnitude of structural and parameter uncertainty and assessed which sub‐model–parameter combinations best reproduced observed C fluxes and community composition. On average, our simulations underestimated observed net primary productivity (NPP) and leaf area index (LAI) after 100 years and predicted complete dominance by a single plant functional type (PFT). Out of 4,000 simulations, only nine fell within the observed range of both NPP and LAI, but these predicted unrealistically complete dominance by either early hardwood or pine PFTs. A different set of seven simulations were ecologically plausible but under‐predicted observed NPP and LAI. Parameter uncertainty was large; NPP and LAI ranged from ~0% to >200% of their mean value, and any PFT could become dominant. The two parameters that contributed most to uncertainty in predicted NPP were plant–soil water conductance and growth respiration, both unobservable empirical coefficients. We conclude that (a) parameter uncertainty is more important than structural uncertainty, at least for ED‐2.2 in Upper Midwest forests and (b) simulating both productivity and plant community composition accurately without physically unrealistic parameters remains challenging for demographic vegetation models.
In this work, the application of the lagged average forecasting (LAF) technique to operational forecasts of the ECMWF is reported. The ECMWF data consist of two 100-day samples of 10-day forecasts of 500-mb geopotential height for winter 1980/81 and summer 1981. the LAF ensemble includes the latest operational forecast, and also forecast for the same verification time started one or more days earlier than the latest one. The focus is on the following two issues: (1) does ensemble averaging improve forecast skill and (2) is the dispersion of the ensemble useful in predicting forecast skill. The LAF technique was used to produce 3, 5, 7, 8, and 9 day forecasts of the 500-mb height field. The results show that the statistically filtered LAF is a marked improvment upon the operational forecast after 5 days. It is found that on a global scale, forecast skill is weakly correlated with the dispersion of the ensemble, as measured by the rms difference between the operational forecast and the statistically filtered LAF.
Climate predictions using coupled models in different time scales, from intraseasonal to decadal, are usually affected by initial shocks, drifts, and biases, which reduce the prediction skill. These arise from inconsistencies between different components of the coupled models and from the tendency of the model state to evolve from the prescribed initial conditions toward its own climatology over the course of the prediction. Aiming to provide tools and further insight into the mechanisms responsible for initial shocks, drifts, and biases, this paper presents a novel data set developed within the Long Range Forecast Transient Intercomparison Project, LRFTIP. This data set has been constructed by averaging hindcasts over available prediction years and ensemble members to form a hindcast climatology, that is a function of spatial variables and lead time, and thus results in a useful tool for characterizing and assessing the evolution of errors as well as the physical mechanisms responsible for them. A discussion on such errors at the different time scales is provided along with plausible ways forward in the field of climate predictions
Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.
The utilization of machine learning techniques has become commonplace in the analysis of optical emission spectra. These methods are often limited to variants of principal components analysis (PCA), partial-least squares (PLS), and artificial neural networks (ANNs). A plethora of other techniques exist and are well established in the world of data science, yet are seldom investigated for their use in spectroscopic problems. In this study, machine learning techniques were used to analyze optical emission spectra of laser-induced plasma from ceria pellets doped with silicon in order to predict silicon content. Additionally, a boosted regression ensemble model was created, and its predictive accuracy was compared to that of traditional PCA, PLS, and ANN regression models. Boosted regression tree ensembles yielded fits with R-squared (R2) values as high as 0.964 and mean-squared errors of prediction (MSEPs) as low as 0.074, providing the most accurate predictive model. Neural networks performed with slightly lower R2 values and higher MSEPs compared to the ensemble methods, thus indicating susceptibility to overfitting.
AI has the power to identify new pathways to extended predictability of the water cycle via its application to model ensembles together with accompanying observations. We discuss both scientific and technological aspects of this challenge, and address the parts of the MODEX approach involving “Model simulations, evaluation, analysis, and benchmarking” and “Identification of key knowledge gaps”. The widespread incorporation of AI into model analysis will significantly advance Earth system predictability and our predictive capabilities.
Foundation species have disproportionately large impacts on ecosystem structure and function. As a result, future changes to their distribution may be important determinants of ecosystem carbon (C) cycling in a warmer world. We assessed the role of a foundation tussock sedge (Eriophorum vaginatum) as a climatically vulnerable C stock using field data, a machine learning ecological niche model, and an ensemble of terrestrial biosphere models (TBMs). Field data indicated that tussock density has decreased by ~0.97 tussocks per m2 over the past ~38 years on Alaska's North Slope from ~1981 to 2019. This declining trend is concerning because tussocks are a large Arctic C stock, which enhances soil organic layer C stocks by 6.9% on average and represents 745 Tg C across our study area. By 2100, we project that changes in tussock density may decrease the tussock C stock by 41% in regions where tussocks are currently abundant (e.g. -0.8 tussocks per m2 and -85 Tg C on the North Slope) and may increase the tussock C stock by 46% in regions where tussocks are currently scarce (e.g. +0.9 tussocks per m2 and +81 Tg C on Victoria Island). These climate-induced changes to the tussock C stock were comparable to, but sometimes opposite in sign, to vegetation C stock changes predicted by an ensemble of TBMs. Our results illustrate the important role of tussocks as a foundation species in determining future Arctic C stocks and highlight the need for better representation of this species in TBMs..
Abstract Coupling between mesoscale models and large‐eddy simulation (LES) models is increasingly used to more realistically represent the wide range of scales of atmospheric motions affecting boundary layer winds and turbulence that need to be simulated accurately for applications such as wind energy. However, such mesoscale‐to‐microscale coupled modeling frameworks are potentially affected by a large number of uncertain closure parameters. Here, we investigate the sensitivity associated with six closure parameters related to a 1.5‐order subgrid‐scale turbulence closure for an ensemble of mesoscale‐coupled LES. The simulations are performed using the Weather Research and Forecasting model nested from horizontal resolutions of greater than a kilometer down to tens of meters. Closure parameters are varied to generate perturbed parameter ensembles for two case studies of highly sheared, convective boundary layers observed in the Columbia Basin of Oregon and Washington during the Second Wind Forecast Improvement Project. Machine learning algorithms are used to explore the sensitivity of LES predictions, considering the effects of the perturbed physical parameters alongside categorical factors such as the case study identity, measurement location, and LES resolution. For the conditions we examine, a single parameter, the eddy viscosity coefficient, is the dominant source of parametric sensitivity and its importance is comparable to the categorical factors for several of the simulation response variables we examine.
Data-driven approaches have the potential to make modeling complex, nonlinear physical phenomena significantly more computationally tractable. For example, computational modeling of fracture is a core challenge where machine learning techniques have the potential to provide a much needed speedup that would enable progress in areas such as multi-scale modeling and uncertainty quantification. Currently, phase field modeling (PFM) of fracture is one such approach that offers a convenient variational formulation to model crack nucleation, branching and propagation. To date, machine learning techniques have shown promise in approximating PFM simulations. While standard fracture benchmarks represent realistic scenarios frequently observed in practice, they typically do not provide sufficiently challenging tests for data-driven methods. Here, to address this gap, we introduce a challenging dataset based on PFM simulations designed to benchmark and advance ML methods for fracture modeling. This dataset includes three energy decomposition methods, two boundary conditions, and 1000 random initial crack configurations for a total of 6000 simulations. Each sample contains 100 time steps capturing the temporal evolution of the crack field. Alongside this dataset, we also implement and evaluate Physics Informed Neural Networks (PINN), Fourier Neural Operators (FNO), and UNet models as baselines, and explore the impact of ensembling strategies on prediction accuracy. With this combination of our dataset and baseline models drawn from the literature we aim to provide a standardized and challenging benchmark for evaluating machine learning approaches to solid mechanics. Our results highlight both the promise and limitations of popular current models, and demonstrate the utility of this dataset as a testbed for advancing machine learning in fracture mechanics research.
Accurate classification of molecular chemical motifs from experimental measurement is an important problem in molecular physics, chemistry, and biology. In this work, we present neural network ensemble classifiers for predicting the presence (or lack thereof) of 41 different chemical motifs on small molecules from simulated C, N, and O K-edge X-ray absorption near-edge structure (XANES) spectra. Our classifiers not only achieve class-balanced accuracies of more than 0.95 but also accurately quantify uncertainty. Here, we also show that including multiple XANES modalities improves predictions notably on average, demonstrating a “multimodal advantage” over any single modality. In addition to structure refinement, our approach can be generalized to broad applications with molecular design pipelines.
Supervised machine learning is the process of using past experience to predict the future. "Ensembles" are a machine-learning meta-method that can be applied to most machine learning algorithms. Ensembles generally greatly improve accuracy, reduce or remove most of the design issues presented by machine learning, and are admirably suited to parallel and distributed computation. The Avatar Tools codes are an implementation of ensembles specifically for decision trees. Some features that distinguish Avatar Tools from other "ensembles for decision trees" codes are: (1) Does the bookkeeping necessary for out of bag (OOB) validation. (2) Can use OOB validation to automatically determine optimal ensemble size. (3) Provides an MPI-based parallel implementation, for distributed operation. (4) Provides convenient tools for cross-validation, to assess the accuracy provided by a training set. SAND2020-3858 M Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.
Focal Area: This white paper responds to Focal area III by exploring data fusion, learning and explainable AI methods in characterizing hydrological extremes and interconnections. It also addresses Focal area II by using probabilistic AI and ensemble ML for predicting extremes and compound extremes. Science Challenge: A key question associated with the integrated water (or hydrological) cycle grand challenge in the Earth and Environmental Systems Sciences Division (EESSD) strategic plan, is how the frequency and intensity of hydrological events will change. Prediction of the tail behavior (extremes) of the hydrological cycle is especially challenging, because of their stochasticity and low probability. These extreme events and their compound impacts have significant societal and economic consequences. It is anticipated for the next-generation Earth System models (ESMs), that model predictability of the water cycle will improve with increased resolution (e.g., regionally refined E3SM), advanced software and computational architectures, and improved model physics based on the data from ARM measurements and high-fidelity models. However, the challenges for predictability of low-probability high-impact extreme events will unlikely be alleviated with conventional modeling and data-driven approaches, as ESMs are calibrated largely for capturing the high-frequency mean climate states. Recent AI and ML applications have shown great potential in quantifying well-defined climate extremes (e.g., supervised learning of tropical cyclones/atmospheric rivers by ClimateNet1) but few efforts are dedicated to compound events, extreme drivers and uncertainty estimation. We envision the opportunity to develop and apply ML and interpretable AI methods extended on the existing efforts, specifically, for: (1) identification of compound extremes, (2) diagnosing drivers of extremes, (3) bias correction in extreme predictions and (4) probabilistic modeling of extremes.