Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Time-Resolved Line Shapes of Single Quantum Emitters via Machine Learned Photon Correlations

Solid-state single-photon emitters (SPEs) are quantum light sources that combine atomlike optical properties with solid-state integration and fabrication capabilities. SPEs are hindered by spectral diffusion, where the emitter’s surrounding environment induces random energy fluctuations. Timescales of spectral diffusion span nanoseconds to minutes and require probing single emitters to remove ensemble averaging. Photon correlation Fourier spectroscopy (PCFS) can be used to measure time-resolved single emitter line shapes, but is hindered by poor signal-to-noise ratio in the measured correlation functions at early times due to low photon counts. Here, we develop a framework to simulate PCFS correlation functions directly from diffusing spectra that match well with experimental data for single colloidal quantum dots. We use these simulated datasets to train a deep ensemble autoencoder machine learning model that outputs accurate, noiseless, and probabilistic reconstructions of the noisy correlations. Using this model, we obtain reconstructed time-resolved single dot emission line shapes at timescales as low as 10 ns, which are otherwise completely obscured by noise. This enables PCFS to extract optical coherence times on the same timescales as Hong-Ou-Mandel two-photon interference, but with the advantage of providing spectral information in addition to estimates of photon indistinguishability. Further, our machine learning approach is broadly applicable to different photon correlation spectroscopy techniques and SPE systems, offering an enhanced tool for probing single emitter line shapes on previously inaccessible timescales.

74 ATOMIC AND MOLECULAR PHYSICS↗

Ensemble Forecasting/Assimilation/Emissions Estimation with WRF-Chem/DART: Accomplishments, Lessons Learned, and Future Plans

Over the past ten years, we have been conducting research on regional ensemble atmospheric composition forecasting/data assimilation/emissions estimation with WRF-Chem/DART. WRF- Chem/DART integrates the Weather Research and Forecasting model (WRF) with online chemistry (WRF-Chem) into the Data Assimilation Research Testbed (DART). DART is an ensemble data assimilation system based on the ensemble adjustment Kalman filter (EAKF) with adaptive inflation, localization (physical and state space), and an optional non-Gaussian formulation of the EAKF. DART includes assimilation of meteorological and limited chemical observations. WRF-Chem/DART extends DART to include assimilation of: MOPITT CO; IASI CO and O3; MODIS AOD; OMI O3, NO2, and SO2; TROPOMI CO, O3, NO2, and SO2, TES CO, CO2 (research mode), O3, NH3, and CH4 (research mode); CrIS CO, O3, NH3, CH4 (research mode), and PAN; SCIAMACHY NO2; GOME2a NO2; MLS O3 and HNO3; and proxy TEMPO O3, and NO2 satellite retrievals as raw retrievals or as ‘compact phase space retrievals’ (CPSRs) for profile retrievals. WRF-Chem/DART also assimilates in situ atmospheric composition measurements and uses the ‘state augmentation method’ for emissions estimation. In our presentation, we will provide an overview of WRF-Chem/DART: • applications and results; • lessons learned from: (i) independent versus joint assimilation; (ii) total/partial column versus profile retrieval assimilation; (iii) joint in situ and retrieval assimilation; (iv) assimilation at grid resolutions ranging from 100 km to 4 km; (v) dynamic emissions estimation; and • future work related to wildfire emissions estimation and intercomparison of CMAQ (online)/JEDI and CMAQ (online and offline)/DART.

WRF-Chem/DART↗

CoRE MOF DB: A curated experimental metal-organic framework database with machine-learned properties for integrated material-process screening

Here, we present an updated version of the Computation-Ready, Experimental (CoRE) Metal-Organic Framework (MOF) database, which includes a curated set of computation-ready MOF crystal structures designed for high-throughput computational materials discovery. Data collection and curation procedures were improved from the previous version to enable more frequent updates in the future. Machine-learning-predicted properties, such as stability metrics and heat capacities, are included in the dataset to streamline screening activities. An updated version of MOFid was developed to provide detailed information on metal nodes, organic linkers, and topologies of an MOF structure. DDEC6 partial atomic charges of MOFs were assigned based on a machine-learning model. Gibbs ensemble Monte Carlo simulations were used to classify the hydrophobicity of MOFs. The finalized dataset was subsequently used to perform integrated material-process screening for various carbon-capture conditions using high-fidelity temperature-swing adsorption (TSA) simulations. Our workflow identified multiple MOF candidates that are predicted to outperform CALF-20 for these applications.

CoRE MOF database↗

Identification of driver genes for critical forms of COVID-19 in a deeply phenotyped young patient cohort

The drivers of critical coronavirus disease 2019 (COVID-19) remain unknown. Given major confounding factors such as age and comorbidities, true mediators of this condition have remained elusive. We used a multi-omics analysis combined with artificial intelligence in a young patient cohort where major comorbidities were excluded at the onset. The cohort included 47 “critical” (in the intensive care unit under mechanical ventilation) and 25 “non-critical” (in a non-critical care ward) patients with COVID-19 and 22 healthy individuals. The analyses included whole-genome sequencing, whole-blood RNA sequencing, plasma and blood mononuclear cell proteomics, cytokine profiling, and high-throughput immunophenotyping. An ensemble of machine learning, deep learning, quantum annealing, and structural causal modeling were used. Patients with critical COVID-19 were characterized by exacerbated inflammation, perturbed lymphoid and myeloid compartments, increased coagulation, and viral cell biology. Among differentially expressed genes, we observed up-regulation of the metalloprotease ADAM9. This gene signature was validated in a second independent cohort of 81 critical and 73 recovered patients with COVID-19 and was further confirmed at the transcriptional and protein level and by proteolytic activity. Ex vivo ADAM9 inhibition decreased severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) uptake and replication in human lung epithelial cells. In conclusion, within a young, otherwise healthy, cohort of individuals with COVID-19, we provide the landscape of biological perturbations in vivo where a unique gene signature differentiated critical from non-critical patients. We further identified ADAM9 as a driver of disease severity and a candidate therapeutic target.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prognostic analysis of high-flow nasal cannula therapy and non-invasive ventilation in mild to moderate hypoxemia patients and construction of a machine learning model for 48-h intubation prediction—a retrospective analysis of the MIMIC database

Background This study aims to investigate the clinical outcome between high-flow nasal cannula (HFNC) and non-invasive ventilation (NIV) therapy in mild to moderate hypoxemic patients on the first ICU day and to develop a predictive model of 48-h intubation. Methods The study included adult patients from the MIMIC III and IV databases who first initiated HFNC or NIV therapy due to mild to moderate hypoxemia (100 < PaO2/FiO2 ≤ 300). The 48-h and 30-day intubation rates were compared using cross-sectional and survival analysis. Nine machine learning and six ensemble algorithms were deployed to construct the 48-h intubation predictive models, of which the optimal model was determined by its prediction accuracy. The top 10 risk and protective factors were identified using the Shapley interpretation algorithm. Result A total of 123,042 patients were screened, of which, 673 were from the MIMIC IV database for ventilation therapy comparison (HFNC n = 363, NIV n = 310) and 48-h intubation predictive model construction (training dataset n = 471, internal validation set n = 202) and 408 were from the MIMIC III database for external validation. The NIV group had a lower intubation rate (23.1% vs. 16.1%, p = 0.001), ICU 28-day mortality (18.5% vs. 11.6%, p = 0.014), and in-hospital mortality (19.6% vs. 11.9%, p = 0.007) compared to the HFNC group. Survival analysis showed that the total and 48-h intubation rates were not significantly different. The ensemble AdaBoost decision tree model (internal and external validation set AUROC 0.878, 0.726) had the best predictive accuracy performance. The model Shapley algorithm showed Sequential Organ Failure Assessment (SOFA), acute physiology scores (APSIII), the minimum and maximum lactate value as risk factors for early failure and age, the maximum PaCO 2 and PH value, Glasgow Coma Scale (GCS), the minimum PaO 2 /FiO 2 ratio, and PaO 2 value as protective factors. Conclusion NIV was associated with lower intubation rate and ICU 28-day and in-hospital mortality. Further survival analysis reinforced that the effect of NIV on the intubation rate might partly be attributed to the other impact factors. The ensemble AdaBoost decision tree model may assist clinicians in making clinical decisions, and early organ function support to improve patients’ SOFA, APSIII, GCS, PaCO 2 , PaO 2 , PH, PaO 2 /FiO 2 ratio, and lactate values can reduce the early failure rate and improve patient prognosis.

Fu, Wei↗

Evaluation of Tropospheric Water Vapor Simulations from the Atmospheric Model Intercomparison Project

Simulations of humidity from 28 general circulation models for the period 1979-88 from the Atmospheric Model Intercomparison Project are compared with observations from radiosondes over North America and the globe and with satellite microwave observations over the Pacific basin. The simulations of decadal mean values of precipitable water (W) integrated over each of these regions tend to be less moist than the real atmosphere in all three cases; the median model values are approximately 5% less than the observed values. The spread among the simulations is larger over regions of high terrain, which suggests that differences in methods of resolving topographic features are important. The mean elevation of the North American continent is substantially higher in the models than is observed, which may contribute to the overall dry bias of the models over that area. The authors do not find a clear association between the mean topography of a model and its mean W simulation, however, which suggests that the bias over land is not purely a matter of orography. The seasonal cycle of W is reasonably well simulated by the models, although over North America they have a tendency to become moister more quickly in the spring than is observed. The interannual component of the variability of W is not well captured by the models over North America. Globally, the simulated W values show a signal correlated with the Southern Oscillation index but the observations do not. This discrepancy may be related to deficiencies in the radiosonde network, which does not sample the tropical ocean regions well. Overall, the interannual variability of W, as well as its climatology and mean seasonal cycle, are better described by the median of the 28 simulations than by individual members of the ensemble. Tests to learn whether simulated precipitable water, evaporation, and precipitation values may be related to aspects of model formulation yield few clear signals, although the authors find, for example, a tendency for the few models that predict boundary layer depth to have large values of evaporation and precipitation. Controlled experiments, in which aspects of model architecture are systematically varied within individual models, may be necessary to elucidate whether and how model characteristics influence simulations.

Gaffen, Dian J.↗

Lessons from Climate Modeling on the Design and Use of Ensembles for Crop Modeling

Working with ensembles of crop models is a recent but important development in crop modeling which promises to lead to better uncertainty estimates for model projections and predictions, better predictions using the ensemble mean or median, and closer collaboration within the modeling community. There are numerous open questions about the best way to create and analyze such ensembles. Much can be learned from the field of climate modeling, given its much longer experience with ensembles. We draw on that experience to identify questions and make propositions that should help make ensemble modeling with crop models more rigorous and informative. The propositions include defining criteria for acceptance of models in a crop MME, exploring criteria for evaluating the degree of relatedness of models in a MME, studying the effect of number of models in the ensemble, development of a statistical model of model sampling, creation of a repository for MME results, studies of possible differential weighting of models in an ensemble, creation of single model ensembles based on sampling from the uncertainty distribution of parameter values or inputs specifically oriented toward uncertainty estimation, the creation of super ensembles that sample more than one source of uncertainty, the analysis of super ensemble results to obtain information on total uncertainty and the separate contributions of different sources of uncertainty and finally further investigation of the use of the multi-model mean or median as a predictor.

Model ensembles↗

Using Machine Learning to Generate a GISS ModelE Calibrated Physics Ensemble (CPE)

A neural network (NN) surrogate of the NASA GISS ModelE atmosphere (version E3) is trained on a perturbed parameter ensemble (PPE) spanning 45 physics parameters and 36 outputs. The NN is leveraged in a Markov Chain Monte Carlo (MCMC) Bayesian parameter inference framework to generate a second posterior constrained ensemble coined a “calibrated physics ensemble,” or CPE. The CPE members are characterized by diverse parameter combinations and are, by definition, close to top-of-atmosphere radiative balance, and must broadly agree with numerous hydrologic, energy cycle and radiative forcing metrics simultaneously. Global observations of numerous cloud, environment, and radiation properties (provided by global satellite products) are crucial for CPE generation. The inference framework explicitly accounts for discrepancies (or biases) in satellite products during CPE generation. We demonstrate that product discrepancies strongly impact calibration of important model parameter settings (e.g., convective plume entrainment rates; fall speed for cloud ice). Structural improvements new to E3 are retained across CPE members (e.g., stratocumulus simulation). Notably, the framework improved the simulation of shallow cumulus and Amazon rainfall while not degrading radiation fields, an upgrade that neither default parameters nor Latin Hypercube parameter searching achieved. Analyses of the initial PPE suggested several parameters were unimportant for output variation. However, many “unimportant” parameters were needed for CPE generation, a result that brings to the forefront how parameter importance should be determined in PPEs. From the CPE, two diverse 45-dimensional parameter configurations are retained to generate radiatively-balanced, auto-tuned atmospheres that were used in two E3 submissions to CMIP6.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Emulation of Spatial Deposition from a Multi-Physics Ensemble of Weather and Atmospheric Transport Models

In the event of an accidental or intentional hazardous material release in the atmosphere, researchers often run physics-based atmospheric transport and dispersion models to predict the extent and variation of the contaminant spread. These predictions are imperfect due to propagated uncertainty from atmospheric model physics (or parameterizations) and weather data initial conditions. Ensembles of simulations can be used to estimate uncertainty, but running large ensembles is often very time consuming and resource intensive, even using large supercomputers. In this paper, we present a machine-learning-based method which can be used to quickly emulate spatial deposition patterns from a multi-physics ensemble of dispersion simulations. We use a hybrid linear and logistic regression method that can predict deposition in more than 100,000 grid cells with as few as fifty training examples. Logistic regression provides probabilistic predictions of the presence or absence of hazardous materials, while linear regression predicts the quantity of hazardous materials. The coefficients of the linear regressions also open avenues of exploration regarding interpretability—the presented model can be used to find which physics schemes are most important over different spatial areas. A single regression prediction is on the order of 10,000 times faster than running a weather and dispersion simulation. However, considering the number of weather and dispersion simulations needed to train the regressions, the speed-up achieved when considering the whole ensemble is about 24 times. Ultimately, this work will allow atmospheric researchers to produce potential contamination scenarios with uncertainty estimates faster than previously possible, aiding public servants and first responders.

97 MATHEMATICS AND COMPUTING↗

Quantifying uncertainty for deep learning based forecasting and flow-reconstruction using neural architecture search ensembles

Classical problems in computational physics such as data-driven forecasting and signal reconstruction from sparse sensors have recently seen an explosion in deep neural network (DNN) based algorithmic approaches. However, most DNN models do not provide uncertainty estimates, which are crucial for establishing the trustworthiness of these techniques in downstream decision making tasks and scenarios. In recent years, ensemble-based methods have achieved significant success for the uncertainty quantification in DNNs on a number of benchmark problems. However, their performance on real-world applications remains under-explored. In this work, we present an automated approach to DNN discovery and demonstrate how this may also be utilized for ensemble-based uncertainty quantification. Specifically, we propose the use of a scalable neural and hyperparameter architecture search for discovering an ensemble of DNN models for complex dynamical systems. We highlight how the proposed method not only discovers high-performing neural network ensembles for our tasks, but also quantifies uncertainty seamlessly. This is achieved by using genetic algorithms and Bayesian optimization for sampling the search space of neural network architectures and hyperparameters. Subsequently, a model selection approach is used to identify candidate models for an ensemble set construction. Afterwards, a variance decomposition approach is used to estimate the uncertainty of the predictions from the ensemble. We demonstrate the feasibility of this framework for two tasks — forecasting from historical data and flow reconstruction from sparse sensors for the sea-surface temperature. In conclusion, we demonstrate superior performance from the ensemble in contrast with individual high-performing models and other benchmarks.

Deep ensembles↗

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Uncertainty based Online Ensemble on Non-Stationary Data for Fusion Science

Machine Learning (ML) is poised to play a pivotal role in the development and operation of next-generation fusion devices. Fusion data shows non-stationary behavior due to drifts in the data. The drifts can arise from both experimental evolution and machine wear-and-tear. ML models assume stationary distribution and fail to maintain performance when encountered with non-stationary data streams.Online learning can be used to continuously adapt the models with new data as it is acquired. However, traditional online learning can suffer from short-term performance degradation, as ground truth are not available before making the prediction. To address this challenge, we propose uncertainty aware ensemble approach for online learning. We use Deep Gaussian Process Approximation (DGPA) technique for calibrated uncertainty estimation and use the uncertainty values to guide a meta-algorithm that produces predictions based on ensemble of learners. Moreover, DGPA also provides uncertainty estimation along with the predictions for decision makers. This paper demonstrates that the proposed method outperforms traditional online learning approach, and a naive ensemble without uncertainty guidance by about 7% and 6%, respectively, on B-coil deflection prediction at DIII-D Fusion Facility.

Rajput, Kishansingh [Thomas Jefferson National Acc↗

Scalable statistical inference of photometric redshift via data subsampling

Handling big data has largely been a major bottleneck in traditional statistical models. Consequently, when accurate point prediction is the primary target, machine learning models are often preferred over their statistical counterparts for bigger problems. But full probabilistic statistical models often outperform other models in quantifying uncertainties associated with model predictions. We develop a data-driven statistical modeling framework that combines the uncertainties from an ensemble of statistical models learned on smaller subsets of data carefully chosen to account for imbalances in the input space. We demonstrate this method on a photometric redshift estimation problem in cosmology, which seeks to infer a distribution of the redshift—the stretching effect in observing the light of far-away galaxies—given multivariate color information observed for an object in the sky. Our proposed method performs balanced partitioning, graph-based data subsampling across the partitions, and training of an ensemble of Gaussian process models.

data subsampling↗

Using Federated Learning to Overcome Data Gravity in Space

Humans intend to take longer missions to outer space. Understanding the impact that space has on human health is paramount to the success of these missions. Controlled experiments with model organisms are run to infer the impact of space conditions on human health, but the data these experiments generate are too large to transfer to Earth for building models. The same is true for space-relevant data generated on Earth. Ideally, these datasets should be combined to improve statistical power and model accuracy without having to transfer data. Federated learning is such a method which trains an algorithm across decentralized computing systems, each of which has their own local copy of training and testing data. In this research, made possible by NASA@Work, the AI for Life in Space group at NASA demonstrates the use of federated learning to train an ensemble of causality inference models on a combination of data residing on the International Space Station (ISS) and in the cloud. Our work leverages CRISP, a causal inference platform developed during the 2020 Frontier Development Lab’s “Astronaut Health Challenge.” We also leverage the OpenFL federated learning library which was collaboratively developed at Intel and UPenn. We used publicly available data from the NASA Ames Life Sciences Data Archive to identify features in ionizing radiation experiments as causal of changes in cardiac blood velocity. This research demonstrates, for the first time, the possibility of running machine learning algorithms on datasets separated by astronomical distances. In this experiment, all the data were generated in terra, half of which were transferred to the ISS and analyzed on the Spaceborne Computer. In the future, our research will leverage federated learning on data generated in situ on the ISS with data generated terrestrially to predict the impact of spaceflight on mammalian female reproductive capacity.

James Casaletto↗

A benchmark dataset for Hydrogen Combustion

The generation of reference data for deep learning models is challenging for reactive systems, and more so for combustion reactions due to the extreme conditions that create radical species and alternative spin states during the combustion process. Here, we extend intrinsic reaction coordinate (IRC) calculations with ab initio MD simulations and normal mode displacement calculations to more extensively cover the potential energy surface for 19 reaction channels for hydrogen combustion. A total of ~290,000 potential energies and ~1,270,000 nuclear force vectors are evaluated with a high quality range-separated hybrid density functional, ωB97X-V, to construct the reference data set, including transition state ensembles, for the deep learning models to study hydrogen combustion reaction.

08 HYDROGEN↗