Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Implementation of stacked ensemble machine learning for the detection of surrogate plutonium contamination in soil via LIBS

Supervised machine learning methods have demonstrated increased utility for the quantification of lanthanide and actinide elements in atomic spectroscopy applications. This study implements laser-induced breakdown spectroscopy (LIBS) for the identification of plutonium surrogate material (CeO 2 ) in soil matrices by training supervised machine learning methods on the recorded spectral data. A bagged ensemble using Random Forest yields the highest sensitivity predictions with a detection limit of 0.015 wt.% CeO 2 . However, high precision in Ce content prediction required the use of a stacked ensemble regression, which provided the superlative Ce quantification model with an error of 0.107% and a detection limit of 0.022 wt.%. Furthermore, the high performance of the stacked ensemble demonstrates its potential to enhance the accuracy and sensitivity of nuclear contaminant detection using field-deployable spectroscopic analyzers in real-world scenarios.

47 OTHER INSTRUMENTATION↗

Statistical upscaling of ecosystem CO 2 fluxes across the terrestrial tundra and boreal domain: Regional patterns and uncertainties

Abstract The regional variability in tundra and boreal carbon dioxide (CO 2 ) fluxes can be high, complicating efforts to quantify sink‐source patterns across the entire region. Statistical models are increasingly used to predict (i.e., upscale) CO 2 fluxes across large spatial domains, but the reliability of different modeling techniques, each with different specifications and assumptions, has not been assessed in detail. Here, we compile eddy covariance and chamber measurements of annual and growing season CO 2 fluxes of gross primary productivity (GPP), ecosystem respiration (ER), and net ecosystem exchange (NEE) during 1990–2015 from 148 terrestrial high‐latitude (i.e., tundra and boreal) sites to analyze the spatial patterns and drivers of CO 2 fluxes and test the accuracy and uncertainty of different statistical models. CO 2 fluxes were upscaled at relatively high spatial resolution (1 km 2 ) across the high‐latitude region using five commonly used statistical models and their ensemble, that is, the median of all five models, using climatic, vegetation, and soil predictors. We found the performance of machine learning and ensemble predictions to outperform traditional regression methods. We also found the predictive performance of NEE‐focused models to be low, relative to models predicting GPP and ER. Our data compilation and ensemble predictions showed that CO 2 sink strength was larger in the boreal biome (observed and predicted average annual NEE −46 and −29 g C m −2 yr −1 , respectively) compared to tundra (average annual NEE +10 and −2 g C m −2 yr −1 ). This pattern was associated with large spatial variability, reflecting local heterogeneity in soil organic carbon stocks, climate, and vegetation productivity. The terrestrial ecosystem CO 2 budget, estimated using the annual NEE ensemble prediction, suggests the high‐latitude region was on average an annual CO 2 sink during 1990–2015, although uncertainty remains high.

Virkkala, Anna‐Maria↗

A minimally invasive, efficient method for propagation of full-field uncertainty in solid dynamics

In this work, we present a minimally invasive method for forward propagation of material property uncertainty to full-field quantities of interest in solid dynamics. Full-field uncertainty quantification enables the design of complex systems where quantities of interest, such as failure points, are not known a priori. The method, motivated by the well-known probability density function (PDF) propagation method of turbulence modeling, uses an ensemble of solutions to provide the joint PDF of desired quantities at every point in the domain. A small subset of the ensemble is computed exactly, and the remainder of the samples are computed with approximation of the evolution equations based on those exact solutions. Although the proposed method has commonalities with traditional interpolatory stochastic collocation methods applied directly to quantities of interest, it is distinct and exploits the parameter dependence and smoothness of the driving term of the evolution equations. The implementation is model independent, storage and communication efficient, and straightforward. We demonstrate its efficiency, accuracy, scaling with dimension of the parameter space, and convergence in distribution with two problems: a quasi-one-dimensional bar impact, and a two material notched plate impact. For the bar impact problem, we provide an analytical solution to PDF of the solution fields for method validation. With the notched plate problem, we also demonstrate good parallel efficiency and scaling of the method.

42 ENGINEERING↗

The WRF-Solar Ensemble Prediction System to Provide Solar Irradiance Probabilistic Forecasts

In this study, we introduce the recently developed WRF-solar ensemble prediction system and a calibration method. The performances of forecast models are evaluated using the National Solar Radiation Database observational analysis for day-ahead solar irradiance predictions. The results demonstrate that the ensemble forecast improves the quality of the forecasts by considering the uncertainty of each ensemble member. The analog ensemble calibration contributed to the reduction of positive bias and an overall improvement in the probabilistic attributes, such as reliability and statistical consistency.

14 SOLAR ENERGY↗

Reducing the complexity of finite-temperature auxiliary-field quantum Monte Carlo

The auxiliary-field quantum Monte Carlo (AFMC) method is a powerful and widely used technique for ground-state and finite-temperature simulations of quantum many-body systems. Here we introduce several algorithmic improvements for finite-temperature AFMC calculations of dilute fermionic systems that reduce the computational complexity of most parts of the algorithm. This is principally achieved by reducing the number of single-particle states that contribute at each configuration of the auxiliary fields to a number that is of the order of the number of fermions. Our methods are applicable for both the canonical and grand-canonical ensembles. We demonstrate the reduced computational complexity of the methods for the homogeneous unitary Fermi gas.

67.85.Lm↗

The role of SAXS and molecular simulations in 3D structure elucidation of a DNA aptamer against lung cancer

Aptamers are short, single-stranded DNA or RNA oligonucleotide molecules that function as synthetic analogs of antibodies and bind to a target molecule with high specificity. Aptamer affinity entirely depends on its tertiary structure and charge distribution. Therefore, length and structure optimization are essential for increasing aptamer specificity and affinity. Here, we present a general optimization procedure for finding the most populated atomistic structures of DNA aptamers. Based on the existed aptamer LC-18 for lung adenocarcinoma, a new truncated LC-18 (LC-18t) aptamer LC-18t was developed. A three-dimensional (3D) shape of LC-18t was reported based on small-angle X-ray scattering (SAXS) experiments and molecular modeling by fragment molecular orbital or molecular dynamic methods. Molecular simulations revealed an ensemble of possible aptamer conformations in solution that were in close agreement with measured SAXS data. The aptamer LC-18t had stronger binding to cancerous cells in lung tumor tissues and shared the binding site with the original larger aptamer. The suggested approach reveals 3D shapes of aptamers and helps in designing better affinity probes.

59 BASIC BIOLOGICAL SCIENCES↗

Pushing the frontiers in climate modelling and analysis with machine learning

Climate modelling and analysis are facing new demands to enhance projections and climate information. Here, in this study, we argue that now is the time to push the frontiers of machine learning beyond state-of-the-art approaches, not only by developing machine-learning-based Earth system models with greater fidelity, but also by providing new capabilities through emulators for extreme event projections with large ensembles, enhanced detection and attribution methods for extreme events, and advanced climate model analysis and benchmarking. Utilizing this potential requires key machine learning challenges to be addressed, in particular generalization, uncertainty quantification, explainable artificial intelligence and causality. This interdisciplinary effort requires bringing together machine learning and climate scientists, while also leveraging the private sector, to accelerate progress towards actionable climate science.

54 ENVIRONMENTAL SCIENCES↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Localization of infrasonic sources via Bayesian back projection

SUMMARY A Bayesian framework is investigated for event-specific localization of infrasonic sources using back projection ray tracing. Direction-of-arrival information from array-based detection analysis is used to initialize a back projection ray path originating from the detecting array location and quantifying propagation characteristics from hypothetical source locations. The Fisher statistic, computed from the array’s beam coherence, is mapped into uncertainty in the launch angles of the ray path. Auxiliary parameters previously introduced for solving the Transport equation to compute geometric spreading along ray paths are used to map uncertainty in the ray launch angles into spatial and temporal uncertainties in the ray path. An atmospheric ensemble approach is applied to account for atmospheric uncertainty, and the relation between uncertainties in the atmospheric state and confidence in estimated localization are evaluated using several ensembles with specified variances. The method is evaluated using a synthetic event in the western United States constructed via forward propagation simulations as well as a single-station, multi-arrival detection from a surface explosion in the western United States. Localization results using this event-specific approach are more accurate and exhibit improved precision than existing Bayesian localization methods that leverage generalized, pre-computed propagation statistics.

58 GEOSCIENCES↗

Designs from Local Random Quantum Circuits with SU ( d ) Symmetry

The generation of k -designs (pseudorandom distributions that emulate the Haar measure up to k moments) with local quantum circuit ensembles is a problem of fundamental importance in quantum information and physics. Despite the extensive understanding of this problem for ordinary random circuits, the crucial situations in which symmetries or conservation laws are in play are known to pose fundamental challenges and remain little understood. Here, we construct explicit local unitary ensembles that can achieve high-order unitary k -designs under transversal continuous symmetry, in the particularly important SU ( d ) case. Specifically, we define the convolutional quantum alternating (CQA) group generated by 4-local SU ( d ) -symmetric Hamiltonians as well as associated 4-local SU ( d ) -symmetric random unitary circuit ensembles and prove that they form and converge to SU ( d ) -symmetric k -designs, respectively, for all k < n ( n − 3 ) / 2 , with n being the number of qudits. A key technique that we employ to obtain the results is the Okounkov-Vershik approach to S n representation theory. To study the convergence time of the CQA ensemble, we develop a numerical method using the Young orthogonal form and the S n branching rule. We provide strong evidence for a subconstant spectral gap and certain convergence time scales of various important circuit architectures, which contrast with the symmetry-free case. We also provide comprehensive explanations of the difficulties and limitations in rigorously analyzing the convergence time using methods that have been effective for cases without symmetries, including Knabe’s local gap threshold and Nachtergaele’s martingale methods. This suggests that a novel approach is likely necessary for understanding the convergence time of SU ( d ) -symmetric local random circuits. Published by the American Physical Society 2024

Li, Zimu (ORCID:0000000314736492)↗

Decomposing Cloud Radiative Feedbacks by Cloud-Top Phase

Changes in cloud scattering properties and emissivity that arise from atmospheric warming cause substantial radiative feedbacks in model projections of anthropogenic climate change, and the relative importance of the underlying mechanisms is poorly understood. One leading hypothesis is that ice-to-liquid conversions cause clouds to optically thicken, producing a major negative feedback. We test this hypothesis by developing a method to decompose cloud radiative feedbacks by cloud-top phase. The method is applied to an ensemble of six state-of-the-art global climate models run with prescribed sea surface temperature. In these simulations, the global mean of the net cloud scattering and emissivity feedback from cloud-phase conversions ranges from −0.17 to −0.01 W m −2 K −1 , while the overall net cloud feedback ranges from 0.02 to 0.91 W m −2 K −1 . The multimodel mean of the cloud scattering and emissivity feedback from cloud-phase conversions is approximately 19% of the magnitude of the multimodel mean of the overall cloud feedback (−0.10 vs 0.52 W m −2 K −1 ). These results indicate that cloud-phase conversions cause a robust negative feedback by changing cloud scattering and emissivity, but this mechanism makes a modest contribution to the overall cloud feedback at the global scale.

Climate change↗

Comprehensive uncertainty quantification (UQ) for full engineering models by solving probability density function (PDF) equation

This report details a new method for propagating parameter uncertainty (forward uncertainty quantification) in partial differential equations (PDE) based computational mechanics applications. The method provides full-field quantities of interest by solving for the joint probability density function (PDF) equations which are implied by the PDEs with uncertain parameters. Full-field uncertainty quantification enables the design of complex systems where quantities of interest, such as failure points, are not known apriori. The method, motivated by the well-known probability density function (PDF) propagation method of turbulence modeling, uses an ensemble of solutions to provide the joint PDF of desired quantities at every point in the domain. A small subset of the ensemble is computed exactly, and the remainder of the samples are computed with approximation of the driving (dynamics) term of the PDEs based on those exact solutions. Although the proposed method has commonalities with traditional interpolatory stochastic collocation methods applied directly to quantities of interest, it is distinct and exploits the parameter dependence and smoothness of the dynamics term of the governing PDEs. The efficacy of the method is demonstrated by applying it to two target problems: solid mechanics explicit dynamics with uncertain material model parameters, and reacting hypersonic fluid mechanics with uncertain chemical kinetic rate parameters. A minimally invasive implementation of the method for representative codes SPARC (reacting hypersonics) and NimbleSM (finite- element solid mechanics) and associated software details are described. For solid mechanics demonstration problems the method shows order of magnitudes improvement in accuracy over traditional stochastic collocation. For the reacting hypersonics problem, the method is implemented as a streamline integration and results show very good accuracy for the approximate sample solutions of re-entry flow past the Apollo capsule geometry at Mach 30.

42 ENGINEERING↗

How well do one-electron self-interaction-correction methods perform for systems with fractional electrons?

Recently developed locally scaled self-interaction correction (LSIC) is a one-electron SIC method that, when used with a ratio of kinetic energy densities (z σ ) as iso-orbital indicator, performs remarkably well for both thermochemical properties as well as for barrier heights overcoming the paradoxical behavior of the well-known Perdew–Zunger self-interaction correction (PZSIC) method. In this work, we examine how well the LSIC method performs for the delocalization error. Our results show that both LSIC and PZSIC methods correctly describe the dissociation of $H$$^{+}_{2}$ and $H$$^{+}_{2}$ but LSIC is overall more accurate than the PZSIC method. Likewise, in the case of the vertical ionization energy of an ensemble of isolated He atoms, the LSIC and PZSIC methods do not exhibit delocalization errors. For the fractional charges, both LSIC and PZSIC significantly reduce the deviation from linearity in the energy vs number of electrons curve, with PZSIC performing superior for C, Ne, and Ar atoms while for Kr they perform similarly. The LSIC performs well at the endpoints (integer occupations) while substantially reducing the deviation. The dissociation of LiF shows both LSIC and PZSIC dissociate into neutral Li and F but only LSIC exhibits charge transfer from Li + to F – at the expected distance from the experimental data and accurate ab initio data. Overall, both the PZSIC and LSIC methods reduce the delocalization errors substantially.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantification of regional net CO 2 flux errors in the Orbiting Carbon Observatory-2 (OCO-2) v10 model intercomparison project (MIP) ensemble using airborne measurements

Inverse model intercomparison projects (MIPs) provide a chance to assess the uncertainties in inversion estimates arising from various sources. However, accurately quantifying ensemble CO 2 flux errors remains challenging and often relies on the ensemble spread. This study proposes a method for quantifying the errors in regional net surface–atmosphere CO 2 flux estimates from models taken from the Orbiting Carbon Observatory-2 (OCO-2) v10 MIP by using independent airborne CO 2 measurements for the period 2015–2017. We first calculate the root mean square error (RMSE) between the ensemble mean of posterior CO 2 concentrations and airborne observations and then isolate the CO 2 concentration errors caused solely by the ensemble mean of posterior net fluxes by subtracting the observation, representation, and transport errors from seven regions. Our analysis reveals that the flux errors projected onto CO 2 space account for 55 %–85 % of the regional average RMSE over the 3 years, ranging from 0.88 to 1.91 ppm. In five regions, the error estimates based on observations exceed those computed from the ensemble spread of posterior fluxes by a factor of 1.33–1.93, implying an underestimation of the actual flux errors, while their magnitudes are comparable in two regions. The adjoint sensitivity analysis identifies that the underestimation of flux errors is prominent where the magnitudes of fossil fuel emissions exceed those of terrestrial-biosphere fluxes by a factor of 3–31 over the 3 years. This suggests the presence of systematic biases in the inversion estimates associated with errors in the prescribed fossil fuel emissions common to all models. Our study emphasizes the value of airborne measurements for quantifying regional errors in ensemble net CO 2 flux estimates.

54 ENVIRONMENTAL SCIENCES↗

Bootstrap-determined p values in lattice QCD

We present a general method to determine the probability that stochastic Monte Carlo data, in particular those generated in a lattice QCD calculation, would have been obtained were that data drawn from the distribution predicted by a given theoretical hypothesis. Such a probability, or p -value, is often used as an important heuristic measure of the validity of that hypothesis. The proposed method offers the benefit that it remains usable in cases where the standard Hotelling T 2 methods based on the conventional χ 2 statistic do not apply, such as for uncorrelated fits. Specifically, we analyze q 2 , defined as the correlated χ 2 statistic obtained using an arbitrary covariance matrix estimator, and show how to use the bootstrap as a data-driven method to determine the expected distribution of q 2 for a given hypothesis with minimal assumptions. This distribution can then be used to determine the p -value for a fit to the data. We also describe a bootstrap approach for quantifying the impact upon this p -value of estimating population parameters from a single ensemble of N samples. The overall method is accurate up to a 1 / N bias which we do not attempt to quantify. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Clarifying trust of materials property predictions using neural networks with distribution-specific uncertainty quantification

It is critical that machine learning (ML) model predictions be trustworthy for high-throughput catalyst discovery approaches. Uncertainty quantification (UQ) methods allow estimation of the trustworthiness of an ML model, but these methods have not been well explored in the field of heterogeneous catalysis. Herein, we investigate different UQ methods applied to a crystal graph convolutional neural network to predict adsorption energies of molecules on alloys from the Open Catalyst 2020 dataset, the largest existing heterogeneous catalyst dataset. We apply three UQ methods to the adsorption energy predictions, namely k-fold ensembling, Monte Carlo dropout, and evidential regression. The effectiveness of each UQ method is assessed based on accuracy, sharpness, dispersion, calibration, and tightness. Evidential regression is demonstrated to be a powerful approach for rapidly obtaining tunable, competitively trustworthy UQ estimates for heterogeneous catalysis applications when using neural networks. Recalibration of model uncertainties is shown to be essential in practical screening applications of catalysts using uncertainties.

36 MATERIALS SCIENCE↗

Discovering the Multisectoral Impacts of Global Energy Sector Outcomes Through Multiple Ensemble Aggregation Measures

Understanding complex human-Earth system interactions often involves analyzing large scenario ensembles that encompass a wide range of plausible futures. These ensembles often require aggregation to summarize information based on specific criteria or conditions. However, previous research using global change scenario ensembles has largely overlooked how the choice of aggregation method influences the interpretation of results. To address this gap, we leverage a large ensemble data set designed to capture broad energy system dynamics generated using the Global Change Analysis Model. We first explore how energy-related uncertainties are propagated to both global and regional water-energy-food sectors. We then conduct a rank correlation analysis across seven ensemble aggregation measures and demonstrate the need to consider multiple measures in global change scenarios. Our results suggest that global water and food sector outcomes in the 21st century vary widely depending on different scenario assumptions. The global energy productivity is projected to improve by the end of the century across all scenarios. Moreover, regions facing water scarcity challenges in 2100 do not always overlap with those facing extreme energy and food sector outcomes. Although rank correlations across seven aggregation measures are relatively stable across sectors, we identify cases where relying on a single measure leads to losing critical information in the full ensemble. Reliance on a single aggregation measure can distort the interpretation of global change scenario outcomes. Instead, adopting multiple ensemble aggregation measures provides a more holistic understanding of global change scenario ensembles.

Kim, Gijoo↗