Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Colloidal quantum dots for optoelectronics

Colloidal quantum dots (QDs) are semiconductor nanocrystals that have unique size-tunable optoelectronic properties and are suitable for wet processing. QD research aims to answer fundamental questions about the chemical and physical properties of nanoscale materials and use these tools for technological applications ranging from bio-imaging to quantum optics. At the core of this field is a set of synthetic, processing and analytical methods designed to produce QDs in uniform ensembles that meet the highest performance standards. Here, this Primer reviews QD fabrication methods with a focus on the applications of QDs in printed optoelectronics and quantum optics. After outlining the current state-of-the-art QD syntheses, the experimental and computational analysis of QDs is discussed. These topics are then connected to the methodologies, processes and concepts required for developing QD-based photodetectors, light-emitting devices and quantum optics applications. Special attention is paid to challenges in reproducibility and current limitations of the field, such as the need to balance non-restricted material composition with high performing technology while achieving long-term stability in QD devices under operating conditions. Finally, the ongoing advancement in QD synthesis, precise atomic-level analysis and computational methodologies are highlighted as key drivers towards rational QD design, particularly in understanding how structural changes under loading impact QD properties.

optical materials↗

Boosting efficiency and reducing graph reliance: Basis adaptation integration in Bayesian multi-fidelity networks

The computational cost of high-fidelity numerical models makes outer-loop analysis, which requires repeated interrogation of the model such as uncertainty quantification, computationally demanding. Multi-fidelity methods, which construct a surrogate model using data from an ensemble of models of varying cost and accuracy, can substantially reduce the cost of outer-loop analysis. However, these methods can be difficult to apply when the model ensemble does not admit a clear hierarchy a priori and the correlations between models are low. Consequently, in this paper, we present a multi-fidelity method that leverages dimension reduction to enhance the correlation between models, thereby reducing the amount of data needed to train a surrogate from an unordered ensemble of models. Our method utilizes basis adaptation to build low-dimensional polynomial chaos expansions of each model and employs Multi-fidelity Networks to encode the relationships among models. We show that the resulting method exhibit two notable advantages over its counterpart: (1) enhanced accuracy (both reduced bias and variance); and (2) reduced dependency on the graph structure encoding relationships among models. We demonstrate the approach on an analytical test problem and a challenging finite element model for a spent nuclear fuel. Our method produces a surrogate model that is significantly more accurate than either a single-fidelity surrogate or a multi-fidelity surrogate constructed without basis adaptation.

42 ENGINEERING↗

Solving high-dimensional inverse problems using amortized likelihood-free inference with noisy and incomplete data

Here, we present a likelihood-free probabilistic inversion method based on normalizing flows for high-dimensional inverse problems. The proposed method is composed of two complementary networks: a summary network for data compression and an inference network for parameter estimation. The summary network encodes raw observations into a fixed-size vector of summary features, while the inference network generates samples of the approximate posterior distribution of the model parameters based on these summary features. The posterior samples are produced in a deep generative fashion by sampling from a latent Gaussian distribution and passing these samples through an invertible transformation. We construct this invertible transformation by sequentially alternating conditional invertible neural network and conditional neural spline flow layers. The summary and inference networks are trained simultaneously. We apply the proposed method to an inversion problem in groundwater hydrology to estimate the posterior distribution of the log-conductivity field conditioned on spatially sparse time-series observations of the system’s hydraulic head responses. The conductivity field is represented with 706 degrees of freedom in the considered problem. Comparison with the likelihood-based iterative ensemble smoother PEST-IES method demonstrates that the proposed method accurately estimates the parameter posterior distribution and the observations’ predictive posterior distribution at a fraction of the inference time of PEST-IES.

conditional invertible neural network↗

Enhancing weak lensing redshift distribution characterization by optimizing the Dark Energy Survey Self-Organizing Map Photo-z method

Characterization of the redshift distribution of ensembles of galaxies is pivotal for large scale structure cosmological studies. In this work, we focus on improving the Self-Organizing Map (SOM) methodology for photometric redshift estimation (SOMPZ), specifically in anticipation of the Dark Energy Survey Year 6 (DES Y6) data. This data set, featuring deeper and fainter galaxies than DES Year 3 (DES Y3), demands adapted techniques to ensure accurate recovery of the underlying redshift distribution. We investigate three strategies for enhancing the existing SOM-based approach used in DES Y3: 1) Replacing the Y3 SOM algorithm with one tailored for redshift estimation challenges; 2) Incorporating $\textit{g}$-band flux information to refine redshift estimates (i.e. using $\textit{griz}$ fluxes as opposed to only $\textit{riz}$); 3) Augmenting redshift data for galaxies where available. These methods are applied to DES Y3 data, and results are compared to the Y3 fiducial ones. Our analysis indicates significant improvements with the first two strategies, notably reducing the overlap between redshift bins. By combining strategies 1 and 2, we have successfully managed to reduce redshift bin overlap in DES Y3 by up to 66$\%$. Conversely, the third strategy, involving the addition of redshift data for selected galaxies as an additional feature in the method, yields inferior results and is abandoned. Our findings contribute to the advancement of weak lensing redshift characterization and lay the groundwork for better redshift characterization in DES Year 6 and future stage IV surveys, like the Rubin Observatory.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Implementation of stacked ensemble machine learning for the detection of surrogate plutonium contamination in soil via LIBS

Supervised machine learning methods have demonstrated increased utility for the quantification of lanthanide and actinide elements in atomic spectroscopy applications. This study implements laser-induced breakdown spectroscopy (LIBS) for the identification of plutonium surrogate material (CeO 2 ) in soil matrices by training supervised machine learning methods on the recorded spectral data. A bagged ensemble using Random Forest yields the highest sensitivity predictions with a detection limit of 0.015 wt.% CeO 2 . However, high precision in Ce content prediction required the use of a stacked ensemble regression, which provided the superlative Ce quantification model with an error of 0.107% and a detection limit of 0.022 wt.%. Furthermore, the high performance of the stacked ensemble demonstrates its potential to enhance the accuracy and sensitivity of nuclear contaminant detection using field-deployable spectroscopic analyzers in real-world scenarios.

47 OTHER INSTRUMENTATION↗

Pushing the frontiers in climate modelling and analysis with machine learning

Climate modelling and analysis are facing new demands to enhance projections and climate information. Here, in this study, we argue that now is the time to push the frontiers of machine learning beyond state-of-the-art approaches, not only by developing machine-learning-based Earth system models with greater fidelity, but also by providing new capabilities through emulators for extreme event projections with large ensembles, enhanced detection and attribution methods for extreme events, and advanced climate model analysis and benchmarking. Utilizing this potential requires key machine learning challenges to be addressed, in particular generalization, uncertainty quantification, explainable artificial intelligence and causality. This interdisciplinary effort requires bringing together machine learning and climate scientists, while also leveraging the private sector, to accelerate progress towards actionable climate science.

54 ENVIRONMENTAL SCIENCES↗

Localization of infrasonic sources via Bayesian back projection

SUMMARY A Bayesian framework is investigated for event-specific localization of infrasonic sources using back projection ray tracing. Direction-of-arrival information from array-based detection analysis is used to initialize a back projection ray path originating from the detecting array location and quantifying propagation characteristics from hypothetical source locations. The Fisher statistic, computed from the array’s beam coherence, is mapped into uncertainty in the launch angles of the ray path. Auxiliary parameters previously introduced for solving the Transport equation to compute geometric spreading along ray paths are used to map uncertainty in the ray launch angles into spatial and temporal uncertainties in the ray path. An atmospheric ensemble approach is applied to account for atmospheric uncertainty, and the relation between uncertainties in the atmospheric state and confidence in estimated localization are evaluated using several ensembles with specified variances. The method is evaluated using a synthetic event in the western United States constructed via forward propagation simulations as well as a single-station, multi-arrival detection from a surface explosion in the western United States. Localization results using this event-specific approach are more accurate and exhibit improved precision than existing Bayesian localization methods that leverage generalized, pre-computed propagation statistics.

58 GEOSCIENCES↗

Designs from Local Random Quantum Circuits with SU ( d ) Symmetry

The generation of k -designs (pseudorandom distributions that emulate the Haar measure up to k moments) with local quantum circuit ensembles is a problem of fundamental importance in quantum information and physics. Despite the extensive understanding of this problem for ordinary random circuits, the crucial situations in which symmetries or conservation laws are in play are known to pose fundamental challenges and remain little understood. Here, we construct explicit local unitary ensembles that can achieve high-order unitary k -designs under transversal continuous symmetry, in the particularly important SU ( d ) case. Specifically, we define the convolutional quantum alternating (CQA) group generated by 4-local SU ( d ) -symmetric Hamiltonians as well as associated 4-local SU ( d ) -symmetric random unitary circuit ensembles and prove that they form and converge to SU ( d ) -symmetric k -designs, respectively, for all k < n ( n − 3 ) / 2 , with n being the number of qudits. A key technique that we employ to obtain the results is the Okounkov-Vershik approach to S n representation theory. To study the convergence time of the CQA ensemble, we develop a numerical method using the Young orthogonal form and the S n branching rule. We provide strong evidence for a subconstant spectral gap and certain convergence time scales of various important circuit architectures, which contrast with the symmetry-free case. We also provide comprehensive explanations of the difficulties and limitations in rigorously analyzing the convergence time using methods that have been effective for cases without symmetries, including Knabe’s local gap threshold and Nachtergaele’s martingale methods. This suggests that a novel approach is likely necessary for understanding the convergence time of SU ( d ) -symmetric local random circuits. Published by the American Physical Society 2024

Li, Zimu (ORCID:0000000314736492)↗

Decomposing Cloud Radiative Feedbacks by Cloud-Top Phase

Changes in cloud scattering properties and emissivity that arise from atmospheric warming cause substantial radiative feedbacks in model projections of anthropogenic climate change, and the relative importance of the underlying mechanisms is poorly understood. One leading hypothesis is that ice-to-liquid conversions cause clouds to optically thicken, producing a major negative feedback. We test this hypothesis by developing a method to decompose cloud radiative feedbacks by cloud-top phase. The method is applied to an ensemble of six state-of-the-art global climate models run with prescribed sea surface temperature. In these simulations, the global mean of the net cloud scattering and emissivity feedback from cloud-phase conversions ranges from −0.17 to −0.01 W m −2 K −1 , while the overall net cloud feedback ranges from 0.02 to 0.91 W m −2 K −1 . The multimodel mean of the cloud scattering and emissivity feedback from cloud-phase conversions is approximately 19% of the magnitude of the multimodel mean of the overall cloud feedback (−0.10 vs 0.52 W m −2 K −1 ). These results indicate that cloud-phase conversions cause a robust negative feedback by changing cloud scattering and emissivity, but this mechanism makes a modest contribution to the overall cloud feedback at the global scale.

Climate change↗

Quantification of regional net CO 2 flux errors in the Orbiting Carbon Observatory-2 (OCO-2) v10 model intercomparison project (MIP) ensemble using airborne measurements

Inverse model intercomparison projects (MIPs) provide a chance to assess the uncertainties in inversion estimates arising from various sources. However, accurately quantifying ensemble CO 2 flux errors remains challenging and often relies on the ensemble spread. This study proposes a method for quantifying the errors in regional net surface–atmosphere CO 2 flux estimates from models taken from the Orbiting Carbon Observatory-2 (OCO-2) v10 MIP by using independent airborne CO 2 measurements for the period 2015–2017. We first calculate the root mean square error (RMSE) between the ensemble mean of posterior CO 2 concentrations and airborne observations and then isolate the CO 2 concentration errors caused solely by the ensemble mean of posterior net fluxes by subtracting the observation, representation, and transport errors from seven regions. Our analysis reveals that the flux errors projected onto CO 2 space account for 55 %–85 % of the regional average RMSE over the 3 years, ranging from 0.88 to 1.91 ppm. In five regions, the error estimates based on observations exceed those computed from the ensemble spread of posterior fluxes by a factor of 1.33–1.93, implying an underestimation of the actual flux errors, while their magnitudes are comparable in two regions. The adjoint sensitivity analysis identifies that the underestimation of flux errors is prominent where the magnitudes of fossil fuel emissions exceed those of terrestrial-biosphere fluxes by a factor of 3–31 over the 3 years. This suggests the presence of systematic biases in the inversion estimates associated with errors in the prescribed fossil fuel emissions common to all models. Our study emphasizes the value of airborne measurements for quantifying regional errors in ensemble net CO 2 flux estimates.

54 ENVIRONMENTAL SCIENCES↗

Bootstrap-determined p values in lattice QCD

We present a general method to determine the probability that stochastic Monte Carlo data, in particular those generated in a lattice QCD calculation, would have been obtained were that data drawn from the distribution predicted by a given theoretical hypothesis. Such a probability, or p -value, is often used as an important heuristic measure of the validity of that hypothesis. The proposed method offers the benefit that it remains usable in cases where the standard Hotelling T 2 methods based on the conventional χ 2 statistic do not apply, such as for uncorrelated fits. Specifically, we analyze q 2 , defined as the correlated χ 2 statistic obtained using an arbitrary covariance matrix estimator, and show how to use the bootstrap as a data-driven method to determine the expected distribution of q 2 for a given hypothesis with minimal assumptions. This distribution can then be used to determine the p -value for a fit to the data. We also describe a bootstrap approach for quantifying the impact upon this p -value of estimating population parameters from a single ensemble of N samples. The overall method is accurate up to a 1 / N bias which we do not attempt to quantify. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Discovering the Multisectoral Impacts of Global Energy Sector Outcomes Through Multiple Ensemble Aggregation Measures

Understanding complex human-Earth system interactions often involves analyzing large scenario ensembles that encompass a wide range of plausible futures. These ensembles often require aggregation to summarize information based on specific criteria or conditions. However, previous research using global change scenario ensembles has largely overlooked how the choice of aggregation method influences the interpretation of results. To address this gap, we leverage a large ensemble data set designed to capture broad energy system dynamics generated using the Global Change Analysis Model. We first explore how energy-related uncertainties are propagated to both global and regional water-energy-food sectors. We then conduct a rank correlation analysis across seven ensemble aggregation measures and demonstrate the need to consider multiple measures in global change scenarios. Our results suggest that global water and food sector outcomes in the 21st century vary widely depending on different scenario assumptions. The global energy productivity is projected to improve by the end of the century across all scenarios. Moreover, regions facing water scarcity challenges in 2100 do not always overlap with those facing extreme energy and food sector outcomes. Although rank correlations across seven aggregation measures are relatively stable across sectors, we identify cases where relying on a single measure leads to losing critical information in the full ensemble. Reliance on a single aggregation measure can distort the interpretation of global change scenario outcomes. Instead, adopting multiple ensemble aggregation measures provides a more holistic understanding of global change scenario ensembles.

Kim, Gijoo↗

Gradient-informed Hamiltonian Monte Carlo for multicomponent CALPHAD model optimization and uncertainty quantification

CALPHAD model parameter optimization is inherently challenging due to non-smooth objective functions, high-dimensional parameter spaces, and the need for uncertainty quantification (UQ). Traditional weighted nonlinear least squares approaches are computationally efficient but local, whereas black-box global optimizers and ensemble Markov Chain Monte Carlo (MCMC) methods provide broader exploration at substantial computational cost. The objective of this work is to combine the global exploration capability of gradient-informed Hamiltonian Monte Carlo – specifically the No-U-Turn Sampler (NUTS) – with local deterministic refinement using BFGS to efficiently optimize multicomponent CALPHAD models with minimal manual intervention. Analytic gradients are computed via the Jansson derivative framework. The methodology is demonstrated on the Cr—Fe binary system and extended to the Cr—Fe—Ni ternary system with 32 degrees of freedom. For Cr—Fe, NUTS achieves comparable or superior optimality relative to ensemble MCMC while requiring over an order-of-magnitude fewer likelihood evaluations. Parameter uncertainties are quantified through NUTS sampling and propagated to thermodynamic observables using local expansion, demonstrating a novel modular approach that combines binary and ternary parameter subsets without requiring global relaxation. These results establish gradient-informed exploration as a scalable strategy for multicomponent CALPHAD optimization and provide a practical route towards efficient higher-order database development with quantified uncertainty.

36 MATERIALS SCIENCE↗

Estimating the CO 2 Fertilization Effect on Extratropical Forest Productivity From Flux‐Tower Observations

Abstract The land sink of anthropogenic carbon emissions, a crucial component of mitigating climate change, is primarily attributed to the CO 2 fertilization effect on global gross primary productivity (GPP). However, direct observational evidence of this effect remains scarce, hampered by challenges in disentangling the CO 2 fertilization effect from other long‐term confounding drivers, particularly climatic changes. Here, we introduce a novel statistical approach to separate the CO 2 fertilization effect on photosynthetic carbon uptake using eddy covariance (EC) records across 38 extratropical forest sites. We find the median stimulation rate of GPP to be 3.2 ± 0.9 gC m −2 yr −1 ppm −1 (or 16.4 ± 4.2% per 100 ppm) under increasing atmospheric CO 2 across these sites, respectively. To validate the robustness of our findings, we test our statistical method using factorial simulations of an ensemble of process‐based land surface models. We address additional factors, including nitrogen deposition and land management, that may impact plant productivity, potentially confounding the attribution to the CO 2 fertilization effect. Assuming these site‐specific effects offset to some extent across sites as random factors, the estimated median value still reflects the strength of the CO 2 fertilization effect. However, disentanglement of these long‐term effects, often inseparable by timescale, requires further causal research. Our study provides direct evidence that the photosynthetic stimulation is maintained under long‐term CO 2 fertilization across multiple EC sites. Such observation‐based quantification is key to constraining the long‐standing uncertainties in the land carbon cycle under rising CO 2 concentrations.

Environmental Sciences & Ecology↗

Effective optimization of atomic decoration in giant and superstructurally ordered crystals with machine learning

Crystals with complicated geometry are often observed with mixed chemical occupancy among Wyckoff sites, presenting a unique challenge for accurate atomic modeling. Similar systems possessing exact occupancy on all the sites can exhibit superstructural ordering, dramatically inflating the unit cell size. In this work, a crystal graph convolutional neural network (CGCNN) is used to predict optimal atomic decorations on fixed crystalline geometries. This is achieved with a site permutation search (SPS) optimization algorithm based on Monte Carlo moves combined with simulated annealing and basin-hopping techniques. Our approach relies on the evidence that, for a given chemical composition, a CGCNN estimates the correct energetic ordering of different atomic decorations, as predicted by electronic structure calculations. This provides a suitable energy landscape that can be optimized according to site occupation, allowing the prediction of chemical decoration in crystals exhibiting mixed or disordered occupancy, or superstructural ordering. Verification of the procedure is carried out on several known compounds, including the superstructurally ordered clathrate compound Rb8Ga27Sb16 and vacancy-ordered perovskite Cs2SnI6, neither of which was previously seen during the neural network training. In addition, the critical temperature of an order–disorder phase transition in solid solution CuZn is probed with our SPS routines by sampling site configuration trajectories in the canonical ensemble. This strategy provides an accurate method for determining favorable decoration in complex crystals and analyzing site occupation at unprecedented speed and scale.

Chemistry↗

A novel conditional generative model for efficient ensemble forecasts of state variables in large-scale geological carbon storage

Integrating monitoring data to efficiently update reservoir pressure and CO 2 plume distribution forecasts presents a significant challenge in geological carbon storage (GCS) applications. Inverse modeling techniques are commonly used to fuse observational data and refine reservoir model parameters, thereby improving state variable forecasts. However, these techniques often rely on linear or Gaussian assumptions, which can limit their effectiveness in accurately predicting state variables. Moreover, simulating large-scale three-dimensional (3D) GCS problems is computationally expensive, making iterative runs in inverse problems prohibitive. To address these challenges, we propose a conditional generative model utilizing the score-based diffusion method for real-time 3D pressure and saturation field distribution predictions. Our approach involves solving the score function with a mini-batch-based Monte Carlo estimator to generate labeled data. This data is subsequently employed to train a fully connected neural network, enabling it to learn the conditional sample generator within a supervised learning framework. This method enables the rapid generation of a large ensemble of predictions, facilitating comprehensive uncertainty quantification of state variables. Here we applied our method to forecast the dynamic 3D distributions of pressure and saturation fields over a 30-year injection period. The statistical assessment with low root mean square error (RMSE) values demonstrates that our method can accurately predict the spatiotemporal distributions of both pressure and saturation fields. Moreover, the developed conditional generative model shows high computational efficiency by generating 100 ensemble forecasts of 3D state variables in less than 10 min. The consistency between ensemble averages and ground truth values further illustrates the model’s capability to capture state variable dynamics during the CO 2 plume injection process. Notably, the ground truth values fall within the ensemble forecasts, indicating that our uncertainty quantification effectively captures variability and potential noise in the observations. Thus, the developed conditional generative model proves to be a more efficient, accurate, and practical tool for GCS applications, facilitating timely risk analysis and informed decision-making.

58 GEOSCIENCES↗

Efficient Unitary Designs from Random Sums and Permutations

A unitary k-design is an ensemble of unitaries that matches the first k moments of the Haar measure. In this work, we provide two efficient constructions of k-designs on n-qubits using new random matrix theory techniques. Our first construction is based on exponentiating sums of random i.i.d. Hermitian matrices and uses O(k2n2)-many gates. In the spirit of central limit theorems, we show that this random sum approximates the Gaussian Unitary Ensemble (GUE). We then show that the product of just two exponentiated GUE matrices is already approximately Haar random. Our second construction is based on products of exponentiated sums of random permutations and uses Õ(k poly (n)) many gates. The k dependence is optimal (up to polylogarithmic factors) and is inherited from the efficiency of existing k-wise independent permutations. Furthermore, replacing random permutations with quantum-secure pseudorandom permutations (PRPs), we also obtain a pseudorandom unitary (PRU) ensemble that is secure under nonadaptive queries. A central feature of both proofs is a new connection between the polynomial method in quantum query complexity and the large-dimension (N) expansion in random matrix theory. In particular, the first construction uses the polynomial method to control high moments of certain random matrix ensembles without requiring delicate Weingarten calculations. In doing so, we define and solve a moment problem on the unit circle, asking whether a finite number of equally weighted points can reproduce a given set of moments. In our second construction, the key step is to exhibit an orthonormal basis for irreducible representations of the partition algebra that has a low-degree large-N expansion. This allows us to show that the distinguishing probability is a low-degree rational polynomial of the dimension N.

algebra↗