Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Interaction of a cumulus cloud ensemble with the large-scale environment. III - Semi-prognostic test of the Arakawa-Schubert cumulus parameterization

The verification of the Arakawa and Schubert (1974) cumulus parameterization is continued using a semiprognostic approach. Observed data from Phase III of GATE are used to provide estimates of the large-scale forcing of a cumulus ensemble at each observation time. Instantaneous values of the precipitation and the warming and drying due to cumulus convection are calculated using the parameterization. The results show that the calculated precipitation agrees very well with estimates from the observed large-scale moisture budget and from radar observations. The calculated vertical profiles of cumulus warming and drying also are quite similar to the observed. It is shown that the closure assumption adopted in the parameterization (the cloud-work function quasi-equilibrium) results in errors of generally less than 10% in the calculated precipitation. The sensitivity of the parameterization to some assumptions of the cloud ensemble model and the solution method for the cloud-base mass flux is investigated.

Lord, S. J.↗

T-Matrix Modeling of Linear Depolarization by Morphologically Complex Soot and Soot-Containing Aerosols

We use state-of-the-art public-domain Fortran codes based on the T-matrix method to calculate orientation and ensemble averaged scattering matrix elements for a variety of morphologically complex black carbon (BC) and BC-containing aerosol particles, with a special emphasis on the linear depolarization ratio (LDR). We explain theoretically the quasi-Rayleigh LDR peak at side-scattering angles typical of low-density soot fractals and conclude that the measurement of this feature enables one to evaluate the compactness state of BC clusters and trace the evolution of low-density fluffy fractals into densely packed aggregates. We show that small backscattering LDRs measured with groundbased, airborne, and spaceborne lidars for fresh smoke generally agree with the values predicted theoretically for fluffy BC fractals and densely packed near-spheroidal BC aggregates. To reproduce higher lidar LDRs observed for aged smoke, one needs alternative particle models such as shape mixtures of BC spheroids or cylinders.

atmospheric radiation↗

Application of Ensemble Detection and Analysis to Modeling Uncertainty in Non Stationary Process

Characterization of non stationary and nonlinear processes is a challenge in many engineering and scientific disciplines. Climate change modeling and projection, retrieving information from Doppler measurements of hydrometeors, and modeling calibration architectures and algorithms in microwave radiometers are example applications that can benefit from improvements in the modeling and analysis of non stationary processes. Analyses of measured signals have traditionally been limited to a single measurement series. Ensemble Detection is a technique whereby mixing calibrated noise produces an ensemble measurement set. The collection of ensemble data sets enables new methods for analyzing random signals and offers powerful new approaches to studying and analyzing non stationary processes. Derived information contained in the dynamic stochastic moments of a process will enable many novel applications.

Racette, Paul↗

Viral coefficient and hidden mass in the galaxy groups

The purpose is the verification of the virial mass estimations for small galaxy groups. The dynamical evolution of triple and quintuple galaxies was studied by the numerical simulations. The dependence of the virial coefficient k(t) versus time was derived. Initial k(O) = O. The function k(t) has some strong oscillations from 0.02 to 0.99. Generally, these oscillations are quasiperiodical ones. Such a behavior of k(t) is caused by formation in a system of close isolated temporary double subsystems. A strong correlation between the virial coefficient and the least mutual distance in the system is observed. Such wide oscillations may add into the estimation of virial mass of the galaxy groups an uncertainty of more than one order. An additional uncertainty is introduced by the projection effect. This uncertainty for the individual estimations of the masses approach three orders. Thus any individual estimation of the virial mass is impossible for small galaxy groups. Some possibility of statistical estimation (median or average) of the total mass, including a hidden mass, is shown for the homogeneous samples. The authors propose a method for these estimations based on a comparison of the medians of dynamical parameters (a mean size in projection and a dispersion of relative radial velocities) for the simulated and observed ensembles of the galaxy groups. This method has been applied to a sample of 46 probably physical triplets of galaxies. The probable median of the hidden mass in a volume of the triplet is about 4 M, where M is the total mass of visible matter.

Anosova, Joanna P.↗

A Particle Batch Smoother Approach to Snow Water Equivalent Estimation

This paper presents a newly proposed data assimilation method for historical snow water equivalent SWE estimation using remotely sensed fractional snow-covered area fSCA. The newly proposed approach consists of a particle batch smoother (PBS), which is compared to a previously applied Kalman-based ensemble batch smoother (EnBS) approach. The methods were applied over the 27-yr Landsat 5 record at snow pillow and snow course in situ verification sites in the American River basin in the Sierra Nevada (United States). This basin is more densely vegetated and thus more challenging for SWE estimation than the previous applications of the EnBS. Both data assimilation methods provided significant improvement over the prior (modeling only) estimates, with both able to significantly reduce prior SWE biases. The prior RMSE values at the snow pillow and snow course sites were reduced by 68%-82% and 60%-68%, respectively, when applying the data assimilation methods. This result is encouraging for a basin like the American where the moderate to high forest cover will necessarily obscure more of the snow-covered ground surface than in previously examined, less-vegetated basins. The PBS generally outperformed the EnBS: for snow pillows the PBSRMSE was approx.54%of that seen in the EnBS, while for snow courses the PBSRMSE was approx.79%of the EnBS. Sensitivity tests show relative insensitivity for both the PBS and EnBS results to ensemble size and fSCA measurement error, but a higher sensitivity for the EnBS to the mean prior precipitation input, especially in the case where significant prior biases exist.

EnBS↗

How well do one-electron self-interaction-correction methods perform for systems with fractional electrons?

Recently developed locally scaled self-interaction correction (LSIC) is a one-electron SIC method that, when used with a ratio of kinetic energy densities (z σ ) as iso-orbital indicator, performs remarkably well for both thermochemical properties as well as for barrier heights overcoming the paradoxical behavior of the well-known Perdew–Zunger self-interaction correction (PZSIC) method. In this work, we examine how well the LSIC method performs for the delocalization error. Our results show that both LSIC and PZSIC methods correctly describe the dissociation of $H$$^{+}_{2}$ and $H$$^{+}_{2}$ but LSIC is overall more accurate than the PZSIC method. Likewise, in the case of the vertical ionization energy of an ensemble of isolated He atoms, the LSIC and PZSIC methods do not exhibit delocalization errors. For the fractional charges, both LSIC and PZSIC significantly reduce the deviation from linearity in the energy vs number of electrons curve, with PZSIC performing superior for C, Ne, and Ar atoms while for Kr they perform similarly. The LSIC performs well at the endpoints (integer occupations) while substantially reducing the deviation. The dissociation of LiF shows both LSIC and PZSIC dissociate into neutral Li and F but only LSIC exhibits charge transfer from Li + to F – at the expected distance from the experimental data and accurate ab initio data. Overall, both the PZSIC and LSIC methods reduce the delocalization errors substantially.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluation of a nonlinear method for the enhancement of tonal signal detection

A method is presented for biasing spectral estimates to enhance detection of tonal signals against a background of broadband noise. In this method, a nonlinear average of an ensemble of individual spectral estimates is made where broadband noise energy is biased downward, pure tone energy is unbiased, and a mixture of the two is biased by an amount that depends on the ratio of tonal energy to broadband energy. The method is analyzed to provide estimates of the extent of tonal signal detection enhancement.

Garber, Donald P.↗

Microphysics, Radiation and Surface Processes in the Goddard Cumulus Ensemble (GCE) Model

One of the most promising methods to test the representation of cloud processes used in climate models is to use observations together with Cloud Resolving Models (CRMs). The CRMs use more sophisticated and realistic representations of cloud microphysical processes, and they can reasonably well resolve the time evolution, structure, and life cycles of clouds and cloud systems (size about 2-200 km). The CRMs also allow explicit interaction between out-going longwave (cooling) and in-coming solar (heating) radiation with clouds. Observations can provide the initial conditions and validation for CRM results. The Goddard Cumulus Ensemble (GCE) Model, a CRM, has been developed and improved at NASA/Goddard Space Flight Center over the past two decades. The GCE model has been used to understand the following: 1) water and energy cycles and their roles in the tropical climate system; 2) the vertical redistribution of ozone and trace constituents by individual clouds and well organized convective systems over various spatial scales; 3) the relationship between the vertical distribution of latent heating (phase change of water) and the large-scale (pre-storm) environment; 4) the validity of assumptions used in the representation of cloud processes in climate and global circulation models; and 5) the representation of cloud microphysical processes and their interaction with radiative forcing over tropical and midlatitude regions. Four-dimensional cloud and latent heating fields simulated from the GCE model have been provided to the TRMM Science Data and Information System (TSDIS) to develop and improve algorithms for retrieving rainfall and latent heating rates for TRMM and the NASA Earth Observing System (EOS). More than 90 referred papers using the GCE model have been published in the last two decades. Also, more than 10 national and international universities are currently using the GCE model for research and teaching. In this talk, five specific major GCE improvements: (1) ice microphysics, (2) longwave and shortwave radiative transfer processes, (3) land surface processes, (4) ocean surface fluxes and (5) ocean mixed layer processes are presented. The performance of these new GCE improvements will be examined. Observations are used for model validation.

Tao, Wei-Kuo↗

Quantification of regional net CO 2 flux errors in the Orbiting Carbon Observatory-2 (OCO-2) v10 model intercomparison project (MIP) ensemble using airborne measurements

Inverse model intercomparison projects (MIPs) provide a chance to assess the uncertainties in inversion estimates arising from various sources. However, accurately quantifying ensemble CO 2 flux errors remains challenging and often relies on the ensemble spread. This study proposes a method for quantifying the errors in regional net surface–atmosphere CO 2 flux estimates from models taken from the Orbiting Carbon Observatory-2 (OCO-2) v10 MIP by using independent airborne CO 2 measurements for the period 2015–2017. We first calculate the root mean square error (RMSE) between the ensemble mean of posterior CO 2 concentrations and airborne observations and then isolate the CO 2 concentration errors caused solely by the ensemble mean of posterior net fluxes by subtracting the observation, representation, and transport errors from seven regions. Our analysis reveals that the flux errors projected onto CO 2 space account for 55 %–85 % of the regional average RMSE over the 3 years, ranging from 0.88 to 1.91 ppm. In five regions, the error estimates based on observations exceed those computed from the ensemble spread of posterior fluxes by a factor of 1.33–1.93, implying an underestimation of the actual flux errors, while their magnitudes are comparable in two regions. The adjoint sensitivity analysis identifies that the underestimation of flux errors is prominent where the magnitudes of fossil fuel emissions exceed those of terrestrial-biosphere fluxes by a factor of 3–31 over the 3 years. This suggests the presence of systematic biases in the inversion estimates associated with errors in the prescribed fossil fuel emissions common to all models. Our study emphasizes the value of airborne measurements for quantifying regional errors in ensemble net CO 2 flux estimates.

54 ENVIRONMENTAL SCIENCES↗

Bootstrap-determined p values in lattice QCD

We present a general method to determine the probability that stochastic Monte Carlo data, in particular those generated in a lattice QCD calculation, would have been obtained were that data drawn from the distribution predicted by a given theoretical hypothesis. Such a probability, or p -value, is often used as an important heuristic measure of the validity of that hypothesis. The proposed method offers the benefit that it remains usable in cases where the standard Hotelling T 2 methods based on the conventional χ 2 statistic do not apply, such as for uncorrelated fits. Specifically, we analyze q 2 , defined as the correlated χ 2 statistic obtained using an arbitrary covariance matrix estimator, and show how to use the bootstrap as a data-driven method to determine the expected distribution of q 2 for a given hypothesis with minimal assumptions. This distribution can then be used to determine the p -value for a fit to the data. We also describe a bootstrap approach for quantifying the impact upon this p -value of estimating population parameters from a single ensemble of N samples. The overall method is accurate up to a 1 / N bias which we do not attempt to quantify. Published by the American Physical Society 2025

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Enhancing Solar Power Forecasting with Regularized Constrained Quantile Regression Averaging and Bootstrapping Techniques

Probabilistic solar power forecasting (SPF) plays an essential role in optimizing power-grid operations by quantifying the forecast uncertainty. To improve the accuracy and robustness of probabilistic SPF, this paper introduces the regularized constrained quantile regression averaging (rCQRA) method to combine outputs from multiple PSPF models. In addition, a bootstrapping method was used to quantify model uncertainty, providing insights into the reliability and significance of each ensemble component. To evaluate its efficacy, the proposed rCQRA method is used to integrate four PSPF methods. The resulting SPF models are trained and validated using a real-world six-year dataset from a rooftop solar plant in the USA. The performance of the proposed rCQRA method is evaluated and compared with two benchmark methods under three categories of weather conditions. It is shown that the rCQRA method has superior performance in its forecast reliability, sharpness, and accuracy.

Ensemble learning, probabilistic solar power forec↗

Clarifying trust of materials property predictions using neural networks with distribution-specific uncertainty quantification

It is critical that machine learning (ML) model predictions be trustworthy for high-throughput catalyst discovery approaches. Uncertainty quantification (UQ) methods allow estimation of the trustworthiness of an ML model, but these methods have not been well explored in the field of heterogeneous catalysis. Herein, we investigate different UQ methods applied to a crystal graph convolutional neural network to predict adsorption energies of molecules on alloys from the Open Catalyst 2020 dataset, the largest existing heterogeneous catalyst dataset. We apply three UQ methods to the adsorption energy predictions, namely k-fold ensembling, Monte Carlo dropout, and evidential regression. The effectiveness of each UQ method is assessed based on accuracy, sharpness, dispersion, calibration, and tightness. Evidential regression is demonstrated to be a powerful approach for rapidly obtaining tunable, competitively trustworthy UQ estimates for heterogeneous catalysis applications when using neural networks. Recalibration of model uncertainties is shown to be essential in practical screening applications of catalysts using uncertainties.

36 MATERIALS SCIENCE↗

Discovering the Multisectoral Impacts of Global Energy Sector Outcomes Through Multiple Ensemble Aggregation Measures

Understanding complex human-Earth system interactions often involves analyzing large scenario ensembles that encompass a wide range of plausible futures. These ensembles often require aggregation to summarize information based on specific criteria or conditions. However, previous research using global change scenario ensembles has largely overlooked how the choice of aggregation method influences the interpretation of results. To address this gap, we leverage a large ensemble data set designed to capture broad energy system dynamics generated using the Global Change Analysis Model. We first explore how energy-related uncertainties are propagated to both global and regional water-energy-food sectors. We then conduct a rank correlation analysis across seven ensemble aggregation measures and demonstrate the need to consider multiple measures in global change scenarios. Our results suggest that global water and food sector outcomes in the 21st century vary widely depending on different scenario assumptions. The global energy productivity is projected to improve by the end of the century across all scenarios. Moreover, regions facing water scarcity challenges in 2100 do not always overlap with those facing extreme energy and food sector outcomes. Although rank correlations across seven aggregation measures are relatively stable across sectors, we identify cases where relying on a single measure leads to losing critical information in the full ensemble. Reliance on a single aggregation measure can distort the interpretation of global change scenario outcomes. Instead, adopting multiple ensemble aggregation measures provides a more holistic understanding of global change scenario ensembles.

Kim, Gijoo↗

Behavioral Ensemble CLM5 Hydrological Parameter Sets

This repository contains hydrological parameter sets derived using the hybrid regionalization method for three distinct streamflow signatures: Streamflow Signatures: Q10: Represents low flow, indicating the nonexceedance probability of 0.1 for daily streamflow. Q90: Represents high flow, with a nonexceedance probability of 0.9 for daily streamflow. Qmean: Indicates the mean annual flow. Parameters for 464 CAMELS Basins: CAMELS_1000_parameters.csv: Contains 1,000 ensemble parameter sets generated using the Latin hypercube sampling method for CLM5, encompassing 15 hydrological parameters. CAMELS_q10_behavioral_parameter_num.csv: Provides the behavioral ensemble parameter sets for the Q10 streamflow signature for each basin. The associated ID number refers to entries in the CAMELS_1000_parameters.csv file. A minimum of 10 ensemble parameter sets are available for each basin. CAMELS_q90_behavioral_parameter_num.csv: Similar to the above file but for the Q90 streamflow signature. CAMELS_qmean_behavioral_parameter_num.csv: Corresponds to the Qmean streamflow signature, similar to the previous files. Parameters for 50,629 1/8° CONUS Land Grid Cells: CONUS_350_parameters.csv: Contains 350 ensemble parameter sets derived using the Latin hypercube sampling method for CLM5's 15 hydrological parameters within 1/8° CONUS land grid cells. CONUS_q10_behavioral_parameter_num.csv: Holds the behavioral ensemble parameter sets for the Q10 streamflow signature, organized for each grid cell. The ID number relates to entries in CONUS_350_parameters.csv. A minimum of 10 ensemble parameter sets are provided for each grid cell. CONUS_q90_behavioral_parameter_num.csv: Similar to the above file but focusing on the Q90 streamflow signature. CONUS_qmean_behavioral_parameter_num.csv: Corresponds to the Qmean streamflow signature, following a similar structure to the previous files.

Yan, Hongxiang↗

Gradient-informed Hamiltonian Monte Carlo for multicomponent CALPHAD model optimization and uncertainty quantification

CALPHAD model parameter optimization is inherently challenging due to non-smooth objective functions, high-dimensional parameter spaces, and the need for uncertainty quantification (UQ). Traditional weighted nonlinear least squares approaches are computationally efficient but local, whereas black-box global optimizers and ensemble Markov Chain Monte Carlo (MCMC) methods provide broader exploration at substantial computational cost. The objective of this work is to combine the global exploration capability of gradient-informed Hamiltonian Monte Carlo – specifically the No-U-Turn Sampler (NUTS) – with local deterministic refinement using BFGS to efficiently optimize multicomponent CALPHAD models with minimal manual intervention. Analytic gradients are computed via the Jansson derivative framework. The methodology is demonstrated on the Cr—Fe binary system and extended to the Cr—Fe—Ni ternary system with 32 degrees of freedom. For Cr—Fe, NUTS achieves comparable or superior optimality relative to ensemble MCMC while requiring over an order-of-magnitude fewer likelihood evaluations. Parameter uncertainties are quantified through NUTS sampling and propagated to thermodynamic observables using local expansion, demonstrating a novel modular approach that combines binary and ternary parameter subsets without requiring global relaxation. These results establish gradient-informed exploration as a scalable strategy for multicomponent CALPHAD optimization and provide a practical route towards efficient higher-order database development with quantified uncertainty.

36 MATERIALS SCIENCE↗

Urban Flood Modeling: Uncertainty Quantification and Physics‐Informed Gaussian Processes Regression Forecasting

Abstract Estimating uncertainty in flood model predictions is important for many applications, including risk assessment and flood forecasting. We focus on uncertainty in physics‐based urban flooding models. We consider the effects of the model's complexity and uncertainty in key input parameters. The effect of rainfall intensity on the uncertainty in water depth predictions is also studied. As a test study, we choose the Interconnected Channel and Pond Routing (ICPR) model of a part of the city of Minneapolis. The uncertainty in the ICPR model's predictions of the floodwater depth is quantified in terms of the ensemble variance using the multilevel Monte Carlo (MC) simulation method. Our results show that uncertainties in the studied domain are highly localized. Model simplifications, such as disregarding the groundwater flow, lead to overly confident predictions, that is, predictions that are both less accurate and uncertain than those of the more complex model. We find that for the same number of uncertain parameters, increasing the model resolution reduces uncertainty in the model predictions (and increases the MC method's computational cost). We employ the multilevel MC method to reduce the cost of estimating uncertainty in a high‐resolution ICPR model. Finally, we use the ensemble estimates of the mean and covariance of the flood depth for real‐time flood depth forecasting using the physics‐informed Gaussian process regression method. We show that even with few measurements, the proposed framework results in a more accurate forecast than that provided by the mean prediction of the ICPR model.

Kohanpur, Amir H.↗

GMRES with embedded ensemble propagation for the efficient solution of parametric linear systems in uncertainty quantification of computational models

In a previous work, embedded ensemble propagation was proposed to improve the efficiency of sampling-based uncertainty quantification methods of computational models on emerging computational architectures. It consists of simultaneously evaluating the model for a subset of samples together, instead of evaluating them individually. A first approach introduced to solve parametric linear systems with ensemble propagation is ensemble reduction. In Krylov methods for example, this reduction consists in coupling the samples together using an inner product that sums the sample contributions. Ensemble reduction has the advantages of being able to use optimized implementations of BLAS functions and having a stopping criterion which involves only one scalar. However, the reduction potentially decreases the rate of convergence due to the gathering of the spectra of the samples. In this paper, we investigate a second approach: ensemble propagation without ensemble reduction in the case of GMRES. This second approach solves each sample simultaneously but independently to improve the convergence compared to ensemble reduction. This raises two new issues which are solved in this paper: the fact that optimized implementations of BLAS functions cannot be used anymore and that ensemble divergence, whereby individual samples within an ensemble must follow different code execution paths, can occur. We tackle those issues by implementing a high-performing ensemble GEMV and by using masks. The proposed ensemble GEMV leads to a similar cost per GMRES iteration for both approaches, i.e. with and without reduction. For illustration, we study the performances of the new linear solver in the context of a mesh tying problem. Furthermore, this example demonstrates improved ensemble propagation speed-up without reduction.

BLAS↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗