Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

High dimensional predictions of suicide risk in 4.2 million US Veterans using ensemble transfer learning

We present an ensemble transfer learning method to predict suicide from Veterans Affairs (VA) electronic medical records (EMR). A diverse set of base models was trained to predict a binary outcome constructed from reported suicide, suicide attempt, and overdose diagnoses with varying choices of study design and prediction methodology. Each model used twenty cross-sectional and 190 longitudinal variables observed in eight time intervals covering 7.5 years prior to the time of prediction. Ensembles of seven base models were created and fine-tuned with ten variables expected to change with study design and outcome definition in order to predict suicide and combined outcome in a prospective cohort. The ensemble models achieved c-statistics of 0.73 on 2-year suicide risk and 0.83 on the combined outcome when predicting on a prospective cohort of ~4.2 M veterans. The ensembles rely on nonlinear base models trained using a matched retrospective nested case-control (Rcc) study cohort and show good calibration across a diversity of subgroups, including risk strata, age, sex, race, and level of healthcare utilization. In addition, a linear Rcc base model provided a rich set of biological predictors, including indicators of suicide, substance use disorder, mental health diagnoses and treatments, hypoxia and vascular damage, and demographics. Similar content being viewed by others

60 APPLIED LIFE SCIENCES↗

Multi-Agency Ensemble Forecast of Wildfire Air Quality in the United States: Toward Community Consensus of Early Warning

Wildfires pose increasing risks to human health and properties in North America. Due to large uncertainties in fire emission, transport, and chemical transformation, it remains challenging to accurately predict air quality during wildfire events, hindering our collective capability to issue effective early warnings to protect public health and welfare. Here we present a new real-time Hazardous Air Quality Ensemble System (HAQES) by leveraging various wildfire smoke forecasts from three U.S. federal agencies (NOAA, NASA, and Navy). Compared to individual models, the HAQES ensemble forecast significantly enhances forecast accuracy. To further enhance forecasting performance, a weighted ensemble forecast approach was introduced and tested. Compared to the unweighted ensemble mean, the multilinear regression weighted ensemble reduced fractional bias by 34% in the major fire regions, false alarm rate by 72%, and increased hit rate by 17%. Finally, we improved the weighted ensemble using quantile regression and weighted regression methods to enhance the forecast of extreme air quality events. The advanced weighted ensemble increased the PM2.5 exceedance hit rate by 55% compared to the ensemble mean. Our findings provide insights into the development of advanced ensemble forecast methods for wildfire air quality, offering a practical way to enhance decision-making support to protect public health.

Yunyao Li↗

Modeling Intercalation Chemistry with Multiredox Reactions by Sparse Lattice Models in Disordered Rocksalt Cathodes

Modern battery materials can contain many elements with substantial site disorder, and their configurational state has been shown to be critical for their performance. The intercalation voltage profile is a critical parameter to evaluate the performance of energy storage. The application of commonly used cluster expansion techniques to model the intercalation thermodynamics of such systems ab initio is challenged by the combinatorial increase in configurational degrees of freedom as the number of species grows. Such challenges necessitate the efficient generation of lattice models without overfitting and proper sampling of the configurational space under the requirement of charge balance in ionic systems. In this work, we introduce a combined approach that addresses these challenges by (1) constructing a robust cluster expansion Hamiltonian using the sparse regression technique, including -norm regularization and structural hierarchy; and (2) implementing semigrand-canonical Monte Carlo to sample charge-balanced ionic configurations using the table-exchange method and an ensemble average approach. These techniques are applied to a disordered rocksalt oxyfluoride (LMNOF) that is part of a family of promising earth-abundant cathode materials. The simulated voltage profile is found to be in good agreement with experimental data and particularly provides a clear demonstration of the and oxygen contributions to the redox potential as a function of content.

25 ENERGY STORAGE↗

Computational investigation of hysteresis and phase equilibria of n-alkanes in a metal-organic framework with both micropores and mesopores

Abstract Adsorption hysteresis is a phenomenon related to phase transitions that can impact applications such as gas storage and separations in porous materials. Computational approaches can greatly facilitate the understanding of phase transitions and phase equilibria in porous materials. In this work, adsorption isotherms for methane, ethane, propane, and n-hexane were calculated from atomistic grand canonical Monte Carlo (GCMC) simulations in a metal-organic framework having both micropores and mesopores to better understand hysteresis and phase equilibria between connected pores of different size and the external bulk fluid. At low temperatures, the calculated isotherms exhibit sharp steps accompanied by hysteresis. As a complementary simulation method, canonical (NVT) ensemble simulations with Widom test particle insertions are demonstrated to provide additional information about these systems. The NVT+Widom simulations provide the full van der Waals loop associated with the sharp steps and hysteresis, including the locations of the spinodal points and points within the metastable and unstable regions that are inaccessible to GCMC simulations. The simulations provide molecular-level insight into pore filling and equilibria between high- and low-density states within individual pores. The effect of framework flexibility on adsorption hysteresis is also investigated for methane in IRMOF-1.

36 MATERIALS SCIENCE↗

Efficient Agent-Based Cluster Ensembles

Numerous domains ranging from distributed data acquisition to knowledge reuse need to solve the cluster ensemble problem of combining multiple clusterings into a single unified clustering. Unfortunately current non-agent-based cluster combining methods do not work in a distributed environment, are not robust to corrupted clusterings and require centralized access to all original clusterings. Overcoming these issues will allow cluster ensembles to be used in fundamentally distributed and failure-prone domains such as data acquisition from satellite constellations, in addition to domains demanding confidentiality such as combining clusterings of user profiles. This paper proposes an efficient, distributed, agent-based clustering ensemble method that addresses these issues. In this approach each agent is assigned a small subset of the data and votes on which final cluster its data points should belong to. The final clustering is then evaluated by a global utility, computed in a distributed way. This clustering is also evaluated using an agent-specific utility that is shown to be easier for the agents to maximize. Results show that agents using the agent-specific utility can achieve better performance than traditional non-agent based methods and are effective even when up to 50% of the agents fail.

Agogino, Adrian↗

Simulated effects of sample size and grain neighborhood on the modeling of extreme value fatigue response

Assessing the size of representative volume elements (RVEs) for fatigue-related applications is challenging. A RVE relevant to random microstructure requires a volume of material that is sufficiently large to capture the grain/phase heterogeneity that captures all statistical moments of the distribution of the driving force for fatigue crack formation at “hot spot” grains. Consequently, the large size of a microstructure RVE required to study fatigue phenomena is largely computationally intractable and difficult to explore. A more realistic objective in this work is to systematically study, as a function of the size of a statistical sample of microstructure, trends towards convergence of the simulated distribution of driving force for fatigue crack formation. Our present work accordingly leverages the recently developed open-source PRISMS-Fatigue framework to examine the trends in convergence of extreme value distributions (EVD) of Fatigue Indicator Parameters (FIPs) in progressively larger polycrystalline microstructure realizations of FCC Al alloy 7075-T6 using crystal plasticity finite element method simulations. The results are compared to the traditional method in which ensembles of statistical volume elements (SVEs) are simulated to build up statistics intended to approximate those associated with a larger volume of material. The convergence of EVDs with increase of size of a SVE of microstructure is closely related to the extent of grain nearest neighbor (NN) interactions. Accordingly, the sensitivity of the local micromechanical response at hot spot grains is quantitatively investigated by systematically varying the orientations of NN grains. Results indicate that SVEs with cubic crystallographic texture tend towards convergence of the EVD of FIPs with tens of thousands of grains while the random and rolled textures require larger volumes. Simple relationships based on microstructure parameters (e.g., Schmid Factor, grain size, NN misorientation) do not completely correlate to fatigue hot spot grains. Finally, the sensitivity of the extreme value fatigue response at hot spot grains extends to the 3rd NN when a single neighborhood grain orientation is altered.

36 MATERIALS SCIENCE↗

Joint state-parameter estimation for the reduced fracture model via the united filter

Here, in this paper, we introduce an effective United Filter method for jointly estimating the solution state and physical parameters in flow and transport problems within fractured porous media. Fluid flow and transport in fractured porous media are critical in subsurface hydrology, geophysics, and reservoir geomechanics. Reduced fracture models, which represent fractures as lower-dimensional interfaces, enable efficient multi-scale simulations. However, reduced fracture models also face accuracy challenges due to modeling errors and uncertainties in physical parameters such as permeability and fracture geometry. To address these challenges, we propose a United Filter method, which integrates the Ensemble Score Filter (EnSF) for state estimation with the Direct Filter for parameter estimation. EnSF, based on a score-based diffusion model framework, produces ensemble representations of the state distribution without deep learning. Meanwhile, the Direct Filter, a recursive Bayesian inference method, estimates parameters directly from state observations. The United Filter combines these methods iteratively: EnSF estimates are used to refine parameter values, which are then fed back to improve state estimation. Numerical experiments demonstrate that the United Filter method surpasses the state-of-the-art Augmented Ensemble Kalman Filter, delivering more accurate state and parameter estimation for reduced fracture models. This framework also provides a robust and efficient solution for PDE-constrained inverse problems with uncertainties and sparse observations.

Bayesian inference↗

Sequential ensemble transform for Bayesian inverse problems

In this work, we present the Sequential Ensemble Transform (SET) method, an approach for generating approximate samples from a Bayesian posterior distribution. The method explores the posterior distribution by solving a sequence of discrete optimal transport problems to produce a series of transport plans which map prior samples to posterior samples. We prove that the sequence of Dirac mixture distributions produced by the SET method converges weakly to the true posterior as the sample size approaches infinity. Furthermore, our numerical results indicate that, when compared to standard Sequential Monte Carlo (SMC) methods, the SET approach is more robust to the choice of Markov mutation kernels and requires less computational efforts to reach a similar accuracy when used to explore complex posterior distributions. Finally, we describe adaptive schemes that allow to completely automate the use of the SET method.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Active operator learning with predictive uncertainty quantification for partial differential equations

With the increased prevalence of neural operators being used to provide rapid solutions to partial differential equations (PDEs), understanding the accuracy of model predictions and the associated error levels is necessary for deploying reliable surrogate models in scientific applications. Existing uncertainty quantification (UQ) frameworks employ ensembles or Bayesian methods, which can incur substantial computational costs during both training and inference. Here, we propose a lightweight predictive UQ method tailored for Deep operator networks (DeepONets) that also generalizes to other operator networks. Numerical experiments on linear and nonlinear PDEs demonstrate that the framework’s uncertainty estimates are unbiased and provide accurate out-of-distribution uncertainty predictions with a sufficiently large training dataset. Our framework provides fast inference and uncertainty estimates that can efficiently drive outer-loop analyses that would be prohibitively expensive with conventional solvers. We demonstrate how predictive uncertainties can be used in the context of Bayesian optimization and active learning problems to yield improvements in accuracy and data-efficiency for outer-loop optimization procedures. In the active learning setup, we extend the framework to Fourier Neural Operators (FNO) and describe a generalized method for other operator networks. To enable real-time deployment, we introduce an inference strategy based on precomputed trunk outputs and a sparse placement matrix, reducing evaluation time by more than a factor of five. Our method provides a practical route to uncertainty-aware operator learning in time-sensitive settings.

97 MATHEMATICS AND COMPUTING↗

Electron Beam Infrared Nano-Ellipsometry of Individual Indium Tin Oxide Nanocrystals

Leveraging recent advances in electron energy monochromation and aberration correction, we record the spatially resolved infrared plasmon spectrum of individual tin-doped indium oxide nanocrystals using electron energy-loss spectroscopy (EELS). Both surface and bulk plasmon responses are measured as a function of tin doping concentration from 1–10 atomic percent. Furthermore, these results are compared to theoretical models, which elucidate the spectral detuning of the same surface plasmon resonance feature when measured from aloof and penetrating probe geometries. We additionally demonstrate a unique approach to retrieving the fundamental dielectric parameters of individual semiconductor nanocrystals via EELS. This method, devoid from ensemble averaging, illustrates the potential for electron-beam ellipsometry measurements on materials that cannot be prepared in bulk form or as thin films.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Determining Infrared Optical Constants for Dolomite using Single Angle Reflectance and Spectroscopic Ellipsometry

The optical constants, namely the real (n) and imaginary (k) parts of the complex refractive index, are of particular interest to generate the infrared (IR) spectra of liquid and solid materials in different morphologies. To obtain n/k, however, most materials are typically not found in the monolithic forms necessary to easily measure n/k, and thus require the use of other methods such as pressing powders to form planar, specularly-reflective pellets. In this work dolomite crystals are measured using fixed-angle IR reflectance spectroscopy in both 1) a monolithic form with the crystal fixed in epoxy and polished, but data are also recorded from 2) pressed-pellet forms of powders of the same mineral using the same spectroscopic protocols. First results for the two methods are compared. It was found that the measured reflectance can vary by as much as a factor of two between the dolomite crystals and the pressed powder forms. For the two-sample preparation approaches the preliminary spectra are compared and the implications and limitations of each method for determining the optical constants of given materials are discussed. Comparison to literature data suggest that polarization effects likely account for the differing amplitude results for reflectance (and hence the k -vectors): Dolomite is known to be biaxial with significantly differing ordinary and extraordinary rays; the pressed pellet method measures an ensemble of microcrystals in randomly oriented positions, whereas the single crystal maintains just one orientation relative to the optical axis.

ellipsometry, dolomite, complex refractive index, ↗

Earth System Reanalysis in Support of Climate Model Improvements

Recent climate model developments, established through increased model resolution, have led to substantial improvements in model simulations of the time-evolving, coupled Earth system and its subcomponents. However, regardless of resolution, climate models will always produce climate features and variability that differ from the real world and will be prone to biases. This is due to many remaining uncertainties, such as in parametric and structural model uncertainty, in the initial conditions prescribed, and in the prescribed (scenario) forcing which varies on decadal to centennial timescales. Further model improvements are expected to arise specifically from improved representation of physical processes realized through model-data fusion. This will create an unprecedented opportunity to better exploit a large array of Earth observations, from in situ measurements to weather radars and satellite observations, as the resolved scales of the models approach those of the observations. For this, climate DA will be the central tool to bring models and observations into consistency, by improving initial conditions, inferring uncertain model parameters and structure, and quantifying uncertainty. Generally, there will be advantages and complementarities of adjoint-based smoother approaches, ensemble-based filter approaches, or new ML-inspired approaches. Yet, the ever-increasing model resolution will present growing challenges arising from computational cost, calling for new ways of performing data assimilation and model optimization. Using the complementarity in a hybrid approach, blending tools and concepts from variational, ensemble and ML methods might be what is required in the future. In this context ML could be important to handle non-linear responses, and to better approximate non-Gaussian distributions.

54 ENVIRONMENTAL SCIENCES↗

The Impact of Constrained Data Assimilation on the Forecasts of Three Convection Systems During the ARM MC3E Field Campaign

A constrained data assimilation (CDA) system based on the ensemble variational (EnVar) method and physical constraints of mass and water conservations is evaluated through three convective cases during the Midlatitude Continental Convective Clouds Experiment (MC3E) of the Atmospheric Radiation Measurement (ARM) program. Compared to the original data assimilation (ODA), the CDA is shown to perform better in the forecasted state variables and simulated precipitation. The CDA is also shown to greatly mitigate the loss of forecast skills in observation denial experiments when radar radial winds are withheld in the assimilation. Modifications to the algorithm and sensitivities of the CDA to the calculation of the time tendencies in the constraints are described.

54 ENVIRONMENTAL SCIENCES↗

Advanced Semi-Supervised Learning with Uncertainty Estimation for Phase Identification in Distribution Systems

The integration of advanced metering infrastructure (AMI) into power distribution networks generates valuable data for tasks such as phase identification; however, the limited and unreliable availability of labeled data in the form of customer phase connectivity presents challenges. To address this issue, we propose a semi-supervised learning (SSL) framework that effectively leverages labeled and unlabeled data. Our approach incorporates self-training, label spreading, and Bayesian neural networks (BNNs) to enhance phase identification with AMI data. Our method uses an ensemble of multilayer perceptron classifiers in a self-training setup, iteratively adding high-confidence pseudo-labels to improve robustness. We also apply label spread to propagate labels based on data similarity, which enhances generalization across diverse distributions. In addition, we employ a BNNs with uncertainty estimation, boosting confidence in predictions and reducing phase identification errors. In our case study, we achieved approximately 98% +/- 0.08 accuracy with uncertainty using minimal and unreliable labeled data from a real U.S. utility, Duquesne Light Company. Our SSL approach, combined with uncertainty estimation, provides an efficient solution for phase identification in AMI data, ultimately improving the reliability of smart grid applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Advanced Semi-Supervised Learning With Uncertainty Estimation for Phase Identification in Distribution Systems

The integration of advanced metering infrastructure (AMI) into power distribution networks generates valuable data for tasks such as phase identification; however, the limited and unreliable availability of labeled data in the form of customer phase connectivity presents challenges. To address this issue, we propose a semi-supervised learning (SSL) framework that effectively leverages labeled and unlabeled data. Our approach incorporates self-training, label spreading, and Bayesian neural networks (BNNs) to enhance phase identification with AMI data. Our method uses an ensemble of multilayer perceptron classifiers in a self-training setup, iteratively adding high-confidence pseudo-labels to improve robustness. We also apply label spread to propagate labels based on data similarity, which enhances generalization across diverse distributions. In addition, we employ a BNNs with uncertainty estimation, boosting confidence in predictions and reducing phase identification errors. In our case study, we achieved approximately 98% +/- 0.08 accuracy with uncertainty using minimal and unreliable labeled data from a real U.S. utility, Duquesne Light Company. Our SSL approach, combined with uncertainty estimation, provides an efficient solution for phase identification in AMI data, ultimately improving the reliability of smart grid applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Ensemble Monte Carlo characterization of graded Al(x)Ga(1-x)As heterojunction barriers

The current-voltage characteristics of graded Al(x)Ga(1-x)As heterojunction barriers were investigated using a self-consistent ensemble Monte Carlo method. Results are presented for barriers with two doping levels (10 to the 15th/cu cm and 10 to the 17th/cu cm) and two barrier heights (100 and 265 meV). It was found that the lower barrier structure exhibited little rectification at room temperature at both doping levels, while the higher barrier exhibited considerable rectification. The structures with the lower doping value exhibited a smaller current in both forward and reverse regions, due to space-charge effect. The results of studies of the energy and momentum distribution functions along the barrier indicate that the assumption of drifted Maxwellian distribution used in energy-momentum models is not justified for Gamma valley electrons.

Kamoua, R.↗

Aerosol Measurements of the Fine and Ultrafine Particle Content of Lunar Regolith

We report the first quantitative measurements of the ultrafine (20 to 100 nm) and fine (100 nm to 20 m) particulate components of Lunar surface regolith. The measurements were performed by gas-phase dispersal of the samples, and analysis using aerosol diagnostic techniques. This approach makes no a priori assumptions about the particle size distribution function as required by ensemble optical scattering methods, and is independent of refractive index and density. The method provides direct evaluation of effective transport diameters, in contrast to indirect scattering techniques or size information derived from two-dimensional projections of high magnification-images. The results demonstrate considerable populations in these size regimes. In light of the numerous difficulties attributed to dust exposure during the Apollo program, this outcome is of significant importance to the design of mitigation technologies for future Lunar exploration.

Greenberg, Paul S.↗