SEARCH · Engineering Papers
Results for “ensemble methods”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
pnnl/SNAP
In this work, we detail two uncertainty quantification (UQ) methods that provide complementary information. Readout ensembling, by finetuning only the readout layers of an ensemble of foundation models, provides information about model uncertainty. Amending the final readout layer to predict upper and lower quantiles replaces point predictions with distributional predictions, which provide information about uncertainty within the underlying training data. We demonstrate our approach with the MACE-MP-0 model, applying UQ to both the foundation model and a series of finetuned models. The uncertainties produced by the ensemble and quantile methods are demonstrated to be distinct measures by which the quality of the NNP output can be judged.
The Atmospheric Carbon and Transport (ACT)-America Mission
The Atmospheric Carbon and Transport (ACT)-America NASA Earth Venture Suborbital Mission set out to improve regional atmospheric greenhouse gas (GHG) inversions by exploring the intersection of the strong GHG fluxes and vigorous atmospheric transport that occurs within the midlatitudes. In this study, two research aircraft instrumented with remote and in situ sensors to measure GHG mole fractions, associated trace gases, and atmospheric state variables collected 1,140.7 flight hours of research data, distributed across 305 individual aircraft sorties, coordinated within 121 research flight days, and spanning five 6-week seasonal flight campaigns in the central and eastern United States. Flights sampled 31 synoptic sequences, including fair-weather and frontal conditions, at altitudes ranging from the atmospheric boundary layer to the upper free troposphere. The observations were complemented with global and regional GHG flux and transport model ensembles. We found that midlatitude weather systems contain large spatial gradients in GHG mole fractions, in patterns that were consistent as a function of season and altitude. We attribute these patterns to a combination of regional terrestrial fluxes and inflow from the continental boundaries. These observations, when segregated according to altitude and air mass, provide a variety of quantitative insights into the realism of regional CO 2 and CH 4 fluxes and atmospheric GHG transport realizations. The ACT-America dataset and ensemble modeling methods provide benchmarks for the development of atmospheric inversion systems. As global and regional atmospheric inversions incorporate ACT-America’s findings and methods, we anticipate these systems will produce increasingly accurate and precise subcontinental GHG flux estimates.
Semantic segmentation of rock images and ensemble approach for deep learning methods.
Abstract not provided.
Modeling the effects of high-G stress on pilots in a tracking task
Air-to-air tracking experiments were conducted at the Aerospace Medical Research Laboratories using both fixed and moving base dynamic environment simulators. The obtained data, which includes longitudinal error of a simulated air-to-air tracking task as well as other auxiliary variables, was analyzed using an ensemble averaging method. In conjunction with these experiments, the optimal control model is applied to model a human operator under high-G stress.
Online Bagging and Boosting
Bagging and boosting are two of the most well-known ensemble learning methods due to their theoretical performance guarantees and strong experimental results. However, these algorithms have been used mainly in batch mode, i.e., they require the entire training set to be available at once and, in some cases, require random access to the data. In this paper, we present online versions of bagging and boosting that require only one pass through the training data. We build on previously presented work by presenting some theoretical results. We also compare the online and batch algorithms experimentally in terms of accuracy and running time.
Tracking Energy Flow Using a Volumetric Acoustic Intensity Imager (VAIM)
A new measurement device has been invented at the Naval Research Laboratory which images instantaneously the intensity vector throughout a three-dimensional volume nearly a meter on a side. The measurement device consists of a nearly transparent spherical array of 50 inexpensive microphones optimally positioned on an imaginary spherical surface of radius 0.2m. Front-end signal processing uses coherence analysis to produce multiple, phase-coherent holograms in the frequency domain each related to references located on suspect sound sources in an aircraft cabin. The analysis uses either SVD or Cholesky decomposition methods using ensemble averages of the cross-spectral density with the fixed references. The holograms are mathematically processed using spherical NAH (nearfield acoustical holography) to convert the measured pressure field into a vector intensity field in the volume of maximum radius 0.4 m centered on the sphere origin. The utility of this probe is evaluated in a detailed analysis of a recent in-flight experiment in cooperation with Boeing and NASA on NASA s Aries 757 aircraft. In this experiment the trim panels and insulation were removed over a section of the aircraft and the bare panels and windows were instrumented with accelerometers to use as references for the VAIM. Results show excellent success at locating and identifying the sources of interior noise in-flight in the frequency range of 0 to 1400 Hz. This work was supported by NASA and the Office of Naval Research.
Interpretable Tree-Based and Graph Neural Network Approaches for Novel Solid State Electrolyte Design
All-solid-state batteries with Li metal anode can address the safety issues surrounding traditional Li-ion batteries as well as the demand for higher energy densities. However, the development of solid electrolytes simultaneously possessing high ionic conductivity and good chemical and electrochemical stabilities has proven to be a challenge. I will present our informatics approach to explore the Li compound space for promising solid electrolytes using high-throughput multi-property screening and interpretable machine learning. This is accomplished through the generation of a large database of battery-related materials properties of Li compounds. We use tree-based ensemble learning methods and graph neural network approaches to accurately learn relationships between crystal structures and corresponding thermodynamic and kinetic properties, with interpretability being a major focus. Our models give us the ability to enable rapid discovery and design of novel solid-state battery chemistries.
Interpretable ML Approaches for Novel Solid State Electrolyte Design
All-solid-state batteries with Li metal anode can address the safety issues surrounding traditional Li-ion batteries as well as the demand for higher energy densities. However, the development of solid electrolytes simultaneously possessing high ionic conductivity and good chemical and electrochemical stabilities has proven to be a challenge. I will present our informatics approach to explore the Li compound space for promising solid electrolytes using high-throughput multi-property screening and interpretable machine learning. This is accomplished through the generation of a large database of battery-related materials properties of Li compounds. We use tree-based ensemble learning methods and graph neural network approaches to accurately learn relationships between crystal structures and corresponding thermodynamic and kinetic properties, with interpretability being a major focus. Our models give us the ability to enable rapid discovery and design of novel solid-state battery chemistries.
Fast HARDI Uncertainty Quantification and Visualization with Spherical Sampling
In this paper, we study uncertainty quantification and visualization of orientation distribution functions (ODF), which corresponds to the diffusion profile of high angular resolution diffusion imaging (HARDI) data. The shape inclusion probability (SIP) function is the state‐of‐the‐art method for capturing the uncertainty of ODF ensembles. The current method of computing the SIP function with a volumetric basis exhibits high computational and memory costs, which can be a bottleneck to integrating uncertainty into HARDI visualization techniques and tools. We propose a novel spherical sampling framework for faster computation of the SIP function with lower memory usage and increased accuracy. In particular, we propose direct extraction of SIP isosurfaces, which represent confidence intervals indicating spatial uncertainty of HARDI glyphs, by performing spherical sampling of ODFs. Our spherical sampling approach requires much less sampling than the state‐of‐the‐art volume sampling method, thus providing significantly enhanced performance, scalability, and the ability to perform implicit ray tracing. Our experiments demonstrate that the SIP isosurfaces extracted with our spherical sampling approach can achieve up to 8164× speedup, 37282× memory reduction, and 50.2% less SIP isosurface error compared to the classical volume sampling approach. We demonstrate the efficacy of our methods through experiments on synthetic and human‐brain HARDI datasets.
Development of Super Ensemble-Based Aviation Turbulence Guidance (SEATG) for Air Traffic Management
A new method for forecasting turbulence is developed and evaluated using the high resolution weather model and in situ turbulence observations from commercial aircraft. The new method is an ensemble of various turbulence metrics from multiple time-lagged ensemble forecasts created using a sequence of four procedures. These include weather modeling, calculation of turbulence metrics, mapping the metrics into a common turbulence-scale, and production of final forecast. The new method uses similar methodology as current operational turbulence forecast with three improvements. First, it uses a higher resolution ((delta)x = 3 km) weather model to capture cloud resolving scale phenomena. Second, it computes the metrics for multiple forecasts that are combined at the same valid time resulting in a time-lagged ensemble of multiple turbulence metrics. Finally, it provides both deterministic and probabilistic turbulence forecasts. Results show the new forecasts match well with observed radar reflectivity along a surface front as well as convectively induced turbulence outside the clouds on research period. Overall performance skill of the new turbulence forecast compared with the observed EDR data during the research period is superior to any single turbulence metric. The probabilistic turbulence forecast is used in an example air traffic management application for creating a wind-optimal route considering turbulence information. The wind-optimal route passing through areas of 50% potential for moderate-or-greater turbulence and the lateral turbulence avoidance routes starting from three different waypoints along the wind-optimal route from Los Angeles international airport to John F. Kennedy international airport are calculated using different turbulence forecasts. This example shows additional flight time is required to avoid potential turbulence encounters.
The 4DEnVar-based weakly coupled land data assimilation system for E3SM version 2
Abstract. A new weakly coupled land data assimilation (WCLDA) system based on the four-dimensional ensemble variational (4DEnVar) method is developed and applied to the fully coupled Energy Exascale Earth System Model version 2 (E3SMv2). The dimension-reduced projection four-dimensional variational (DRP-4DVar) method is employed to implement 4DVar using the ensemble technique instead of the adjoint technique. With an interest in providing initial conditions for decadal climate predictions, monthly mean anomalies of soil moisture and temperature from the Global Land Data Assimilation System (GLDAS) reanalysis from 1980 to 2016 are assimilated into the land component of E3SMv2 within the coupled modeling framework with a 1-month assimilation window. The coupled assimilation experiment is evaluated using multiple metrics, including the cost function, assimilation efficiency index, correlation, root-mean-square error (RMSE), and bias, and compared with a control simulation without land data assimilation. The WCLDA system yields improved simulation of soil moisture and temperature compared with the control simulation, with improvements found throughout the soil layers and in many regions of the global land. In terms of both soil moisture and temperature, the assimilation experiment outperforms the control simulation with reduced RMSE and higher temporal correlation in many regions, especially in South America, central Africa, Australia, and large parts of Eurasia. Furthermore, significant improvements are also found in reproducing the time evolution of the 2012 US Midwest drought, highlighting the crucial role of land surface in drought lifecycle. The WCLDA system is intended to be a foundational resource for research to investigate land-derived climate predictability.
A Pattern-Recognition-Based Ensemble Data Imputation Framework for Sensors from Building Energy Systems
Building operation data are important for monitoring, analysis, modeling, and control of building energy systems. However, missing data is one of the major data quality issues, making data imputation techniques become increasingly important. There are two key research gaps for missing sensor data imputation in buildings: the lack of customized and automated imputation methodology, and the difficulty of the validation of data imputation methods. In this paper, a framework is developed to address these two gaps. First, a validation data generation module is developed based on pattern recognition to create a validation dataset to quantify the performance of data imputation methods. Second, a pool of data imputation methods is tested under the validation dataset to find an optimal single imputation method for each sensor, which is termed as an ensemble method. The method can reflect the specific mechanism and randomness of missing data from each sensor. The effectiveness of the framework is demonstrated by 18 sensors from a real campus building. The overall accuracy of data imputation for those sensors improves by 18.2% on average compared with the best single data imputation method.
Study of Overfitting by Machine Learning Methods Using Generalization Equations
The training error of Machine Learning (ML) methods has been extensively used for performance assessment, and its low values have been used as a main justification for complex methods such as estimator fusion and ensembles, and hyper parameter tuning. We present two practical cases where independent tests indicate that the low training error is more of a reflection of over-fitting rather than the generalization ability. We derive a generic form of the generalization equations that separates the training error terms of ML methods from their epistemic terms that correspond to approximation and learnability properties. It provides a framework to separately account for both terms to ensure an overall high generalization performance. For regression estimation tasks, we derive conditions for performance enhancements achieved by hyper parameter tuning, and fusion and ensemble methods over their constituent methods. We present experimental measurements and ML estimates that illustrate the analytical results for the throughput profile estimation of a data transport infrastructure.
A Gridded Solar Irradiance Ensemble Prediction System Based on WRF-Solar EPS and the Analog Ensemble
The WRF-Solar Ensemble Prediction System (WRF-Solar EPS) and a calibration method, the analog ensemble (AnEn), are used to generate calibrated gridded ensemble forecasts of solar irradiance over the contiguous United States (CONUS). Global horizontal irradiance (GHI) and direct normal irradiance (DNI) retrievals, based on geostationary satellites from the National Solar Radiation Database (NSRDB) are used for both calibrating and verifying the day-ahead GHI and DNI predictions (GDIP). A 10-member ensemble of WRF-Solar EPS is run in a re-forecast mode to generate day-ahead GDIP for three years. The AnEn is used to calibrate GDIP at each grid point independently using the NSRDB as the “ground truth”. Performance evaluations of deterministic and probabilistic attributes are carried out over the whole CONUS. The results demonstrate that using the AnEn calibrated ensemble forecast from WRF-Solar EPS contributes to improving the overall quality of the GHI predictions with respect to an AnEn calibrated system based only on the deterministic run of WRF-Solar. In fact, the calibrated WRF-Solar EPS’s mean exhibits a lower bias and RMSE than the calibrated deterministic WRF-Solar. Moreover, using the ensemble mean and spread as predictors for the AnEn allows a more effective calibration than using variables only from the deterministic runs. Finally, it has been shown that the recently introduced algorithm of correction for rare events is of paramount importance to obtain the lowest values of GHI from the calibrated ensemble (WRF-Solar EPS AnEn), qualitatively consistent with those observed from the NSRDB.
Ensemble cure kinetics network (ECK-Net): A method to derive cure kinetics of thermosetting resin
This paper introduces an Ensemble Cure Kinetics Network (ECK-Net), a neural network (NN)–based framework for modeling the cure kinetics of thermosetting resins within a phenomenological context. ECK-Net replaces traditional analytic models, which require extensive chemical insight and multiple isothermal/non-isothermal experiments, with a data-driven surrogate that maps nonlinear relationships between temperature, degree of cure, and reaction rate from differential scanning calorimetry data. The proposed approach predicts input-dependent kinetic coefficients of a generalized nth-order reaction equation rather than reaction rates directly, enabling a single unified model to represent various epoxy systems without relying on iso-conversional analysis or predefined functional forms. To ensure robustness, multiple independently trained networks under different random initializations are blended through an ensemble strategy, effectively mitigating the stochastic variability inherent to neural networks. The framework is validated using experimental datasets from multiple resin systems, including aerospace-grade materials (Toray 3900-2, Cycom 5320-1, and Hexcel 8552) and a windmill-grade resin (RIMR 035c). The model accurately reproduces the temporal evolution of the degree of cure under manufacturers’ recommended cure cycles across all tested resins systems, yielding Pearson’s correlation coefficients of 0.992, 0.994, 0.993, 0.997, respectively. To demonstrate process-level applicability, the trained network was implemented within the Abaqus environment to simulate out-of-autoclave (OOA) curing process of the CFRP panel composed of Toray T830H-6K/3900-2D prepreg. The simulation results showed excellent agreement with experimental temperature response (maximum peak temperature, simulation: 189.6 °C, experiment: 188.5 °C) and the final degree of cure (simulation: 0.948, experiment: 0.960 ± 0.013), confirming ECK-Net’s capability as a reliable alternative to conventional cure kinetics modeling methods.
Improving the Representation of Land Surface Processes Using the Data Assimilation Research Testbed (DART)
The land surface is a critical part of the earth system as processes related to water, carbon, energy and nitrogen cycling have important implications for climate forcing, air quality, water availability and seasonal atmospheric forecasting. Despite advances in land surface modeling, land surface model performance is often limited because of errors related to initial and boundary conditions, model structure, and parameters. Data assimilation (DA) techniques combined with an expanding network of earth system observations present an opportunity to reduce these errors and improve simulations. Here, we emphasize the implementation of tools and approaches to overcome challenges related to land DA to constrain carbon and water cycling. In particular, we discuss the implementation of adaptive inflation to modify ensemble spread in response to time-varying networks of gridded observations. We also discuss methods to generate ensemble spread through boundary condition (meteorology) forcing that can be applied to site-level applications. Next, we describe the application of vertical localization upon surface soil moisture observations, and forward operators specifically designed for the assimilation of snow and solar-induced fluorescence observations. Finally, we discuss the potential benefit of a quantile conserving filter used to update bounded quantities (state or parameter values).
Ensemble Spread Behavior in Coupled Climate Models: Insights From the Energy Exascale Earth System Model Version 1 Large Ensemble
AbstractAssessing uncertainty in future climate projections requires understanding both internal climate variability and external forcing. For this reason, single‐model initial condition large ensembles (SMILEs) run with Earth System Models (ESMs) have recently become popular. Here we present a new 20‐member SMILE with the Energy Exascale Earth System Model version 1 (E3SMv1‐LE), which uses a “macro” initialization strategy choosing coupled atmosphere/ocean states based on inter‐basin contrasts in ocean heat content (OHC). The E3SMv1‐LE simulates tropical climate variability well, albeit with a muted warming trend over the twentieth century due to overly strong aerosol forcing. The E3SMv1‐LE's initial climate spread is comparable to other (larger) SMILEs, suggesting that maximizing inter‐basin ocean heat contrasts may be an efficient method of generating ensemble spread. We also compare different ensemble spread across multiple SMILEs, using surface air temperature and OHC. The Community Earth system Model version 1, the only ensemble which utilizes a “micro” initialization approach perturbing only atmospheric initial conditions, yields lower spread in the first ∼30 years. The E3SMv1‐LE exhibits a relatively large spread, with some evidence for anthropogenic forcing influencing spread in the late twentieth century. However, systematic effects of differing “macro” initialization strategies are difficult to detect, possibly resulting from differing model physics or responses to external forcing. Notably, the method of standardizing results affects ensemble spread: control simulations for most models have either large background trends or multi‐centennial variability in OHC. This spurious disequlibrium behavior is a substantial roadblock to understanding both internal climate variability and its response to forcing.