Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Data from: Coupled machine learning-ecosystem ensemble models substantially improve predictions of nitrous oxide (N 2 O) fluxes from US croplands

Nitrous oxide (N₂O) is a potent and persistent greenhouse gas, with rising atmospheric concentrations driven in part by inefficient use of synthetic nitrogen (N) fertilizers in agriculture. Predicting soil N₂O emissions is challenging due to high spatial and temporal variability arising from complex soil biogeochemical processes. Process-based ecosystem models and standalone machine learning (ML) approaches without extensive site-specific calibration often miss high emission episodes. Here, we show how an Ensemble Modeling System (EMS) based on outputs from an ensemble of ecosystem models coupled to an ensemble of ML models can improve predictions and understanding of N2O fluxes from US cropland. Trained and validated on approximately 12,000 N2O chamber measurements at 17 U.S. Midwest sites (six crops, 35 management practices), the EMS accurately predicted daily fluxes of N2O at both training (R² = 0.84, RMSE = 16.4 g N ha⁻¹ d⁻¹) and held-out testing sites (R² = 0.84, RMSE = 6.2 g N ha⁻¹ d⁻¹). Analyses identified six dominant N₂O drivers: soil organic carbon (SOC), NH₄⁺, NO₃⁻, water-filled pore space (WFPS), soil temperature, and biomass production. Wet, warm soils produced large N₂O peaks only with sufficient SOC and mineral N; in low-SOC soils, fluxes remained low. Incorporating these drivers into process-based models might significantly improve their predictive capacity. The EMS demonstrates a strong potential to predict N₂O fluxes at unseen sites, enabling more reliable regional inventories, improved gap-filling where measurements are sparse, and enhanced understanding of mechanisms to advance targeted mitigation strategies in food, feed, and bioenergy crops.

agricultural sciences↗

Sensitivity Analysis of Wind and Turbulence Predictions With Mesoscale‐Coupled Large Eddy Simulations Using Ensemble Machine Learning

Abstract Coupling between mesoscale models and large‐eddy simulation (LES) models is increasingly used to more realistically represent the wide range of scales of atmospheric motions affecting boundary layer winds and turbulence that need to be simulated accurately for applications such as wind energy. However, such mesoscale‐to‐microscale coupled modeling frameworks are potentially affected by a large number of uncertain closure parameters. Here, we investigate the sensitivity associated with six closure parameters related to a 1.5‐order subgrid‐scale turbulence closure for an ensemble of mesoscale‐coupled LES. The simulations are performed using the Weather Research and Forecasting model nested from horizontal resolutions of greater than a kilometer down to tens of meters. Closure parameters are varied to generate perturbed parameter ensembles for two case studies of highly sheared, convective boundary layers observed in the Columbia Basin of Oregon and Washington during the Second Wind Forecast Improvement Project. Machine learning algorithms are used to explore the sensitivity of LES predictions, considering the effects of the perturbed physical parameters alongside categorical factors such as the case study identity, measurement location, and LES resolution. For the conditions we examine, a single parameter, the eddy viscosity coefficient, is the dominant source of parametric sensitivity and its importance is comparable to the categorical factors for several of the simulation response variables we examine.

54 ENVIRONMENTAL SCIENCES↗

Indirect Tool Condition Monitoring Using Ensemble Machine Learning Techniques

Abstract Tool condition monitoring (TCM) has become a research area of interest due to its potential to significantly reduce manufacturing costs while increasing process visibility and efficiency. Machine learning (ML) is one analysis technique which has demonstrated advantages for TCM applications. However, the commonly studied individual ML models lack generalizability to new machining and environmental conditions, as well as robustness to the unbalanced datasets which are common in TCM. Ensemble ML models have demonstrated superior performance in other fields, but have only begun to be evaluated for TCM. As a result, it is not well understood how their TCM performance compares to that of individual models, or how homogeneous and heterogeneous ensemble models’ performances compare to one another. To fill in these research gaps, milling experiments were conducted using various cutting conditions, and the model groups were compared across several performance metrics. Statistical t-tests were also used to evaluate the significance of model performance differences. Through the analysis of four individual ML models and five ensemble models, all based on the processes’ sound, spindle power, and axial load signals, it was found that on average, the ensemble models performed better than the individual models, and that the homogeneous ensembles outperformed the heterogeneous ensembles.

Engineering↗

Bridging Hydrological Ensemble Simulation and Learning Using Deep Neural Operators

Ensemble-based simulation and learning (ESnL) has long been used in hydrology for parameter inference, but computational demands of process-based ESnL can be quite high. To address this issue, we propose a deep neural operator learning approach. Neural operators are generic machine learning algorithms that can learn functional mappings between infinite-dimensional spaces, providing a highly flexible tool for scientific machine learning. Our approach is built upon DeepONet, a specific deep neural operator, and is designed to address several common problems in hydrology, namely, model parameter estimation, prediction at ungaged locations, and uncertainty quantification. Here we demonstrate the effectiveness of our DeepONet-based workflow using an existing large model ensemble created for an eastern U.S. watershed that is instrumented with 10 streamflow gages. Results suggest DeepONet achieves high efficiency in learning an ML surrogate model from the model ensemble, with the modified Kling-Gupta Efficiency exceeding 0.9 on holdout test sets. Parameter inference, carried out using the trained DeepONet surrogate model and genetic algorithm, also yields robust results. Additionally, we formulate and train a separate DeepONet model for physics-informed, seq-to-seq streamflow forecasting, which further reduces biases in the pre-trained DeepONet surrogate model. While this study focuses primarily on a single watershed, our approach is general and may be extended to enable learning from model ensembles across multiple basins or models. Thus, this research represents a significant contribution to the application of hybrid machine learning in hydrology.

54 ENVIRONMENTAL SCIENCES↗

Collaborative Supervised Learning for Sensor Networks

Collaboration methods for distributed machine-learning algorithms involve the specification of communication protocols for the learners, which can query other learners and/or broadcast their findings preemptively. Each learner incorporates information from its neighbors into its own training set, and they are thereby able to bootstrap each other to higher performance. Each learner resides at a different node in the sensor network and makes observations (collects data) independently of the other learners. After being seeded with an initial labeled training set, each learner proceeds to learn in an iterative fashion. New data is collected and classified. The learner can then either broadcast its most confident classifications for use by other learners, or can query neighbors for their classifications of its least confident items. As such, collaborative learning combines elements of both passive (broadcast) and active (query) learning. It also uses ideas from ensemble learning to combine the multiple responses to a given query into a single useful label. This approach has been evaluated against current non-collaborative alternatives, including training a single classifier and deploying it at all nodes with no further learning possible, and permitting learners to learn from their own most confident judgments, absent interaction with their neighbors. On several data sets, it has been consistently found that active collaboration is the best strategy for a distributed learner network. The main advantages include the ability for learning to take place autonomously by collaboration rather than by requiring intervention from an oracle (usually human), and also the ability to learn in a distributed environment, permitting decisions to be made in situ and to yield faster response time.

Wagstaff, Kiri L.↗

Learning electric vehicle driver range anxiety with an initial state of charge-oriented gradient boosting approach

This manuscript focuses on the modeling of electric vehicle (EV) driver’s range anxiety, a fear that a vehicle does not have sufficient range, or state of charge (SOC) of the battery pack, to reach its destination and would strand its occupants. Despite numerous research studies on the modeling of charging behaviors, modeling efforts to understand at what battery percentages do EV drivers charge their vehicles, and what are the associated contributing factors, are rather limited. To this end, an ensemble learning model based on gradient boosting is developed. The model sequentially fits new predictors to new residuals of the previous prediction and, then, minimizes the loss when adding the latest prediction. A total of 18 features are defined and extracted from the multisource data, which cover information on driver, vehicles, stations, traffic conditions, as well as spatial-temporal context information of the charging events. The analyzed dataset includes 4.5-year’s charging event log data from 3,096 users and 468 public charging stations in Kansas City Missouri, and the macroscopic travel demand model maintained by the metropolitan planning organization. Here, the result shows the proposed model achieved a satisfactory result with a R square value of 0.54 and root mean square error of 0.14, both better than multiple linear regression model and random forest model. To reduce range anxiety, it is suggested that the priorities of deploying new charging facilities should be given to the areas with higher daily traffic prediction, with more conservative EV users or that are further from residential areas.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning for postprocessing ensemble streamflow forecasts

Skillful streamflow forecasts can inform decisions in various areas of water policy and management. We integrate numerical weather prediction ensembles, distributed hydrological model, and machine learning to generate ensemble streamflow forecasts at medium-range lead times (1–7 days). We demonstrate the application of machine learning as postprocessor for improving the quality of ensemble streamflow forecasts. Our results show that the machine learning postprocessor can improve streamflow forecasts relative to low-complexity forecasts (e.g., climatological and temporal persistence) as well as standalone hydrometeorological modeling and neural network. The relative gain in forecast skill from postprocessor is generally higher at medium-range timescales compared to shorter lead times; high flows compared to low–moderate flows, and the warm season compared to the cool ones. Overall, our results highlight the benefits of machine learning in many aspects for improving both the skill and reliability of streamflow forecasts.

54 ENVIRONMENTAL SCIENCES↗

High-throughput screening of tribological properties of monolayer films using molecular dynamics and machine learning

Monolayer films have shown promise as a lubricating layer to reduce friction and wear of mechanical devices with separations on the nanoscale. These films have a vast design space with many tunable properties that can affect their tribological effectiveness. For example, terminal group chemistry, film composition, and backbone chemistry can all lead to films with significantly different tribological properties. This design space, however, is very difficult to explore without a combinatorial approach and an automatable, reproducible, and extensible workflow to screen for promising candidate films. Here, using the Molecular Simulation Design Framework (MoSDeF), a combinatorial screening study was performed to explore 9747 unique monolayer films (116 964 total simulations) and a machine learning (ML) model using a random forest regressor, an ensemble learning technique, to explore the role of terminal group chemistry and its effect on tribological effectiveness. The most promising films were found to contain small terminal groups such as cyano and ethylene. The ML model was subsequently applied to screen terminal group candidates identified from the ChEMBL small molecule library. Approximately 193 131 unique film candidates were screened with approximately a five order of magnitude speed-up in analysis compared to simulation alone. The ML model was thus able to be used as a predictive tool to greatly speed up the initial screening of promising candidate films for future simulation studies, suggesting that computational screening in combination with ML can greatly increase the throughput in combinatorial approaches to generate in silico data and then train ML models in a controlled, self-consistent fashion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Inverse design of photonic surfaces via multi fidelity ensemble framework and femtosecond laser processing

We demonstrate a multi-fidelity (MF) machine learning ensemble framework for the inverse design of photonic surfaces, trained on a dataset of 11,759 samples that we fabricate using high throughput femtosecond laser processing. The MF ensemble combines an initial low fidelity model for generating design solutions, with a high fidelity model that refines these solutions through local optimization. The combined MF ensemble can generate multiple disparate sets of laser-processing parameters that can each produce the same target input spectral emissivity with high accuracy (root mean squared errors < 2%). SHapley Additive exPlanations analysis shows transparent model interpretability of the complex relationship between laser parameters and spectral emissivity. Finally, the MF ensemble is experimentally validated by fabricating and evaluating photonic surface designs that it generates for improved efficiency energy harvesting devices. Our approach provides a powerful tool for advancing the inverse design of photonic surfaces in energy harvesting applications.

97 MATHEMATICS AND COMPUTING↗

Assessing boundary condition and parametric uncertainty in numerical-weather-prediction-modeled, long-term offshore wind speed through machine learning and analog ensemble

To accurately plan and manage wind power plants, not only does the time-varying wind resource at the site of interest need to be assessed but also the uncertainty connected to this estimate. Numerical weather prediction (NWP) models at the mesoscale represent a valuable way to characterize the wind resource offshore, given the challenges connected with measuring hub-height wind speed. The boundary condition and parametric uncertainty associated with modeled wind speed is often estimated by running a model ensemble. However, creating an NWP ensemble of long-term wind resource data over a large region represents a computational challenge. Here, we propose two approaches to temporally extrapolate wind speed boundary condition and parametric uncertainty using a more convenient setup in which a mesoscale ensemble is run over a short-term period (1 year), and only a single model covers the desired long-term period (20 year). We quantify hub-height wind speed boundary condition and parametric uncertainty from the short-term model ensemble as its normalized across-ensemble standard deviation. Then, we develop and apply a gradient-boosting model and an analog ensemble approach to temporally extrapolate such uncertainty to the full 20-year period, for which only a single model run is available. As a test case, we consider offshore wind resource characterization in the California Outer Continental Shelf. Both of the proposed approaches provide accurate estimates of the long-term wind speed boundary condition and parametric uncertainty across the region (R 2 >0.75), with the gradient-boosting model slightly outperforming the analog ensemble in terms of bias and centered root-mean-square error. At the three offshore wind energy lease areas in the region, we find a long-term median hourly uncertainty between 10 % and 14 % of the mean hub-height wind speed values. Finally, we assess the physical variability in the uncertainty estimates. In general, we find that the wind speed uncertainty increases closer to land. Also, neutral conditions have smaller uncertainty than the stable and unstable cases, and the modeled wind speed in winter has less boundary condition and parametric sensitivity than summer.

17 WIND ENERGY↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

A Novel Machine Learning Method for Surface PM2.5 Estimations from Geostationary Satellites

Particulate matter (PM) with a diameter of less or equal to 2.5 μm, known as PM , affects human health as it penetrates the respiratory system. The Environmental Protection Agency (EPA) measures the atmospheric concentration of PM using air quality monitors stationed throughout the Continental United States (CONUS). Such measurements are points on a spatial domain and therefore, might not be representative of the air quality at nearby areas considering that the composition of the atmosphere is highly variable from place to place. Satellite based AOD permits a spatially uniform means of estimating PM and new geostationary satellites provide high temporal and spatial resolution estimation of AOD. However, the concentration of PM is non-linearly dependent on other atmospheric parameters that include relative humidity, temperature, and height of the planetary boundary layer. This information may be estimated at similar spatial and temporal resolutions as AOD from numerical modeling such as from the National Oceanic and Atmospheric Administration’s (NOAA) High Resolution Rapid Refresh (HRRR) model which resolves near real-time atmospheric conditions over the CONUS. The estimation of PM concentration is a multi-parametric problem that considers the effect of temporal dependencies among the different parameters. Deep learning approaches are appropriate for such complex estimation problems as they intrinsically capture relations among multiple non-linear parameters. This study compares deep-learning methods to traditional regression analysis to demonstrate the capabilities of these methods in predicting PM2.5 concentrations. Additionally, a novel ensemble learning approach is employed to identify scientific processes that could further improve the estimation of PM concentration. Utilizing Long Short-Term Memory (LSTM) neural networks, which are suitable for multivariate time series estimation problems as they are capable of learning long-term dependencies, individual models are created for each EPA station and trained on the aforementioned dataset collocated over each station. Individual station models are merged if the model's performance is improved by reducing the root mean squared error (RMSE) metric. This ensemble training method ultimately reduces the RMSE value. Evaluation of these results provide insights into physical processes and related observable parameters that may contribute to PM concentrations. Identified parameters evaluated to be statistically different between the merged and unmerged models are expected to improve overall performance. These new parameters are then utilized for reevaluation of the deep learning methods with an extreme gradient boosting model with an RMSE of 5.5 providing the best results.

George Priftis↗

Interpretable Tree-Based and Graph Neural Network Approaches for Novel Solid State Electrolyte Design

All-solid-state batteries with Li metal anode can address the safety issues surrounding traditional Li-ion batteries as well as the demand for higher energy densities. However, the development of solid electrolytes simultaneously possessing high ionic conductivity and good chemical and electrochemical stabilities has proven to be a challenge. I will present our informatics approach to explore the Li compound space for promising solid electrolytes using high-throughput multi-property screening and interpretable machine learning. This is accomplished through the generation of a large database of battery-related materials properties of Li compounds. We use tree-based ensemble learning methods and graph neural network approaches to accurately learn relationships between crystal structures and corresponding thermodynamic and kinetic properties, with interpretability being a major focus. Our models give us the ability to enable rapid discovery and design of novel solid-state battery chemistries.

Materials discovery↗

Interpretable ML Approaches for Novel Solid State Electrolyte Design

All-solid-state batteries with Li metal anode can address the safety issues surrounding traditional Li-ion batteries as well as the demand for higher energy densities. However, the development of solid electrolytes simultaneously possessing high ionic conductivity and good chemical and electrochemical stabilities has proven to be a challenge. I will present our informatics approach to explore the Li compound space for promising solid electrolytes using high-throughput multi-property screening and interpretable machine learning. This is accomplished through the generation of a large database of battery-related materials properties of Li compounds. We use tree-based ensemble learning methods and graph neural network approaches to accurately learn relationships between crystal structures and corresponding thermodynamic and kinetic properties, with interpretability being a major focus. Our models give us the ability to enable rapid discovery and design of novel solid-state battery chemistries.

Shreyas J Honrao↗

Using Multimodal Input for Autonomous Decision Making for Unmanned Systems

Autonomous decision making in the presence of uncertainly is a deeply studied problem space particularly in the area of autonomous systems operations for land, air, sea, and space vehicles. Various techniques ranging from single algorithm solutions to complex ensemble classifier systems have been utilized in a research context in solving mission critical flight decisions. Realized systems on actual autonomous hardware, however, is a difficult systems integration problem, constituting a majority of applied robotics development timelines. The ability to reliably and repeatedly classify objects during a vehicles mission execution is vital for the vehicle to mitigate both static and dynamic environmental concerns such that the mission may be completed successfully and have the vehicle operate and return safely. In this paper, the Autonomy Incubator proposes and discusses an ensemble learning and recognition system planned for our autonomous framework, AEON, in selected domains, which fuse decision criteria, using prior experience on both the individual classifier layer and the ensemble layer to mitigate environmental uncertainty during operation.

Neilan, James H.↗

Adaptive Fault Detection on Liquid Propulsion Systems with Virtual Sensors: Algorithms and Architectures

Prior to the launch of STS-119 NASA had completed a study of an issue in the flow control valve (FCV) in the Main Propulsion System of the Space Shuttle using an adaptive learning method known as Virtual Sensors. Virtual Sensors are a class of algorithms that estimate the value of a time series given other potentially nonlinearly correlated sensor readings. In the case presented here, the Virtual Sensors algorithm is based on an ensemble learning approach and takes sensor readings and control signals as input to estimate the pressure in a subsystem of the Main Propulsion System. Our results indicate that this method can detect faults in the FCV at the time when they occur. We use the standard deviation of the predictions of the ensemble as a measure of uncertainty in the estimate. This uncertainty estimate was crucial to understanding the nature and magnitude of transient characteristics during startup of the engine. This paper overviews the Virtual Sensors algorithm and discusses results on a comprehensive set of Shuttle missions and also discusses the architecture necessary for deploying such algorithms in a real-time, closed-loop system or a human-in-the-loop monitoring system. These results were presented at a Flight Readiness Review of the Space Shuttle in early 2009.

Matthews, Bryan L.↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗