Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Predicting peak day and peak hour of electricity demand with ensemble machine learning

Battery energy storage systems can be used for peak demand reduction in power systems, leading to significant economic benefits. Two practical challenges are 1) accurately determining the peak load days and hours and 2) quantifying and reducing uncertainties associated with the forecast in probabilistic risk measures for dispatch decision-making. In this study, we develop a supervised machine learning approach to generate 1) the probability of the next operation day containing the peak hour of the month and 2) the probability of an hour to be the peak hour of the day. Guidance is provided on preparation and augmentation of data as well as selection of machine learning models and decision-making thresholds. The proposed approach is applied to the Duke Energy Progress system and successfully captures 69 peak days out of 72 testing months with a 3% exceedance probability threshold. On 90% of the peak days, the actual peak hour is among the 2 h with the highest probabilities.

25 ENERGY STORAGE↗

Data from: Coupled machine learning-ecosystem ensemble models substantially improve predictions of nitrous oxide (N 2 O) fluxes from US croplands

Nitrous oxide (N₂O) is a potent and persistent greenhouse gas, with rising atmospheric concentrations driven in part by inefficient use of synthetic nitrogen (N) fertilizers in agriculture. Predicting soil N₂O emissions is challenging due to high spatial and temporal variability arising from complex soil biogeochemical processes. Process-based ecosystem models and standalone machine learning (ML) approaches without extensive site-specific calibration often miss high emission episodes. Here, we show how an Ensemble Modeling System (EMS) based on outputs from an ensemble of ecosystem models coupled to an ensemble of ML models can improve predictions and understanding of N2O fluxes from US cropland. Trained and validated on approximately 12,000 N2O chamber measurements at 17 U.S. Midwest sites (six crops, 35 management practices), the EMS accurately predicted daily fluxes of N2O at both training (R² = 0.84, RMSE = 16.4 g N ha⁻¹ d⁻¹) and held-out testing sites (R² = 0.84, RMSE = 6.2 g N ha⁻¹ d⁻¹). Analyses identified six dominant N₂O drivers: soil organic carbon (SOC), NH₄⁺, NO₃⁻, water-filled pore space (WFPS), soil temperature, and biomass production. Wet, warm soils produced large N₂O peaks only with sufficient SOC and mineral N; in low-SOC soils, fluxes remained low. Incorporating these drivers into process-based models might significantly improve their predictive capacity. The EMS demonstrates a strong potential to predict N₂O fluxes at unseen sites, enabling more reliable regional inventories, improved gap-filling where measurements are sparse, and enhanced understanding of mechanisms to advance targeted mitigation strategies in food, feed, and bioenergy crops.

agricultural sciences↗

Sensitivity Analysis of Wind and Turbulence Predictions With Mesoscale‐Coupled Large Eddy Simulations Using Ensemble Machine Learning

Abstract Coupling between mesoscale models and large‐eddy simulation (LES) models is increasingly used to more realistically represent the wide range of scales of atmospheric motions affecting boundary layer winds and turbulence that need to be simulated accurately for applications such as wind energy. However, such mesoscale‐to‐microscale coupled modeling frameworks are potentially affected by a large number of uncertain closure parameters. Here, we investigate the sensitivity associated with six closure parameters related to a 1.5‐order subgrid‐scale turbulence closure for an ensemble of mesoscale‐coupled LES. The simulations are performed using the Weather Research and Forecasting model nested from horizontal resolutions of greater than a kilometer down to tens of meters. Closure parameters are varied to generate perturbed parameter ensembles for two case studies of highly sheared, convective boundary layers observed in the Columbia Basin of Oregon and Washington during the Second Wind Forecast Improvement Project. Machine learning algorithms are used to explore the sensitivity of LES predictions, considering the effects of the perturbed physical parameters alongside categorical factors such as the case study identity, measurement location, and LES resolution. For the conditions we examine, a single parameter, the eddy viscosity coefficient, is the dominant source of parametric sensitivity and its importance is comparable to the categorical factors for several of the simulation response variables we examine.

54 ENVIRONMENTAL SCIENCES↗

Indirect Tool Condition Monitoring Using Ensemble Machine Learning Techniques

Abstract Tool condition monitoring (TCM) has become a research area of interest due to its potential to significantly reduce manufacturing costs while increasing process visibility and efficiency. Machine learning (ML) is one analysis technique which has demonstrated advantages for TCM applications. However, the commonly studied individual ML models lack generalizability to new machining and environmental conditions, as well as robustness to the unbalanced datasets which are common in TCM. Ensemble ML models have demonstrated superior performance in other fields, but have only begun to be evaluated for TCM. As a result, it is not well understood how their TCM performance compares to that of individual models, or how homogeneous and heterogeneous ensemble models’ performances compare to one another. To fill in these research gaps, milling experiments were conducted using various cutting conditions, and the model groups were compared across several performance metrics. Statistical t-tests were also used to evaluate the significance of model performance differences. Through the analysis of four individual ML models and five ensemble models, all based on the processes’ sound, spindle power, and axial load signals, it was found that on average, the ensemble models performed better than the individual models, and that the homogeneous ensembles outperformed the heterogeneous ensembles.

Engineering↗

Bridging Hydrological Ensemble Simulation and Learning Using Deep Neural Operators

Ensemble-based simulation and learning (ESnL) has long been used in hydrology for parameter inference, but computational demands of process-based ESnL can be quite high. To address this issue, we propose a deep neural operator learning approach. Neural operators are generic machine learning algorithms that can learn functional mappings between infinite-dimensional spaces, providing a highly flexible tool for scientific machine learning. Our approach is built upon DeepONet, a specific deep neural operator, and is designed to address several common problems in hydrology, namely, model parameter estimation, prediction at ungaged locations, and uncertainty quantification. Here we demonstrate the effectiveness of our DeepONet-based workflow using an existing large model ensemble created for an eastern U.S. watershed that is instrumented with 10 streamflow gages. Results suggest DeepONet achieves high efficiency in learning an ML surrogate model from the model ensemble, with the modified Kling-Gupta Efficiency exceeding 0.9 on holdout test sets. Parameter inference, carried out using the trained DeepONet surrogate model and genetic algorithm, also yields robust results. Additionally, we formulate and train a separate DeepONet model for physics-informed, seq-to-seq streamflow forecasting, which further reduces biases in the pre-trained DeepONet surrogate model. While this study focuses primarily on a single watershed, our approach is general and may be extended to enable learning from model ensembles across multiple basins or models. Thus, this research represents a significant contribution to the application of hybrid machine learning in hydrology.

54 ENVIRONMENTAL SCIENCES↗

Learning electric vehicle driver range anxiety with an initial state of charge-oriented gradient boosting approach

This manuscript focuses on the modeling of electric vehicle (EV) driver’s range anxiety, a fear that a vehicle does not have sufficient range, or state of charge (SOC) of the battery pack, to reach its destination and would strand its occupants. Despite numerous research studies on the modeling of charging behaviors, modeling efforts to understand at what battery percentages do EV drivers charge their vehicles, and what are the associated contributing factors, are rather limited. To this end, an ensemble learning model based on gradient boosting is developed. The model sequentially fits new predictors to new residuals of the previous prediction and, then, minimizes the loss when adding the latest prediction. A total of 18 features are defined and extracted from the multisource data, which cover information on driver, vehicles, stations, traffic conditions, as well as spatial-temporal context information of the charging events. The analyzed dataset includes 4.5-year’s charging event log data from 3,096 users and 468 public charging stations in Kansas City Missouri, and the macroscopic travel demand model maintained by the metropolitan planning organization. Here, the result shows the proposed model achieved a satisfactory result with a R square value of 0.54 and root mean square error of 0.14, both better than multiple linear regression model and random forest model. To reduce range anxiety, it is suggested that the priorities of deploying new charging facilities should be given to the areas with higher daily traffic prediction, with more conservative EV users or that are further from residential areas.

33 ADVANCED PROPULSION SYSTEMS↗

Machine learning for postprocessing ensemble streamflow forecasts

Skillful streamflow forecasts can inform decisions in various areas of water policy and management. We integrate numerical weather prediction ensembles, distributed hydrological model, and machine learning to generate ensemble streamflow forecasts at medium-range lead times (1–7 days). We demonstrate the application of machine learning as postprocessor for improving the quality of ensemble streamflow forecasts. Our results show that the machine learning postprocessor can improve streamflow forecasts relative to low-complexity forecasts (e.g., climatological and temporal persistence) as well as standalone hydrometeorological modeling and neural network. The relative gain in forecast skill from postprocessor is generally higher at medium-range timescales compared to shorter lead times; high flows compared to low–moderate flows, and the warm season compared to the cool ones. Overall, our results highlight the benefits of machine learning in many aspects for improving both the skill and reliability of streamflow forecasts.

54 ENVIRONMENTAL SCIENCES↗

High-throughput screening of tribological properties of monolayer films using molecular dynamics and machine learning

Monolayer films have shown promise as a lubricating layer to reduce friction and wear of mechanical devices with separations on the nanoscale. These films have a vast design space with many tunable properties that can affect their tribological effectiveness. For example, terminal group chemistry, film composition, and backbone chemistry can all lead to films with significantly different tribological properties. This design space, however, is very difficult to explore without a combinatorial approach and an automatable, reproducible, and extensible workflow to screen for promising candidate films. Here, using the Molecular Simulation Design Framework (MoSDeF), a combinatorial screening study was performed to explore 9747 unique monolayer films (116 964 total simulations) and a machine learning (ML) model using a random forest regressor, an ensemble learning technique, to explore the role of terminal group chemistry and its effect on tribological effectiveness. The most promising films were found to contain small terminal groups such as cyano and ethylene. The ML model was subsequently applied to screen terminal group candidates identified from the ChEMBL small molecule library. Approximately 193 131 unique film candidates were screened with approximately a five order of magnitude speed-up in analysis compared to simulation alone. The ML model was thus able to be used as a predictive tool to greatly speed up the initial screening of promising candidate films for future simulation studies, suggesting that computational screening in combination with ML can greatly increase the throughput in combinatorial approaches to generate in silico data and then train ML models in a controlled, self-consistent fashion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Inverse design of photonic surfaces via multi fidelity ensemble framework and femtosecond laser processing

We demonstrate a multi-fidelity (MF) machine learning ensemble framework for the inverse design of photonic surfaces, trained on a dataset of 11,759 samples that we fabricate using high throughput femtosecond laser processing. The MF ensemble combines an initial low fidelity model for generating design solutions, with a high fidelity model that refines these solutions through local optimization. The combined MF ensemble can generate multiple disparate sets of laser-processing parameters that can each produce the same target input spectral emissivity with high accuracy (root mean squared errors < 2%). SHapley Additive exPlanations analysis shows transparent model interpretability of the complex relationship between laser parameters and spectral emissivity. Finally, the MF ensemble is experimentally validated by fabricating and evaluating photonic surface designs that it generates for improved efficiency energy harvesting devices. Our approach provides a powerful tool for advancing the inverse design of photonic surfaces in energy harvesting applications.

97 MATHEMATICS AND COMPUTING↗

Assessing boundary condition and parametric uncertainty in numerical-weather-prediction-modeled, long-term offshore wind speed through machine learning and analog ensemble

To accurately plan and manage wind power plants, not only does the time-varying wind resource at the site of interest need to be assessed but also the uncertainty connected to this estimate. Numerical weather prediction (NWP) models at the mesoscale represent a valuable way to characterize the wind resource offshore, given the challenges connected with measuring hub-height wind speed. The boundary condition and parametric uncertainty associated with modeled wind speed is often estimated by running a model ensemble. However, creating an NWP ensemble of long-term wind resource data over a large region represents a computational challenge. Here, we propose two approaches to temporally extrapolate wind speed boundary condition and parametric uncertainty using a more convenient setup in which a mesoscale ensemble is run over a short-term period (1 year), and only a single model covers the desired long-term period (20 year). We quantify hub-height wind speed boundary condition and parametric uncertainty from the short-term model ensemble as its normalized across-ensemble standard deviation. Then, we develop and apply a gradient-boosting model and an analog ensemble approach to temporally extrapolate such uncertainty to the full 20-year period, for which only a single model run is available. As a test case, we consider offshore wind resource characterization in the California Outer Continental Shelf. Both of the proposed approaches provide accurate estimates of the long-term wind speed boundary condition and parametric uncertainty across the region (R 2 >0.75), with the gradient-boosting model slightly outperforming the analog ensemble in terms of bias and centered root-mean-square error. At the three offshore wind energy lease areas in the region, we find a long-term median hourly uncertainty between 10 % and 14 % of the mean hub-height wind speed values. Finally, we assess the physical variability in the uncertainty estimates. In general, we find that the wind speed uncertainty increases closer to land. Also, neutral conditions have smaller uncertainty than the stable and unstable cases, and the modeled wind speed in winter has less boundary condition and parametric sensitivity than summer.

17 WIND ENERGY↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

Robust design of semi-automated clustering models for 4D-STEM datasets

Materials discovery and design require characterizing material structures at the nanometer and sub-nanometer scale. Four-Dimensional Scanning Transmission Electron Microscopy (4D-STEM) resolves the crystal structure of materials, but many 4D-STEM data analysis pipelines are not suited for the identification of anomalous and unexpected structures. This work introduces improvements to the iterative Non-Negative Matrix Factorization (NMF) method by implementing consensus clustering for ensemble learning. We evaluate the performance of models during parameter tuning and find that consensus clustering improves performance in all cases and is able to recover specific grains missed by the best performing model in the ensemble. The methods introduced in this work can be applied broadly to materials characterization datasets to aid in the design of new materials.

Bruefach, Alexandra (ORCID:0000000209323477)↗

Ensemble voting-based fault classification and location identification for a distribution system with microgrids using smart meter measurements

This study presents an ensemble learning approach for fault classification and location identification in a smart distribution network containing photovoltaics (PV)-based microgrid. Lack of available data points and the unbalanced nature of the distribution system make fault handling a challenging task for utilities. The proposed method uses event-driven voltage data from smart meters to classify and locate faults. The ensemble voting classifier is composed of three base learners; random forest, k-nearest neighbours, and artificial neural network. The fault location (FL) task has been formulated as a classification problem where the fault type is classified in the first step and based on the fault type, the faulty bus is identified. The method is tested on IEEE-123 bus system modified with added PV-based microgrid along with dynamic loading conditions and varying fault resistances from 0 to 20 Ω for both unbalanced and balanced fault types. A further sensitivity analysis has been done to test the robustness of the proposed method under various noise levels and data loss errors in the smart meter measurements. The ensemble method shows improved performance and robustness compared to some previously proposed FL methods. Finally, the proposed method has been experimentally validated on a real-time simulation-based testbed using a state-of-the-art digital real-time simulator, industry standard DNP3 communication protocol and a cpu-based control centre running the FL algorithm.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multi-Agent Hierarchical Deep Reinforcement Learning for HVAC Control With Flexible DERs

As electricity consumption in commercial and residential buildings continues to rise, reducing energy costs presents an increasing challenge. Heating, ventilating, and air-conditioning (HVAC) systems, which typically account for 40%-50% of a building's energy use, are prime targets for energy savings. Intelligent control of HVAC temperature through the exploitation of HVAC load flexibility brings significant potential to reduce energy consumption and electricity expenses. The nonlinear models of HVAC systems challenge traditional control methods, while the uncertainty introduced by HVAC load flexibility complicates distributed energy resource (DER) management using conventional optimal dispatch techniques. In response to these challenges, we propose a hierarchical multi-agent deep reinforcement learning (DRL) approach. The lower-level agents focus on balancing comfort and energy conservation, while the upper-level DRL agents optimize the use of DERs to reduce peak demand based on the control outcomes of the HVAC by the lower-level agents. Here, in the upper-level agents, we incorporate a multi-agent structure based on ensemble learning, which acts based on historical and current data without relying on precise load forecasting to address the delayed rewarding issue in DRL. This allows for the effective reduction of energy costs. The proposed method is tested using a real-world microgrid comprising 413 buildings in Southern California, and the results demonstrate that our approach can significantly reduce overall electricity bills while ensuring the comfort of consumers and residents.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Boosting Barlow Twins Reduced Order Modeling for Machine Learning‐Based Surrogate Models in Multiphase Flow Problems

Abstract We present an innovative approach called boosting Barlow Twins reduced order modeling (BBT‐ROM) to enhance the reliability of machine learning surrogate models for multiphase flow problems. BBT‐ROM builds upon Barlow Twins reduced order modeling that leverages self‐supervised learning to effectively handle linear and nonlinear manifolds by constructing well‐structured latent spaces of input parameters and output quantities. To address the challenge of high contrast data in multiphase flow problems due to injection wells and faults, we employ a boosting algorithm within BBT‐ROM. This algorithm sequentially trains a set of weak models (i.e., inaccurate models), improving prediction accuracy through ensemble learning. To evaluate the performance of BBT‐ROM, we conduct three three‐dimensional multiphase flow problems, including waterflooding and geologic carbon storage (GCS), with varying numbers of input parameter cases and model domain features. The results demonstrate that BBT‐ROM excels at predicting non‐wetting phase saturation (e.g., oil or saturation) and fluid pressure, with average relative errors ranging from 0.5% to 3%. Importantly, BBT‐ROM showcases robustness when faced with limited input parameter space during GCS testing.

58 GEOSCIENCES↗

Short-term solar radiation forecast using total sky imager via transfer learning

Ground-based sky cameras, which capture hemispherical images, have been extensively used for localized monitoring of clouds. This paper proposes a short-term forecasting approach based on transfer learning using Total Sky-Imager (TSI) images of the Southern Great Plains (SGP) site obtained from the Atmospheric Radiation Measurement (ARM) dataset. An accurate estimation of solar irradiance using TSI is key for short-term solar energy generation forecasting and optimal energy consumption planning. We make use of deep neural network architectures such as AlexNet and ResNet-101 to extract the underlying deep convolution features from TSI images and then train using an ensemble learning approach to model and forecast solar radiation. We demonstrate the performance of the proposed approach by showcasing the best and worst cases. Thus, the transfer learning approach significantly reduces the time and resources required for modeling solar radiation. We outperform with reference to another state-of-art technique for solar modeling using TSI images at different forecast lead times.

54 ENVIRONMENTAL SCIENCES↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗