Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Prediction of Weather Impacted Airport Capacity using Ensemble Learning

Ensemble learning with the Bagging Decision Tree (BDT) model was used to assess the impact of weather on airport capacities at selected high-demand airports in the United States. The ensemble bagging decision tree models were developed and validated using the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) data and weather forecast at these airports. The study examines the performance of BDT, along with traditional single Support Vector Machines (SVM), for airport runway configuration selection and airport arrival rates (AAR) prediction during weather impacts. Testing of these models was accomplished using observed weather, weather forecast, and airport operation information at the chosen airports. The experimental results show that ensemble methods are more accurate than a single SVM classifier. The airport capacity ensemble method presented here can be used as a decision support model that supports air traffic flow management to meet the weather impacted airport capacity in order to reduce costs and increase safety.

Weather impact↗

Crack fault diagnosis of rotating machine in nuclear power plant based on ensemble learning

Crack faults in rotating machines can cause machine shutdown or scrapping, endangering the normal operation and safety of nuclear power plants. Intelligent diagnostic techniques based on machine learning have the potential to diagnose crack faults. However, problems such as scarcity of field fault data and high noise of plant measurements pose challenges to the application of machine learning. Here this study proposes an ensemble learning approach to mitigate the negative impacts of the problems. Ensemble learning is a strategy for combining multiple machine learning models into a composite model. The basic idea of ensemble learning is that even if one model makes a mistake, other models can correct it. Case studies based on bearing and gear system fault experiments show that the proposed ensemble learning models have better diagnostic results than the single model in the presence of noise and small data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Large-Scale High-Resolution Coastal Mangrove Forests Mapping Across West Africa With Machine Learning Ensemble and Satellite Big Data

Coastal mangrove forests provide important ecosystem goods and services, including carbon sequestration, biodiversity conservation, and hazard mitigation. However, they are being destroyed at an alarming rate by human activities. To characterize mangrove forest changes, evaluate their impacts, and support relevant protection and restoration decision making, accurate and up-to-date mangrove extent mapping at large spatial scales is essential. Available large-scale mangrove extent data products use a single machine learning method commonly with 30 m Landsat imagery, and significant inconsistencies remain among these data products. With huge amounts of satellite data involved and the heterogeneity of land surface characteristics across large geographic areas, finding the most suitable method for large-scale high-resolution mangrove mapping is a challenge. The objective of this study is to evaluate the performance of a machine learning ensemble for mangrove forest mapping at 20 m spatial resolution across West Africa using Sentinel-2 (optical) and Sentinel-1 (radar) imagery. The machine learning ensemble integrates three commonly used machine learning methods in land cover and land use mapping, including Random Forest (RF), Gradient Boosting Machine (GBM), and Neural Network (NN). The cloud-based big geospatial data processing platform Google Earth Engine (GEE) was used for pre-processing Sentinel-2 and Sentinel-1 data. Extensive validation has demonstrated that the machine learning ensemble can generate mangrove extent maps at high accuracies for all study regions in West Africa (92%–99% Producer’s Accuracy, 98%–100% User’s Accuracy, 95%–99% Overall Accuracy). This is the first-time that mangrove extent has been mapped at a 20 m spatial resolution across West Africa. The machine learning ensemble has the potential to be applied to other regions of the world and is therefore capable of producing high-resolution mangrove extent maps at global scales periodically.

coastal environment↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Interpretable ensemble learning unveils main aerosol optical properties in predicting cloud condensation nuclei number concentration

Variations in cloud condensation nuclei number concentration (N CCN ) significantly influence cloud microphysics, yet direct N CCN measurements remain challenging. Here, we present an N CCN ensemble learning (NEL) model utilizing ensemble learning and interpretability analysis on aerosol optical parameters. Validated at two land sites, two ocean sites and one polar site within the Atmospheric Radiation Measurement program, the mean absolute percentage error range of the NEL model across different environments is from 12% to 36%, demonstrating high accuracy. Key findings reveal that aerosol optical parameters can serve as predictors for N CCN . Aerosol scattering and backscattering coefficients, absorption coefficient, backscatter fraction (BSF), and Ångström exponent (AE) are positively correlated with N CCN , while single scattering albedo shows negative correlations. N CCN prediction at land sites is highly sensitive to BSF, largely driven by the backscattering coefficient, as fine particles dominate in these sites. At ocean sites, N CCN prediction is more sensitive to AE, primarily influenced by the scattering coefficient, due to the higher proportion of larger particles. At the polar site, N CCN prediction shows sensitivity to both BSF and AE, mainly driven by the scattering coefficient, as polar sites are cleaner and contain larger particles. These differences reflect the variation in particle size and number concentration across different environments.

Atmospheric science↗

An adaptive knowledge-based data-driven approach for turbulence modeling using ensemble learning technique under complex flow configuration: 3D PWR sub-channel with DNS data

This work describes a new approach to increase the accuracy of Reynolds-averaged Navier–Stokes (RANS) in modeling turbulence flow leveraging the machine learning technique. Traditionally, different turbulence models for Reynolds stress are developed for different flow patterns based on human knowledge. Each turbulence model has a certain application domain and prediction uncertainty. In recent years, with the rapid improvements of machine learning techniques, researchers start to develop an approach to compensate for the prediction discrepancy of traditional turbulence models with statistical models and data. However, the approach has deficiencies in several aspects. For example, the amount of human knowledge introduced to the statistical model couldn’t be controlled, which makes the statistical model learn from a very naïve stage and limits its application. In this work, a new approach is developed to address those deficiencies. Here, the new approach uses the “ensemble learning” technique to control the amount of human knowledge introduced into the statistical model. Therefore, the new approach could be adaptive to the multiple application domains. In conclusion, according to the results of case study, the new approach shows higher accuracy than both traditional turbulence models and the previous machine learning approach.

42 ENGINEERING↗

Developing a heterogeneous ensemble learning framework to evaluate Alkali-silica reaction damage in concrete using acoustic emission signals

The monitoring and evaluation of Alkali-silica reaction (ASR) damage in concrete structures are required to ensure the serviceability and integrity of concrete infrastructures such as bridges and dams. The innovation of this paper lies in the development of an automatic ASR monitoring and evaluation approach by leveraging acoustic emission (AE) and a heterogeneous ensemble learning framework. Here, in this paper, ASR was monitored by AE sensors attached to a concrete specimen, which was placed in a chamber with high humidity and temperature. The recorded AE signals were filtered and divided by four ASR phases according to signal strength, crack width and expansion strain. A heterogeneous ensemble network including convolutional neural networks (CNN) and random forest models was employed to learn different features from AE signals and classify the AE signals into their corresponding phases. The results suggest that the proposed model has a high performance and classifies the signals into the assigned phases with high accuracy.

42 ENGINEERING↗

High-Resolution PM2.5 Concentrations Estimation Based on Stacked Ensemble Learning Model Using Multi-Source Satellite TOA Data

Nepal has experienced severe fine particulate matter (PM2.5) pollution in recent years. However, few studies have focused on the distribution of PM2.5 and its variations in Nepal. Although many researchers have developed PM2.5 estimation models, these models have mainly focused on the kilometer scale, which cannot provide accurate spatial distribution of PM2.5 pollution. Based on Gaofen-1/6 and Landsat-8/9 satellite data, we developed a stacked ensemble learning model (named XGBLL) combined with meteorological data, ground PM2.5 concentrations, ground elevation, and population data. The model includes two layers: a XGBoost and Light GBM model in the first layer, and a linear regression model in the second layer. The accuracy of XGBLL model is better than that of a single model, and the fusion of multi-source satellite remote sensing data effectively improves the spatial coverage of PM2.5 concentrations. Besides, the spatial distribution of the daily mean PM2.5 concentrations in the Kathmandu region under different air conditions was analyzed. The validation results showed that the monthly averaged dataset was accurate (R2 = 0.80 and root mean square error = 7.07). In addition, compared to previous satellite PM2.5 datasets in Nepal, the dataset produced in this study achieved superior accuracy and spatial resolution.

Environmental Sciences & Ecology↗

Early Fault Detection in Particle Accelerator Power Electronics Using Ensemble Learning

Early fault detection and fault prognosis are crucial to ensure efficient and safe operations of complex engineering systems such as the Spallation Neutron Source (SNS) and its power electronics (high voltage converter modulators). Following an advanced experimental facility setup that mimics SNS operating conditions, the authors successfully conducted 21 early fault detection experiments, where fault precursors are introduced in the system to a degree enough to cause degradation in the waveform signals, but not enough to reach a real fault. Nine different machine learning techniques based on ensemble trees, convolutional neural networks, support vector machines, and hierarchical voting ensembles are proposed to detect the fault precursors. Although all 9 models have shown a perfect and identical performance during the training and testing phase, the performance of most models has decreased in the next test phase once they got exposed to realworld data from the 21 experiments. The hierarchical voting ensemble, which features multiple layers of diverse models, maintains a distinguished performance in early detection of the fault precursors with 95% success rate (20/21 tests), followed by adaboost and extremely randomized trees with 52% and 48% success rates, respectively. The support vector machine models were the worst with only 24% success rate (5/21 tests). The study concluded that a successful implementation of machine learning in the SNS or particle accelerator power systems would require a major upgrade in the controller and the data acquisition system to facilitate streaming and handling big data for the machine learning models. In addition, this study shows that the best performing models were diverse and based on the ensemble concept to reduce the bias and hyperparameter sensitivity of individual models.

43 PARTICLE ACCELERATORS↗

Improving the accuracy of freight mode choice models: A case study using the 2017 CFS PUF data set and ensemble learning techniques

Here, the US Census Bureau has collected two rounds of experimental data from the Commodity Flow Survey, providing shipment-level characteristics of nationwide commodity movements, published in 2012 (i.e., Public Use Microdata) and in 2017 (i.e., Public Use File). With this information, data-driven methods have become increasingly valuable for understanding detailed patterns in freight logistics. In this study, we used the 2017 Commodity Flow Survey Public Use File data set to explore building a high-performance freight mode choice model, considering three main improvements: (1) constructing local models for each separate commodity/industry category; (2) extracting useful geographical features, particularly the derived distance of each freight mode between origin/destination zones; and (3) applying additional ensemble learning methods such as stacking or voting to combine results from local and unified models for improved performance. The proposed method achieved over 92% accuracy without incorporating external information, an over 19% increase compared to directly fitting Random Forests models over 10,000 samples. Furthermore, SHAP (Shapely Additive Explanations) values were computed to explain the outputs and major patterns obtained from the proposed model. The model framework could enhance the performance and interpretability of existing freight mode choice models.

42 ENGINEERING↗

Inferring the Redshift of More than 150 GRBs with a Machine-learning Ensemble Model

Gamma-ray bursts (GRBs), due to their high luminosities, are detected up to a redshift of 10, and thus have the potential to be vital cosmological probes of early processes in the Universe. Fulfilling this potential requires a large sample of GRBs with known redshifts, but due to observational limitations, only 11% have known redshifts (z). There have been numerous attempts to estimate redshifts via correlation studies, most of which have led to inaccurate predictions. To overcome this, we estimated GRB redshift via an ensemble-supervised machine-learning (ML) model that uses X-ray afterglows of long-duration GRBs observed by the Neil Gehrels Swift Observatory. The estimated redshifts are strongly correlated (a Pearson coefficient of 0.93) and have an rms error, namely, the square root of the average squared error $\langle$Δz 2 $\rangle$, of 0.46 with the observed redshifts showing the reliability of this method. The addition of GRB afterglow parameters improves the predictions considerably by 63% compared to previous results in peer-reviewed literature. Finally, we use our ML model to infer the redshifts of 154 GRBs, which increase the known redshifts of long GRBs with plateaus by 94%, a significant milestone for enhancing GRB population studies that require large samples with redshift.

79 ASTRONOMY AND ASTROPHYSICS↗

A Predictive Prescription Framework for Stochastic Unit Commitment Using Boosting Ensemble Learning Algorithms

To take unit commitment (UC) decisions under uncertain load, most existing stochastic optimization (SO) frameworks adopt a generic representation of uncertainty. While load levels that materialize on a particular day are influenced by various covariates (such as the day of the week or temperature), SO frameworks typically disregard such side observations, wasting actionable information that could significantly enhance decision quality. Here, this article proposes a contextual SO (CSO) framework for UC under uncertain load, which can effectively exploit covariate observations in conjunction with a class of machine learning (ML) algorithms to improve the out-of-sample performance of UC decisions. It shows how three ML algorithms, adaptive boosting, gradient boosted trees, and extreme gradient boosting, can be used to this end, constituting the first application of these algorithms in any CSO framework. Using real-world data harvested from the New York ISO grid, we measure the out-of-sample performance of the framework in terms of total operation cost, shed load values, locational marginal prices, and total payments by the loads, against several benchmark methods proposed in the literature. The article has an online companion (Yurdakul et al.), wherein we present additional results and lay out further mathematical formulations used in this work.

42 ENGINEERING↗

Understanding and control of Zener pinning via phase field and ensemble learning

Zener pinning refers to the dispersion of fine particles which influences grain size distribution via movement of grain boundaries in a polycrystalline material. Grain size distribution in polycrystals has a significant impact on their properties including physical, chemical, mechanical, and optical to name a few. We explore the use of Phase-field modeling and machine-learning techniques to understand and improve the control of grain size distribution via Zener pinning in polycrystalline materials. We develop a machine learning model that determines the relative importance of various parameters to exercise microstructure control via Zener pinning. Our workflow combines high-throughput phase-field simulations and machine learning to address the computational bottlenecks associated with large-scale simulations as well as identify features necessary for microstructure control in polycrystals. A random forest (RF) regression model was developed to predict grain sizes based on five Phase-field model parameters, achieving an average prediction error of 0.72 nm for the training data and 1.44 nm for the test data. The importance of the input parameters is analyzed using the SHapley Additive exPlanations (SHAP) approach which reveals that diffusivity, volume fraction, and particle diameter are the most important parameters in determining the final grain size. These findings will allow us to select the best second-phase particles, optimize grain size distributions and thus design microstructures with the desired properties. The developed method is a highly versatile and generalizable approach that can be used to assess the combined effects of individual features in the presence of multiple variables.

36 MATERIALS SCIENCE↗

Explosion Detection Using Smartphones: Ensemble Learning with the Smartphone High-Explosive Audio Recordings Dataset and the ESC-50 Dataset

Explosion monitoring is performed by infrasound and seismoacoustic sensor networks that are distributed globally, regionally, and locally. However, these networks are unevenly and sparsely distributed, especially at the local scale, as maintaining and deploying networks is costly. With increasing interest in smaller-yield explosions, the need for more dense networks has increased. To address this issue, we propose using smartphone sensors for explosion detection as they are cost-effective and easy to deploy. Although there are studies using smartphone sensors for explosion detection, the field is still in its infancy and new technologies need to be developed. We applied a machine learning model for explosion detection using smartphone microphones. The data used were from the Smartphone High-explosive Audio Recordings Dataset (SHAReD), a collection of 326 waveforms from 70 high-explosive (HE) events recorded on smartphones, and the ESC-50 dataset, a benchmarking dataset commonly used for environmental sound classification. Two machine learning models were trained and combined into an ensemble model for explosion detection. The resulting ensemble model classified audio signals as either “explosion”, “ambient”, or “other” with true positive rates (recall) greater than 96% for all three categories.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

In Silico Prediction of the Toxicity of Nitroaromatic Compounds: Application of Ensemble Learning QSAR Approach

In this work, a dataset of more than 200 nitroaromatic compounds is used to develop Quantitative Structure–Activity Relationship (QSAR) models for the estimation of in vivo toxicity based on 50% lethal dose to rats (LD 50 ). An initial set of 4885 molecular descriptors was generated and applied to build Support Vector Regression (SVR) models. The best two SVR models, SVR_A and SVR_B, were selected to build an Ensemble Model by means of Multiple Linear Regression (MLR). The obtained Ensemble Model showed improved performance over the base SVR models in the training set (R 2 = 0.88), validation set (R 2 = 0.95), and true external test set (R 2 = 0.92). The models were also internally validated by 5-fold cross-validation and Y-scrambling experiments, showing that the models have high levels of goodness-of-fit, robustness and predictivity. The contribution of descriptors to the toxicity in the models was assessed using the Accumulated Local Effect (ALE) technique. The proposed approach provides an important tool to assess toxicity of nitroaromatic compounds, based on the ensemble QSAR model and the structural relationship to toxicity by analyzed contribution of the involved descriptors.

54 ENVIRONMENTAL SCIENCES↗

Simultaneous quantification of uranium( VI ), samarium, nitric acid, and temperature with combined ensemble learning, laser fluorescence, and Raman scattering for real-time monitoring

In this work, laser-induced fluorescence spectroscopy (LIFS), Raman spectroscopy, and a stacked regression ensemble was developed for near real-time quantification of uranium(VI) (1–100 μg mL –1 ), samarium (0–200 μg mL –1 ) and nitric acid (0.1–4 M) with varying temperature (20 °C–45 °C). LIFS applications range from fundamental lab-scale studies to real-time process monitoring at industrial levels, such as nuclear reprocessing applications, provided the phenomena affecting the fluorescence spectrum are accounted for (e.g., absorption, quenching, complexation). Multiple chemometric models were examined and compared to a more traditional multivariate regression approach called partial least squares (PLS). Results obtained on synthetic samples selected using D-optimal experimental design indicated that a stacked regression method, which included ridge regression, random forest, PLS, and an eXtreme gradient boost algorithm, successfully measured uranium(VI) concentrations directly in nitric acid without measuring luminescence lifetimes or standard addition. The top model resulted in percent root-mean-square error of prediction values of 5.2, 1.9, 3.0, and 2.3% for U(VI), Sm 3+ , HNO 3 , and temperature, respectively. The approach may be useful for quantifying fluorescent fission products (e.g., Sm 3+ ) to provide information on burnup of irradiated nuclear fuel. This novel framework reinforces the applicability of LIFS for real-time applications in nuclear fuel cycle applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

EnZymClass: Substrate specificity prediction tool of plant acyl-ACP thioesterases based on ensemble learning

Characterizing the functional properties of plant acyl-ACP thioesterases (TEs), a key enzyme class used in the production of renewable oleochemicals in microbial hosts, experimentally, can be an expensive and time consuming process since it requires manual screening of thousands of candidates in a database. Using amino acid sequence to computationally predict an enzyme’s function might accelerate this process; however obtaining the necessary amount of information on previously characterized enzymes and their respective sequences required by standard Machine Learning (ML) based approaches to accurately infer sequence-function relationships can be prohibitive, especially with a low-throughput testing cycle. Experimental noise, unbalanced dataset where high sequence similarity does not always imply identical functional properties will further prevent robust prediction performance. Herein we present a ML method, Ensemble method for enZyme Classification (EnZymClass), that is specifically designed to address these issues. We used EnZymClass to classify TEs into short, long and mixed free fatty acid substrate specificity categories. While general guidelines for inferring substrate specificity have been proposed before, prediction of chain-length preference from primary sequence has remained elusive for plant acyl-ACP TEs. By applying EnZymClass to a subset of TEs in the ThYme database, we identified two medium chain TEs, ClFatB3 and CwFatB2, with previously uncharacterized activity in E. coli fatty acid production hosts.

59 BASIC BIOLOGICAL SCIENCES↗

Photometric redshift-aided classification using ensemble learning

We present SHEEP, a new machine learning approach to the classic problem of astronomical source classification, which combines the outputs from the XGBoost, LightGBM, and CatBoost learning algorithms to create stronger classifiers. A novel step in our pipeline is that prior to performing the classification, SHEEP first estimates photometric redshifts, which are then placed into the data set as an additional feature for classification model training; this results in significant improvements in the subsequent classification performance. SHEEP contains two distinct classification methodologies: (i) Multi-class and (ii) one versus all with correction by a meta-learner. We demonstrate the performance of SHEEP for the classification of stars, galaxies, and quasars using a data set composed of SDSS and WISE photometry of 3.5 million astronomical sources. The resulting F1 -scores are as follows: 0.992 for galaxies; 0.967 for quasars; and 0.985 for stars. In terms of the F1-scores for the three classes, SHEEP is found to outperform a recent RandomForest-based classification approach using an essentially identical data set. Our methodology also facilitates model and data set explainability via feature importances; it also allows the selection of sources whose uncertain classifications may make them interesting sources for follow-up observations.

79 ASTRONOMY AND ASTROPHYSICS↗