Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ensemble learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Short-term solar radiation forecast using total sky imager via transfer learning

Ground-based sky cameras, which capture hemispherical images, have been extensively used for localized monitoring of clouds. This paper proposes a short-term forecasting approach based on transfer learning using Total Sky-Imager (TSI) images of the Southern Great Plains (SGP) site obtained from the Atmospheric Radiation Measurement (ARM) dataset. An accurate estimation of solar irradiance using TSI is key for short-term solar energy generation forecasting and optimal energy consumption planning. We make use of deep neural network architectures such as AlexNet and ResNet-101 to extract the underlying deep convolution features from TSI images and then train using an ensemble learning approach to model and forecast solar radiation. We demonstrate the performance of the proposed approach by showcasing the best and worst cases. Thus, the transfer learning approach significantly reduces the time and resources required for modeling solar radiation. We outperform with reference to another state-of-art technique for solar modeling using TSI images at different forecast lead times.

54 ENVIRONMENTAL SCIENCES↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Linking repeat lidar with Landsat products for large scale quantification of fire-induced permafrost thaw settlement in interior Alaska

The permafrost–fire–climate system has been a hotspot in research for decades under a warming climate scenario. Surface vegetation plays a dominant role in protecting permafrost from summer warmth, thus, any alteration of vegetation structure, particularly following severe wildfires, can cause dramatic top–down thaw. A challenge in understanding this is to quantify fire-induced thaw settlement at large scales (>1000 km 2 ). In this study, we explored the potential of using Landsat products for a large-scale estimation of fire-induced thaw settlement across a well-studied area representative of ice-rich lowland permafrost in interior Alaska. Six large fires have affected ~1250 km 2 of the area since 2000. We first identified the linkage of fires, burn severity, and land cover response, and then developed an object-based machine learning ensemble approach to estimate fire-induced thaw settlement by relating airborne repeat lidar data to Landsat products. The model delineated thaw settlement patterns across the six fire scars and explained ~65% of the variance in lidar-detected elevation change. Our results indicate a combined application of airborne repeat lidar and Landsat products is a valuable tool for large scale quantification of fire-induced thaw settlement.

54 ENVIRONMENTAL SCIENCES↗

Unraveling Adsorbate-Induced Structural Evolution of Iron Carbide Nanoparticles

Iron carbide (Fe x C y ) nanoparticles (NPs) are promising candidates for replacing platinum group metals in industrial applications, such as high-temperature Fischer–Tropsch synthesis. However, due to their amorphous nature, characterization of the active sites has been challenging experimentally and computationally. Here, using a combined density functional theory (DFT), neural network interatomic potential-assisted global optimization, and ensemble learning study, we evaluate dynamic surface changes associated with syngas (H and CO) interactions. For this purpose, we have developed a general procedure that we use to model an experimentally relevant 270-atom Fe 182 C 88 NP using the neural network-assisted stochastic surface walk global optimization algorithm (SSW-NN). Once generated, the Fe 182 C 88 NP active sites and particle morphology are thoroughly characterized before the effects of syngas adsorbate interactions are explored by using DFT and molecular dynamics simulations. Lastly, we explore correlations between geometric and electronic features of the active sites and the adsorption of H (H ads ), using a regularized random forest machine learning algorithm. In doing so, we identified the Fe–C coordination number and p orbital occupancy as the most important descriptors affecting H ads . Furthermore, using a combined ML and quantum chemistry approach, our work demonstrates a general and efficient procedure for generating and probing complex surface phenomena on binary nanoparticles.

Adsorption↗

Probing Electron Beam Induced Transformations on a Single-Defect Level via Automated Scanning Transmission Electron Microscopy

A robust approach for real-time analysis of the scanning transmission electron microscopy (STEM) data streams, based on ensemble learning and iterative training (ELIT) of deep convolutional neural networks, is implemented on an operational microscope, enabling the exploration of the dynamics of specific atomic configurations under electron beam irradiation via an automated experiment in STEM. Combined with beam control, this approach allows studying beam effects on selected atomic groups and chemical bonds in a fully automated mode. Here, in this study, we demonstrate atomically precise engineering of single vacancy lines in transition metal dichalcogenides and the creation and identification of topological defects in graphene. The ELIT-based approach facilitates direct on-the-fly analysis of the STEM data and engenders real-time feedback schemes for probing electron beam chemistry, atomic manipulation, and atom by atom assembly.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Development of Machine Learning Algorithm for Pebble Bed Modular Reactor Misuse Detection

The objective of this work was to develop a machine learning ensemble that could assist pebble bed reactor verification by evaluating whether a given pebble circulating through a PBR was normal or anomalous using gamma spectroscopy measurements from a notional PBR burnup measurement system. Using a PBR reference design, data sets of synthetic gamma spectra representative of BUMS measurements of normal and anomalous pebbles that may be used to produce special fissile material were generated to train and test an ML anomaly detection ensemble on two reference scenarios – substitution of normal pebbles with target pebbles for production of Pu or 233 U. The ML ensemble correctly identified all anomalous pebbles in the testing data set, and while perfect ensemble performance is normally indicative of overfitting, it was concluded that significantly lower photon intensity of target pebbles produced distinctly less intense photon spectra to where perfect ensemble performance was expected.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A computational pipeline to generate a synthetic dataset of metal ion sorption to oxides for AI/ML exploration

The charged mineral/electrolyte interfaces are ubiquitous in the surface and subsurface–including the surroundings of the geological disposal sites for radioactive waste. Therefore, understanding how ions interact with charged surfaces is critically important for predicting radionuclide mobility in the case of waste leakage. At present, the Surface Complexation Models (SCMs) are the most successful thermodynamic frameworks to describe ion retention by mineral surfaces. SCMs are interfacial speciation models that account for the effect of the electric field generated by charged surfaces on sorption equilibria. These models have been successfully used to analyze and interpret a broad range of experimental observations including potentiometric and electrokinetic titrations or spectroscopy. Unfortunately, many of the current procedures to solve and fit SCM to experimental data are not optimal, which leads to a non-transferable or non-unique description of interfacial electrostatics and consequently of the strength and extent of ion retention by mineral surfaces. Recent developments in Artificial Intelligence (AI) offer a new avenue to replace SCM solvers and fitting algorithms with trained AI surrogates. Unfortunately, there is a lack of a standardized dataset covering a wide range of SCM parameter values available for AI exploration and training–a gap filled by this study. Here, we described the computational pipeline to generate synthetic SCM data and discussed approaches to transform this dataset into AI-learnable input. First, we used this pipeline to generate a synthetic dataset of electrostatic properties for a broad range of the prototypical oxide/electrolyte interfaces. The next step is to extend this dataset to include complex radionuclide sorption and complexation, and finally, to provide trained AI architectures able to infer SCMs parameter values rapidly from experimental data. Here, we illustrated the AI-surrogate development using the ensemble learning algorithms, such as Random Forest and Gradient Boosting. These surrogate models allow a rapid prediction of the SCM model parameters, do not rely on an initial guess, and guarantee convergence in all cases.

Li, Chunhui↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

A physics-based ensemble machine-learning approach to identifying a relationship between lightning indices and binary lightning hazard

To convert lightning indices generated by numerical weather prediction experiments into binary lightning hazard, a machine-learning tool was developed. This tool, consisting of parallel multilayer perceptron classifiers, was trained on an ensemble of planetary boundary layer schemes and microphysics parameterizations that generated four different lightning indices over 1 week. In a subsequent week, the multi-physics ensemble was applied and the machine-learning tool was used to evaluate the accuracy. Unintuitively, the machine-learning tool performed better on the testing dataset than the training dataset. Much of the error may be attributed to mischaracterizing the convection. The combination of the machine learning model and simulations could not differentiate between cloud-to-cloud lightning and cloud-to-ground lightning, despite being trained on cloud-to-ground lightning. It was found that the simulation most representative of the local operational model was the most accurate simulation tested.

54 ENVIRONMENTAL SCIENCES↗

Resilience-based Recovery Scheduling of Transportation Network in Mixed Traffic Environment: A Deep-Ensemble-Assisted Active Learning Approach

Devising effective post-hazard recovery strategies is critical in enhancing the resilience of transportation networks (TNs). However, existing work does not consider the multiclass users’ travel behavior in network functionality quantification and the metaheuristic solution procedures often suffer from extensive computational burden due to the exploration need in large solution space and the expensive functionality quantification. Here, this study develops a bilevel decision-making framework for the resilience-based recovery scheduling of the TN in a mixed traffic environment with connected and autonomous vehicles (CAVs) and human-driven vehicles (HDVs). The lower level quantifies the TN's functionality over time considering different travel behavior of CAV and HDV users arisen from their different levels of information perception. The upper level presents a novel deep-ensemble-assisted active learning approach to balance optimization performance and computational cost. This framework can help decision makers better quantify the TN's functionality to support effective recovery scheduling of TN with different mixed traffic scenarios ranging from HDV-only to future CAV-dominant traffic. The optimization approach bears the potential to be extended to solving general large-scale network recovery scheduling problems effectively and efficiently. The proposed methodology is demonstrated using a real-world traffic network in Southern California under earthquake considering deterministic and stochastic repair durations.

42 ENGINEERING↗

Predicting concrete compressive strength using hybrid ensembling of surrogate machine learning models

This study aims to implement a hybrid ensemble surrogate machine learning technique in predicting the compressive strength (CS) of concrete, an important parameter used for durability design and service life prediction of concrete structures in civil engineering projects. For this purpose, an experimental database consisting of 1030 records has been compiled from the machine learning repository of the University of California, Irvine. The database was used to train and validate four conventional machine learning (CML) models, namely Artificial Neural Network (ANN), Linear and Non-Linear Multivariate Adaptive Regression Splines (MARS-L and MARS-C), Gaussian Process Regression (GPR), and Minimax Probability Machine Regression (MPMR). Subsequently, the predicted outputs of CML models were combined and trained using ANN to construct the Hybrid Ensemble Model (HENSM). It is observed that the proposed HENSM produces higher predictive accuracy compared to the CML models used in the present study. The predictive performance of all models for CS prediction was compared using the testing dataset and it is found that the HENSM model attained the highest predictive accuracy in both phases. Based on the experimental results, the newly constructed HENSM model is very potential to be a new alternative in handling the overfitting issues of CML models and hence, can be used to predict the concrete CS, including the design of less polluting and more sustainable concrete constructions.

36 MATERIALS SCIENCE↗

Ensemble Federated Machine Learning‐Based Cybersecurity Situational Awareness in Microgrid Network

Cyber-physical microgrids are vulnerable to stealthy cybersecurity threats that disguise their actions through the exploitation of system knowledge. Such actions can severely impacts microgrids deployed in defense bases, slowing the response time of military forces during national emergencies. Several machine-learning algorithms have been proposed to detect intrusions in the grid networks; however, these traditional machine-learning algorithms lack data privacy and are subject to several adversarial machine-learning threats. This paper proposes a novel federated machine learning (FML)-based three-model framework to detect and identify stealthy data-integrity attacks while ensuring data privacy in microgrid networks. The proposed architecture uses a variational mode decomposition technique to extract derived features from incoming measurement and control datasets. The extraction of these derived features allows FML models to learn minute variations in data patterns that allow them to perform significantly better than the models trained with generic datasets consisting of raw features. Our experimental results show the efficient performance of the proposed methodology against different types of data integrity attacks while considering primary and secondary controllers in microgrids. Further, the applied FML-integrated random forest ensemble algorithm outperforms the existing generic FML algorithms during noisy and noise-free datasets with prediction latencies of only 91–134 µs per sample within the 0.1 s sampling interval and requires communication bandwidth of around ∼8.25 KB/s at the control center and ∼2.7 KB/s per edge client for communication.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Utilization of Synthetic Near-Infrared Spectra via Generative Adversarial Network to Improve Wood Stiffness Prediction

Near-infrared (NIR) spectroscopy is widely used as a nondestructive evaluation (NDE) tool for predicting wood properties. When deploying NIR models, one faces challenges in ensuring representative training data, which large datasets can mitigate but often at a significant cost. Machine learning and deep learning NIR models are at an even greater disadvantage because they typically require higher sample sizes for training. In this study, NIR spectra were collected to predict the modulus of elasticity (MOE) of southern pine lumber (training set = 573 samples, testing set = 145 samples). To account for the limited size of the training data, this study employed a generative adversarial network (GAN) to generate synthetic NIR spectra. The training dataset was fed into a GAN to generate 313, 573, and 1000 synthetic spectra. The original and enhanced datasets were used to train artificial neural networks (ANNs), convolutional neural networks (CNNs), and light gradient boosting machines (LGBMs) for MOE prediction. Overall, results showed that data augmentation using GAN improved the coefficient of determination (R 2 ) by up to 7.02% and reduced the error of predictions by up to 4.29%. ANNs and CNNs benefited more from synthetic spectra than LGBMs, which only yielded slight improvement. All models showed optimal performance when 313 synthetic spectra were added to the original training data; further additions did not improve model performance because the quality of the datapoints generated by GAN beyond a certain threshold is poor, and one of the main reasons for this can be the size of the initial training data fed into the GAN. LGBMs showed superior performances than ANNs and CNNs on both the original and enhanced training datasets, which highlights the significance of selecting an appropriate machine learning or deep learning model for NIR spectral-data analysis. The results highlighted the positive impact of GAN on the predictive performance of models utilizing NIR spectroscopy as an NDE technique and monitoring tool for wood mechanical-property evaluation. Further studies should investigate the impact of the initial size of training data, the optimal number of generated synthetic spectra, and machine learning or deep learning models that could benefit more from data augmentation using GANs.

59 BASIC BIOLOGICAL SCIENCES↗

Sampling lattices in semi-grand canonical ensemble with autoregressive machine learning

Calculating thermodynamic potentials and observables efficiently and accurately is key for the application of statistical mechanics simulations to materials science. However, naive Monte Carlo approaches, on which such calculations are often dependent, struggle to scale to complex materials in many state-of-the-art disciplines such as the design of high entropy alloys or multi-component catalysts. To address this issue, we adapt sampling tools built upon machine learning-based generative modeling to the materials space by transforming them into the semi-grand canonical ensemble. Furthermore, we show that the resulting models are transferable across wide ranges of thermodynamic conditions and can be implemented with any internal energy model U, allowing integration into many existing materials workflows. We demonstrate the applicability of this approach to the simulation of benchmark systems (AgPd, CuAu) that exhibit diverse thermodynamic behavior in their phase diagrams. Finally, we discuss remaining challenges in model development and promising research directions for future improvements.

36 MATERIALS SCIENCE↗

A Machine Learning Framework for Modeling Ensemble Properties of Atomically Disordered Materials

Atomic disorder can strongly influence material properties such as charge transport, optical response, and catalytic activity. However, efficiently modeling these disorder effects remains challenging for first-principles methods due to the cost of sampling large configurational spaces and computing complex physical quantities. Recent advances of machine learning techniques, particularly graph neural networks (GNNs), has enabled the efficient and accurate predictions of complex material properties, offering promising tools for studying disordered systems. In this work, we present a general machine-learning-assisted computational framework that integrates equivariant GNNs with Monte Carlo simulations to compute the thermodynamic and ensemble-averaged functional properties of disordered materials. Using the surface-termination-disordered MXene monolayer Ti 3 C 2 T 2–x as a representative system, we find that electrical conductivity exhibits an emergent peak near the order–disorder phase transition temperature due to the interplay between electron scattering and doping. In contrast, optical conductivity remains largely insensitive to local atomic disorder and reflects the global surface chemical composition. These results highlight the role of atomic disorder in affecting material properties and demonstrate the potential of our approach for statistically modeling disorder effects in a wide range of materials such as high-entropy alloys and spin liquids.

MXene↗

Variational deep learning of equilibrium transition path ensembles

Here, we present a time-dependent variational method to learn the mechanisms of equilibrium reactive processes and efficiently evaluate their rates within a transition path ensemble. This approach builds off of the variational path sampling methodology by approximating the time-dependent commitment probability within a neural network ansatz. The reaction mechanisms inferred through this approach are elucidated by a novel decomposition of the rate in terms of the components of a stochastic path action conditioned on a transition. This decomposition affords an ability to resolve the typical contribution of each reactive mode and their couplings to the rare event. The associated rate evaluation is variational and systematically improvable through the development of a cumulant expansion. We demonstrate this method in both over- and under-damped stochastic equations of motion, in low-dimensional model systems, and in the isomerization of a solvated alanine dipeptide. In all examples, we find that we can obtain quantitatively accurate estimates of the rates of the reactive events with minimal trajectory statistics and gain unique insights into transitions through the analysis of their commitment probability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗