Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boost machine”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A novel improved model for building energy consumption prediction based on model integration

Building energy consumption prediction plays an irreplaceable role in energy planning, management, and conservation. Constantly improving the performance of prediction models is the key to ensuring the efficient operation of energy systems. Moreover, accuracy is no longer the only factor in revealing model performance, it is more important to evaluate the model from multiple perspectives, considering the characteristics of engineering applications. Based on the idea of model integration, this paper proposes a novel improved integration model (stacking model) that can be used to forecast building energy consumption. The stacking model combines advantages of various base prediction algorithms and forms them into “meta-features” to ensure that the final model can observe datasets from different spatial and structural angles. Two cases are used to demonstrate practical engineering applications of the stacking model. A comparative analysis is performed to evaluate the prediction performance of the stacking model in contrast with existing well-known prediction models including Random Forest, Gradient Boosted Decision Tree, Extreme Gradient Boosting, Support Vector Machine, and K-Nearest Neighbor. The results indicate that the stacking method achieves better performance than other models, regarding accuracy (improvement of 9.5%–31.6% for Case A and 16.2%–49.4% for Case B), generalization (improvement of 6.7%–29.5% for Case A and 7.1%-34.6% for Case B), and robustness (improvement of 1.5%–34.1% for Case A and 1.8%–19.3% for Case B). The proposed model enriches the diversity of algorithm libraries of empirical models.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Gaining Perspective on Unconventional Well Design Choices through Play-level Application of Machine Learning Modeling

The recent development of unconventional oil and gas (O&G) reservoirs has led to an abundant hydrocarbon supply, both domestically and globally. However, there is a continued push to develop new and innovative approaches to improve exploration and extraction efficiencies and overall well productivity moving forward. Substantial improvements in unconventional O&G development are expected through optimized well completion and stimulation strategies aimed at maximizing well productivity. Optimizing well designs will require tailoring to the distinctive geologic conditions present for any newly placed well. To better evaluate the impact of well design attributes and their associated interactions on productivity in a major unconventional play, multivariate machine learning-based models that use empirical datasets were developed. A gradient boosted regression tree (GBRT) algorithm was applied. GBRT has been narrowly investigated for O&G applications but enables straightforward parametric importance and influence evaluation, as well as assessment of parameter interaction effects. Models were trained on well design and locational parameters that serve as a proxy for variable geologic conditions to estimate two types of productivity indicator response variables strongly correlated to estimated ultimate recovery (EUR). The dataset utilized consists of over 7,000 well observations that cover the majority of the productive region of the Marcellus Shale. Model performance was evaluated and algorithm parameters tuned by analyzing the goodness-of-fit for simulated results against observed data in a cross-validation approach. Models were found capable of 73–79 percent prediction accuracy on held out testing data of gas equivalent production and can be used to inform future well design and placement decisions for increasing EUR per well and improving overall field-level recovery. Study results indicate that Marcellus well performance improves most with upscaling perforated interval lengths and water and proppant volumes per foot; but relative productivity improvements are spatially dependent across the play. Finally, optimal combinations of water and proppant on well performance were found to vary depending on well location, emphasizing the utility of data-driven models capable of broad application across a play of interest for informing tailored well design approaches prior to their field deployment.

04 OIL SHALES AND TAR SANDS↗

Explaining System-Level Prognostics with Established Machine Learning Methods

System-level prognostics is crucial for ensuring reliability and enabling predictive maintenance in complex systems with interconnected components. This study presents a framework that integrates data-driven methods to predict the remaining useful life (RUL) of a subsystem under multiple and concurrent faults within a nuclear power plant system with explainable artificial intelligence (XAI). A nuclear power plant (NPP) operation was simulated to model the degradation behavior of NPP components, and four machine learning models—Gradient Boosting Regressor (GBR), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory (LSTM)—were evaluated for prognostics with a novel system RUL parameter. The LSTM model demonstrated potential superior repeatability, while SHAP (SHapley Additive exPlanations) for explainability provided consistent and trustworthy global explanations. In contrast, LIME (Local Interpretable Model-agnostic Explanations) offered localized interpretability but showed reduced stability for sequential data. Key findings include the interplay between component-level degradation and system-wide performance, with LSTM effectively capturing these dynamics through sequence-level predictions. The XAI techniques enhanced transparency by identifying critical features influencing model predictions and aligning with domain knowledge. Furthermore, this framework has significant implications for improving trust and understanding in predictive maintenance, particularly in safety-critical industries like nuclear energy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Predicting Airport Runway Configurations for Decision-Support Using Supervised Learning

One of the most challenging tasks for air traffic controllers is runway configuration management (RCM). It deals with the optimal selection of runways to operate on (for arrivals and departures) based on traffic, surface wind speed, wind direction, other environmental variables, noise constraints, and several other airport-specific factors. It affects the efficiency of the National Airspace System (NAS) and both surface and airspace operations can benefit from better understanding future runway configurations. In this paper, we present a comprehensive implementation of predictive models for runway configuration estimation from large volumes of historical data. Specifically, operational data from two full years (2018 and 2019) is collected, analyzed, and fused together to build the data product used in this work. The data set differs from prior work in the field in terms of its scope, resolution, and variety of factors collected and considered. Meteorological data is collected from two different sources – current weather conditions from METAR (Meteorological Terminal Aviation Routine Weather Report) and forecast weather conditions from Localized Aviation MOS Program (LAMP). Operational data from the Federal Aviation Administration (FAA) Aviation System Performance Metrics (ASPM) related to scheduled and actual number of arrivals and departures, average taxi times, etc. are collected. NASA’s Sherlock Data Warehouse is used to identify critical information such as go-arounds, and other events that might impact RCM decision-making. All data is collected and aggregated over 15-minute intervals throughout the two years. This provides a resolution like the timescales that might be necessary for runway configuration management decision-making. A variety of supervised learning algorithms are tested including Support Vector Machine, Random Forest, Gradient Boosting, etc. including tuning of the model hyperparameters. The modeling process is applied and presented on two representative U.S. airports – Charlotte Douglas International Airport (KCLT) and Denver International Airport (KDEN). The two airports present different levels of complexity in terms of the total number of configurations used and provide a balanced perspective on the generalizability of the developed approach to other airports in the NAS. Initial results are promising (F1 score of 0.91 at KCLT and 0.83 at KDEN) for data in the test set. The final paper will contain a comprehensive comparison between different models and model building strategies as well as further refined results. Most important predictors for each airport will be identified along with a discussion and recommendations on adapting the framework to other scenarios.

Tejas G Puranik↗

Investigating boosted decision trees as a guide for inertial confinement fusion design

Inertial confined fusion experiments at the National Ignition Facility have recently entered a new regime approaching ignition. Improved modeling and exploration of the experimental parameter space were essential to deepening our understanding of the mechanisms that degrade and amplify the neutron yield. The growing prevalence of machine learning in fusion studies opens a new avenue for investigation. Here in this paper, we have applied the Gradient-Boosted Decision Tree machine-learning architecture to further explore the parameter space and find correlations with the neutron yield, a key performance indicator. We find reasonable agreement between the measured and predicted yield, with a mean absolute percentage error on a randomly assigned test set of 35.5%. This model finds the characteristics of the laser pulse to be the most influential in prediction, as well as the hohlraum laser entrance hole diameter and an enhanced capsule fabrication technique. We used the trained model to scan over the design space of experiments from three different campaigns to evaluate the potential of this technique to provide design changes that could improve the resulting neutron yield. While these data-driven model cannot predict ignition without examples of ignited shots in the training set, it can be used to indicate that an unseen shot design will at least be in the upper range of previously observed neutron yields.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine learning-based prediction of enzyme substrate scope: Application to bacterial nitrilases

Predicting the range of substrates accepted by an enzyme from its amino acid sequence is challenging. Although sequenc- and structure-based annotation approaches are often accurate for predicting broad categories of substrate specificity, they generally cannot predict which specific molecules will be accepted as substrates for a given enzyme, particularly within a class of closely related molecules. Combining targeted experimental activity data with structural modeling, ligand docking, and physicochemical properties of proteins and ligands with various machine learning models provides complementary information that can lead to accurate predictions of substrate scope for related enzymes. Here we describe such an approach that can predict the substrate scope of bacterial nitrilases, which catalyze the hydrolysis of nitrile compounds to the corresponding carboxylic acids and ammonia. Each of the four machine learning models (logistic regression, random forest, gradient-boosted decision trees, and support vector machines) performed similarly (average ROC = 0.9, average accuracy = ~82%) for predicting substrate scope for this dataset, although random forest offers some advantages. Finally, this approach is intended to be highly modular with respect to physicochemical property calculations and software used for structural modeling and docking.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning to Predict Joint Performance in Epoxy Composites Based on Process Parameters

Polymer matrix composites are gaining popularity in the aerospace industry due to their high specific strength, fatigue properties, and processability. However, based on current FAA certification guidelines, manufacturers utilizing current state-of-the-art composites made with adhesive bonds commonly install redundant fasteners to guarantee the strength of these adhesively bonded composite parts. The number of fasteners in a single-aisle commercial transport aircraft is typically on the order of 105, which reduces manufacturing rate, increases cost tremendously, and reduces the advantage of the specific strength composites provide. Due to this, the Adhesive Free Bonding of Composites (AERoBOND) project at NASA Langley Research Center has developed a novel assembly process to manufacture complex composite parts without the use of adhesives and fasteners. However, optimization of the process is currently challenging due to the complex and interdependent process parameters. To assist with the optimization, four machine learning algorithms utilizing gradient boosting decision trees were created to provide predictions for the mechanical and characterization properties of the composite parts. Approximately 200 random states from each algorithm were tested, and the models from each state were isolated and analyzed based on their accuracy, a validation process, and their feature importance. This analysis concluded that the models created from the machine learning algorithms could accelerate a parametric study for the AERoBOND process by rapidly optimizing process parameters to achieve desired performance characteristics.

Brennen Michael Middleton↗

Machine Learning to Predict Joint Performance in Epoxy Composites Based on Process Parameters

Polymer matrix composites are gaining popularity in the aerospace industry due to their high specific strength, fatigue properties, and processability. However, based on current FAA certification guidelines, manufacturers utilizing current state-of-the art composites made with adhesive bonds commonly install redundant fasteners to guarantee the strength of these adhesively bonded composite parts.1,2 The number of fasteners in a single-aisle commercial transport aircraft is typically on the order of 105, which reduces manufacturing rate, increases cost tremendously, and reduces the advantage of the specific strength composites provide. Due to this, the Adhesive Free Bonding of Composites (AERoBOND) project at NASA Langley Research Center has developed a novel assembly process to manufacture complex composite parts without the use of adhesives and fasteners.1 However, optimization of the process is currently challenging due to the complex and interdependent process parameters. To assist with the optimization, four machine learning algorithms utilizing gradient boosting decision trees were created to provide predictions for the mechanical and characterization properties of the composite parts. Approximately 200 random states from each algorithm were tested, and the models from each state were isolated and analyzed based on their accuracy, a validation process, and their feature importance. This analysis concluded that the models created from the machine learning algorithms could accelerate a parametric study for the AERoBOND process by rapidly optimizing process parameters to achieve desired performance characteristics.

Brennen M Middleton↗

Tree-Based Ensemble Learning Models for Wall Temperature Predictions in Post-Critical Heat Flux Flow Regimes at Subcooled and Low-Quality Conditions

Accurately predicting post-critical heat flux (CHF) heat transfer is an important but challenging task in water-cooled reactor design and safety analysis. Although numerous heat transfer correlations have been developed to predict post-CHF heat transfer, these correlations are only applicable to relatively narrow ranges of flow conditions due to the complex physical nature of the post-CHF heat transfer regimes. In this paper, a large quantity of experimental data is collected and summarized from the literature for steady-state subcooled and low-quality film boiling regimes with water as the working fluid in vertical tubular test sections. In addition, a low-quality water film boiling (LWFB) database is consolidated with a total of 22,813 experimental data points, which cover a wide flow range of the system pressure from 0.1 to 9.0 MPa, mass flux from 25 to 2750 kg/m 2 s, and inlet subcooling from 1 to 70 °C. Two machine learning (ML) models, based on random forest (RF) and gradient boosted decision tree (GBDT), are trained and validated to predict wall temperatures in post-CHF flow regimes. The trained ML models demonstrate significantly improved accuracies compared to conventional empirical correlations. To further evaluate the performance of these two ML models from a statistical perspective, three criteria are investigated and three metrics are calculated to quantitatively assess the accuracy of these two ML models. For the full LWFB database, the root-mean-square errors between the measured and predicted wall temperatures by the GBDT and RF models are 5.7% and 6.2%, respectively, confirming the accuracy of the two ML models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Learning-based CO 2 concentration prediction: Application to indoor air quality control using demand-controlled ventilation

There have been increasing concerns over the air quality inside buildings as high levels of bio-effluents can cause nausea, dizziness, headaches, and fatigue to the people working in those spaces. First published in 2004 as Standard 62.1, ASHRAE Standard 62.2-2019 requires highly occupied spaces to implement heating, ventilation, and air conditioning (HVAC) that can dilute contaminants produced by occupants. In this regard, occupant-centric ventilation control has been regarded as an effective practice to maintain a satisfactory indoor air quality (IAQ) when dealing with highly variable occupancy environments. However, few established models in current literature and practice consider dynamic occupancy behavior and adaptive IAQ control. To address this gap, a dynamic indoor CO2 model is constructed using machine learning algorithms to forecast CO2concentrations across a range of forecasting horizons. Herein, we tuned and compared six state-of-the-algorithms—including Support Vector Machine, Ada Boost, Random Forest, Gradient Boosting, Logistic Regression, and Multilayer Perceptron. The algorithms’ performances are validated using CO 2 and historical meteorological data collected from a campus classroom with a variable occupancy rate. Simulation results showed that Multilayer Perceptron can strongly predict the volatile CO 2 behavior and also outperforms other algorithms in terms of accuracy. Furthermore, a control strategy capable of modeling and detecting dynamic patterns of CO 2 level is utilized to modulate the ventilation rate in real-time and also reduce the energy consumption. The proposed controller reduced the HVAC fan’s energy consumption by 51.4% and provide ventilation as needed per the ASHRAE standards.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Detection of topological materials with machine learning

Databases compiled using ab initio and symmetry-based calculations now contain tens of thousands of topological insulators and topological semimetals. This makes the application of modern machine learning methods to topological materials possible. Using gradient boosted trees, we show how to construct a machine learning model which can predict the topology of a given existent material with an accuracy of 90%. Such predictions are orders of magnitude faster than actual ab initio calculations. In this work, we use machine learning models to probe how different material properties affect topological features. Notably, we observe that topology is mostly determined by the “coarse-grained” chemical composition and crystal symmetry and depends little on the particular positions of atoms in the crystal lattice. We identify the sources of our model's errors and we discuss approaches to overcome them.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Composite Qdrift-product formulas for quantum and classical simulations in real and imaginary time

Recent study has shown that it can be advantageous to implement a composite channel that partitions the Hamiltonian H for a given simulation problem into subsets A and B such that H = A + B , where the terms in A are simulated with a Trotter-Suzuki channel and the B terms are randomly sampled via the Qdrift algorithm. Here we extend Qdrift and composite product formulas to imaginary time, formulating candidate classical algorithms for quantum Monte Carlo calculations. We upper bound the induced Schatten- 1 → 1 norm on both imaginary-time Qdrift and composite channels. Another recent result demonstrated that simulations of lattice Hamiltonians containing geometrically local interactions can be improved using a Lieb-Robinson argument to decompose H into subsets that contain only terms supported on that subset of the lattice. Here, we provide a quantum algorithm by unifying this result with the composite approach into “local composite channels” and we upper bound the diamond distance. We provide exact numerical simulations of algorithmic cost by counting the number of gates of the form e − i H j t and e − H j β to meet a certain error tolerance ε . In doing so, we optimize the partitioning into sets A and B using gradient boosted tree models from machine learning. These numerical studies are important given that product formulas have been historically known to outperform analytic upper bounds. We show constant factor advantages for a variety of interesting Hamiltonians, the maximum of which is a ≈ 20 -fold speedup that occurs in the simulation of Jellium. Published by the American Physical Society 2024

Pocrnic, Matthew (ORCID:0000000203089376)↗

The LSST AGN Data Challenge: Selection Methods

Abstract Development of the Rubin Observatory Legacy Survey of Space and Time (LSST) includes a series of Data Challenges (DCs) arranged by various LSST Scientific Collaborations that are taking place during the project's preoperational phase. The AGN Science Collaboration Data Challenge (AGNSC-DC) is a partial prototype of the expected LSST data on active galactic nuclei (AGNs), aimed at validating machine learning approaches for AGN selection and characterization in large surveys like LSST. The AGNSC-DC took place in 2021, focusing on accuracy, robustness, and scalability. The training and the blinded data sets were constructed to mimic the future LSST release catalogs using the data from the Sloan Digital Sky Survey Stripe 82 region and the XMM-Newton Large Scale Structure Survey region. Data features were divided into astrometry, photometry, color, morphology, redshift, and class label with the addition of variability features and images. We present the results of four submitted solutions to DCs using both classical and machine learning methods. We systematically test the performance of supervised models (support vector machine, random forest, extreme gradient boosting, artificial neural network, convolutional neural network) and unsupervised ones (deep embedding clustering) when applied to the problem of classifying/clustering sources as stars, galaxies, or AGNs. We obtained classification accuracy of 97.5% for supervised models and clustering accuracy of 96.0% for unsupervised ones and 95.0% with a classic approach for a blinded data set. We find that variability features significantly improve the accuracy of the trained models, and correlation analysis among different bands enables a fast and inexpensive first-order selection of quasar candidates.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems: Preprint

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different sets of preprocess data for the prediction of the frequency nadir in the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. The training and testing data sets are collected by extensive generation scheduling simulations on the Multi-timescale Integrated Dynamic and Scheduling (MIDAS) toolbox. Numerical results show that all five machine learning methods can achieve high performance accuracy for power system nadir frequency estimation. Among them, the gradient boosting and the XGBoost are clear winners by providing the best prediction accuracy.

data driven↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different datasets, i.e., 1) the unit generation dataset and 2) the system total inertia and headroom dataset, for the prediction of the frequency nadir. The training and testing datasets are generated through extensive generation scheduling simulations using Multi-timescale Integrated Dynamic and Scheduling (MI-DAS) toolbox on the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. Numerical results show that all five machine learning methods perform well in predicting the nadir frequency of the system. Among them, the gradient boosting and the XGBoost are clear winners yielding the best prediction accuracy in terms of four evaluation metrics.

data driven↗

Reliable photometric membership (RPM) of galaxies in clusters – I. A machine learning method and its performance in the local universe

ABSTRACT We introduce a new method to determine galaxy cluster membership based solely on photometric properties. We adopt a machine learning approach to recover a cluster membership probability from galaxy photometric parameters and finally derive a membership classification. After testing several machine learning techniques (such as stochastic gradient boosting, model averaged neural network and k-nearest neighbours), we found the support vector machine algorithm to perform better when applied to our data. Our training and validation data are from the Sloan Digital Sky Survey main sample. Hence, to be complete to $M_r^* + 3$, we limit our work to 30 clusters with $z$phot-cl ≤ 0.045. Masses (M200) are larger than $\sim 0.6\times 10^{14} \, \mathrm{M}_{\odot }$ (most above $3\times 10^{14} \, \mathrm{M}_{\odot }$). Our results are derived taking in account all galaxies in the line of sight of each cluster, with no photometric redshift cuts or background corrections. Our method is non-parametric, making no assumptions on the number density or luminosity profiles of galaxies in clusters. Our approach delivers extremely accurate results (completeness, C $\sim 92{\rm{ per\ cent}}$ and purity, P $\sim 87{\rm{ per\ cent}}$) within R200, so that we named our code reliable photometric membership. We discuss possible dependencies on magnitude, colour, and cluster mass. Finally, we present some applications of our method, stressing its impact to galaxy evolution and cosmological studies based on future large-scale surveys, such as eROSITA, EUCLID, and LSST.

Lopes, Paulo A. A.↗