Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient boosting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Machine Learning Benchmarks for the Classification of Equivalent Circuit Models from Electrochemical Impedance Spectra

Analysis of Electrochemical Impedance Spectroscopy (EIS) data for electrochemical systems often consists of defining an Equivalent Circuit Model (ECM) using expert knowledge and then optimizing the model parameters to deconvolute various resistance, capacitive, inductive, or diffusion responses. For small data sets, this procedure can be conducted manually; however, it is not feasible to manually define a proper ECM for extensive data sets with a wide range of EIS responses. Automatic identification of an ECM would substantially accelerate the analysis of large sets of EIS data. We showcase machine learning methods to classify the ECMs of 9,300 impedance spectra provided by QuantumScape for the BatteryDEV hackathon. The best-performing approach is a gradient-boosted tree model utilizing a library to automatically generate features, followed by a random forest model using the raw spectral data. A convolutional neural network using boolean images of Nyquist representations is presented as an alternative, although it achieves a lower accuracy. We publish the data and open source the associated code. The approaches described in this article can serve as benchmarks for further studies. A key remaining challenge is the identifiability of the labels, underlined by the model performances and the comparison of misclassified spectra.

25 ENERGY STORAGE↗

PyTREES

PyTREES (Python tool for Training/Testing Robust Explainable Ensembles on Spectra) is software that implements a data-driven approach to predicting the amount of specific oxides present in materials samples of laser-induced breakdown spectroscopy (LIBS); such as from the ChemCam instrument suite onboard the NASA Curiosity rover. PyTREES is designed to input LIBS data in the format provided by the ChemCam team [1]. PyTREES then applies appropriate pre-processing to this data [2], and implements several regression methods for predicting oxides from spectra. The regression methods include: ensemble methods (random forest, extra trees, and gradient boosting regression) and blended submodels using the “double blending” technique. PyTREES additionally implements methods for quantifying the importance of features in regression model: (1) mean decrease in impurity (MDI) and (2) permutation importance to investigate the wavelengths used by the regression methods. [1] Gasda et al. (2021). Spectrochim Acta B, 181, 106223. [2] Clegg et al. (2017). Spectrochim Acta B , 129, 64–85.

Oyen, Diane↗

PV Generation and Load Forecasting for Adjuntas PR Community Microgrids

Existing frameworks to forecast time-series photovoltaic (PV) output power and consumer load for microgrid operations and controls assume a near-continuous availability of real-time input features from the field assets such as PV inverters, energy meters, and weather station. These incoming data points are used to periodically retrain models and update forecast snapshots over a moving horizon window, be it one hour-ahead, one-day ahead, or one-week ahead. However, such frameworks are not resilient to disruptions in data availability caused by losses in communications between the field sensors and data loggers. Hence, there is a need for programs that assume no availability of real-time microgrid asset data and still make reliable forecasts that can be used for decision-making. Such programs would be apt to function in extreme weather events such as hurricanes and would use lightweight recursive time-series models to independently forecast solar irradiance and ambient temperature, then compute PV power from those forecasts, as well as independently forecast consumer load. The codebase performs forecasting for the scenario of when the microgrid does not have a reliable access to forecasts or real-time observations of solar irradiance (I) and ambient temperature (AT) and load (Load) to be able to adequately forecast, in real-time, the PV power production or a business' load. In this case, using historical values of PV power and load, a univariate forecasting of generation and consumption are respectively made. The use-case in particular has two sub-scenarios: one, a normal 7-day ahead forecast where the unavailability of real-time data is assumed due to infrastructure issues such as loss of communication or sensor maintenance or service downtimes. Whereas a hurricane-caused unavailability of real-time data requires a second model trained specifically on historical hurricane days to be able to capture the extreme day behavior of generation in particular, and load if applicable. A gradient boosted regression tree comprises an ensemble of additive models that map between the input of historical values (be it irradiance, temperature, or load) and their corresponding output forecasts of a given horizon such that the individual learner predictions are summed up over the total number of such learners in the ensemble to produce an aggregate forecast. A weighting mechanism is applied to the training data in each iteration, where actual and forecast values are compared to penalize incorrect forecasts by increasing the weight and reducing it to reward correct forecasts. The code's benefits are that it: (a) accounts for a contingency where communication loss renders newly measured real-time data unavailable for model tuning and snapshot updates; (b) presents blind forecasting that recursively determines the next time-step value in a horizon using the forecast of the same attribute from a prior step; and (c) employs lightweight models that, once trained, can reliably generalize for different horizons, which make them suitable for enhancing the resilience of field microgrids prone to extreme events that encounter disruptions to data availability.

Sundararajan, Aditya [Oak Ridge National Laborator↗

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

fSCAml

A repository for predicting fractional snow covered area using a gradient boosted trees machine learning.

Crumley, Ryan L [Self-Employed]↗

Reducing Southern Ocean Shortwave Radiation Errors in the ERA5 Reanalysis with Machine Learning and 25 Years of Surface Observations

Earth system models struggle to simulate clouds and their radiative effects over the Southern Ocean, partly due to a lack of measurements and targeted cloud microphysics knowledge. We have evaluated biases of downwelling shortwave radiation in the ERA5 climate reanalysis using 25 years (1995–2019) of summertime surface measurements, collected on the Research and Supply Vessel (RSV) Aurora Australis, the Research Vessel (R/V) Investigator, and at Macquarie Island. During October–March daylight hours, the ERA5 simulation of SW down exhibited large errors (mean bias = 54 W m -2 , mean absolute error = 82 W m -2 , root-mean-square error = 132 W m -2 , and R 2 = 0.71). To determine whether we could improve these statistics, we bypassed ERA5’s radiative transfer model for SW down with machine learning–based models using a number of ERA5’s gridscale meteorological variables as predictors. These models were trained and tested with the surface measurements of SW down using a 10-fold shuffle split. An extreme gradient boosting (XGBoost) and a random forest–based model setup had the best performance relative to ERA5, both with a near complete reduction of the mean bias error, a decrease in the mean absolute error and root-mean-square error by 25% ± 3%, and an increase in the R 2 value of 5% ± 1% over the 10 splits. Large improvements occurred at higher latitudes and cyclone cold sectors, where ERA5 performed most poorly. We further interpret our methods using Shapley additive explanations. Our results indicate that data-driven techniques could have an important role in simulating surface radiation fluxes and in improving reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

Estimating Fine-Resolution Shortwave Broadband Albedo of Croplands from Harmonized Landsat and Sentinel-2 Data

Altered surface albedo due to land-cover conversions and management is a significant driver of global climate change. Albedo can be directly measured at ground stations, and remote sensing data can be used to scale-up albedo values to regional and global levels. Some previous studies have retrieved fine-resolution (10–30 m) instantaneous albedo and coarse-resolution (500–1000 m) daily mean albedo from remote sensing data, but they all required the input of Moderate Resolution Imaging Spectroradiometer (MODIS) albedo information at 500-m resolution, and none have assembled both instantaneous and daily albedo based exclusively on fine-resolution satellite data. Here, to address this issue, we compiled 387 instantaneous and 346 daily albedo records using field net radiometer measurements from the bioenergy croplands at the W. K. Kellogg Biological Station in southwest Michigan. We then connected these albedo records with a suite of variables derived from harmonized Landsat and Sentinel-2 data through two machine learning algorithms (random forest regression and extreme gradient boosting) to retrieve clear-sky instantaneous and daily shortwave broadband albedo. The performance statistics indicate reasonable accuracy of model results [root-mean-square error (RMSE)] around or below 0.03 except for snow-covered surfaces), suggesting that the retrieval of both instantaneous and daily albedo based exclusively on fine-resolution satellite data is promising. To facilitate the use of fine-resolution albedo products at the global level, future efforts need to include more albedo records of diverse surface cover types, as well as to accurately model daily albedo for cloudy days to address the “clear-sky bias.”

Harmonized Landsat and Sentinel-2↗

Learning curves for drug response prediction in cancer cell lines

Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.

60 APPLIED LIFE SCIENCES↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

Mitigating Algorithmic Bias in Cancer Site Classification Models

Purpose Integrating artificial intelligence in cancer diagnostics has improved tumor classification beyond rule-based systems. Despite these advancements, these models may still encode demographic biases. We conducted a large-scale, applied bias-probing study of a deep learning–based cancer site classifier to quantify race information encoded in document embeddings. We then evaluated how performance changes when race-correlated embedding dimensions are removed in a post-training sensitivity analysis. Methods The cancer site classifier was trained using 3.5 million electronic cancer pathology reports from six of the National Cancer Institute's SEER registries. We trained a hierarchical self-attention network to generate 400-dimensional document embeddings. These embeddings were used to train two downstream, gradient-boosted decision tree classifiers: one to classify the cancer sites and another to predict racial categories. We identified overlapping features by intersecting the top 50 feature-importance rankings from the site and race models and computed their cumulative feature importance in each model. As a post hoc sensitivity analysis, we progressively pruned these overlapping dimensions, retrained the site model, and compared overall macro-F1 and accuracy, race-stratified macro-F1, and group fairness metrics on the basis of demographic parity and equalized odds before and after pruning. Results The analysis revealed minimal feature overlap between the cancer site and race prediction models, and the cumulative importance scores indicated a negligible influence of racial information on clinical predictions. Post-training pruning of overlapping features did not compromise the models' diagnostic accuracy, with a 0.07% loss in accuracy. Conclusion Our findings demonstrate that HiSAN-generated embeddings from SEER data can be used effectively in cancer site classification without significant demographic bias influencing the outcomes. Post-training pruning therefore functions as a practical audit and sensitivity check.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

UNNT: A novel Utility for comparing Neural Net and Tree-based models

The use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Dataset for 'Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning', Water 2022

This data package presents forcing data, model code, and model output for classical machine learning models that predict monthly stream water temperature as presented in the manuscript ‘Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning’, Water (Weierbach et al., 2022). Specifically, for input forcing datasets we include two files each generated using the BASIN-3D data integration tool (Varadharajan et al., 2022) for stations in the Pacific Northwest and Mid Atlantic Hydrologic regions. Model code (written in python with the use of jupyter notebooks) includes codes for data preprocessing, training Multiple Linear Regression, Support Vector Regression, and Extreme Gradient Boosted Tree models, and additional notebooks for analysis of model output. We include specific model output files which represent modeling configurations presented in the manuscript also presented in an hdf5 format. Together, these data make up the workflow for predictions across three scenarios (single station, regional, and predictions in unmonitored basins) presented in the manuscript and allow for reproducibility of modeling procedures.

54 ENVIRONMENTAL SCIENCES↗

Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification" Willard et al. (2025).

This data release provides all data and code used in the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025)" to model stream temperature, evaluate, and assess results. The associated manuscript explores the effect of different ensemble construction techniques across different common machine learning (ML) architectures for predictions in unmonitored basins. Modeling was done using long short-term memory (LSTM), gated recurrent unit (GRU), temporal convolution network (TCN), and extreme gradient boosting (XGBoost) models, and stream site coverage spans 1362 locations across the conterminous United States. The ensemble construction techniques investigated include ensemble by random weight initialization, differing hyperparameters, different random subsets of training data, different subselections of input features, different architectures, and Monte Carlo Dropout. The data is organized into these items items:Code repository and data for the paper " "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantifications" Willard et al. (2025).Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code:- data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repositoryData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2025streamensembles,author = {Jared Willard and Charuleka Varadharajan},title = {Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification"},year = {2024},doi = {10.15485/2527393},publisher = {ESS-DIVE Repository},url = {https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2527393}}MLA: Willard, Jared, et al. Dataset for "Machine Learning Ensembles Can Enhance Hydrologic Predictions and Uncertainty Quantification". 2025. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Reference data, predictors, and probability grids for forest degradation classes in three sites in the Brazilian Amazon

Forest degradation by fires and selective logging is widespread in the Amazon region. We implemented a gradient boosted classification modeling framework to classify intact, logged, and burned forests at three Amazonian sites: Feliz Natal Municipality and Xingu Indigenous Territory in Mato Grosso State, and Saracá-Taquera National Forest in Pará State. We used forest degradation history from Landsat time-series as reference data and textural metrics derived from PlanetScope images as predictors. Textural metrics were computed using the Gray-Level Co-Occurrence Matrix (GLCM) textural technique. Included in the attached zip file are ten files: - a shapefile containing the reference data (fire and selective logging polygons and year of event) for each site; - a multiband tif file containing the 8 GLCM metrics used as predictors (Mean, Variance, Homogeneity, Contrast, Dissimilarity, Entropy, Angular Second Moment, Correlation) at the original PlanetScope resolution (3.125m) for each site; - a multiband tif file containing the 72 aggregated GLCM metrics used as predictors (Mean, Variance, Homogeneity, Contrast, Dissimilarity, Entropy, Angular Second Moment, and Correlation aggregated using the mean, first quartile, third quartile, maximum, median, minimum, root mean square, standard deviation, and skewness statistics) at 562m resolution for each site; - a multiband tif file containing the 3 probability grids for either intact, logged, or burned forests at the aggregation resolution (562m) for each site.

54 ENVIRONMENTAL SCIENCES↗

Integrating Experiments and Well Logs to Predict Caney Shale Static Mechanical Properties During Production with Supervised Machine Learning

Caney shale is one of the emerging oil reservoirs in Oklahoma. Understanding the impact of effective stress on its mechanical properties is critical for predicting hydraulic fracture geometry and overall hydrocarbon production. The objective of our study is to evaluate the impact of effective stress on the dynamic Young’s modulus using ultrasonic velocity measurements for Caney shale samples. A triaxial cell was utilized to measure ultrasonic (P-wave and S-wave) velocities for ten downhole Caney shale samples under various effective stresses to indirectly assess the impact of pore pressure change. The dynamic Young’s moduli estimated from these measurements were integrated with available conventional well logs (excluding sonic logs) and triaxial test results from Benge et al. (2021) to predict the static Young’s modulus using Random Forest (RF) and Extreme Gradient Boosting (XGBoost) models. The results showed that the estimated dynamic Young’s moduli from ultrasonic measurements were higher than the corresponding static Young’s modulus of cores from the same vertical well at similar depths. With increasing effective stress, the dynamic Young’s modulus increased for all samples. The estimated dynamic-to-static correction factor tended to be higher in zones of high neutron porosity (PHIN) and low density compared to other zones. Finally, SHapley Additive exPlanations (SHAP) for RF and XGBoost models identified depth, gamma ray (GR), and PHIN as key features for predicting the static Young’s modulus. This study enhances our understanding of the dynamic and static Young’s moduli for the Caney shale interval, as a function of effective stress and conventional well logs. The findings from this study can improve predictions of production throughout the well's lifespan by offering insights into the mechanical property degradation resulting from pore pressure depletion.

Kholy, Sherif M.↗

Simultaneous prediction of structural properties in epitaxially–grown GaN with quantum and conventional multi–output learning algorithms

Hundreds of GaN thin film crystal plasma–assisted molecular beam epitaxy synthesis experiment records spanning two decades were organized into a dataset correlating the growth experiment design parameters with discrete, binary determinations of crystallinity and surface morphology. Conventional data science techniques as well as both quantum and classical multi–output supervised machine learning algorithms were implemented to investigate the relationships between the operating parameter data and the structural figures of merit. Correlation coefficients, decision tree nodes, p–values, and SHAP values all support substrate temperature and gallium effusion cell conditions as being statistically significant for simultaneously influencing GaN crystallinity and surface morphology. Here, a conventional deep neural network learned best from the data, followed by a quantum–classical hybrid gradient boosting algorithm. When combined with calculations of uncertainty intervals based on VennAbers predictors, machine learning predictions of both structural properties show good agreement with results reported in published experimental literature.

36 MATERIALS SCIENCE↗

Monitoring Operational States of a Nuclear Reactor Using Seismoacoustic Signatures and Machine Learning

Monitoring nuclear reactors is an important safety and security task with growing requirements. We explore the possibility of using seismic and acoustic data for inferring the power level of an operating reactor. Continuous data recorded at a single seismoacoustic station that is located about 50 m away from a research reactor was visualized and analyzed. The data show a clear correlation between seismoacoustic features and reactor main operational states. We designed a workflow that includes two machine learning (ML) models to classify the reactor operational states (OFF, transition, and ON) and estimate reactor power levels (10%, 30%, 50%, 70%, and 90%). We applied and compared five ML algorithms for the reactor OFF-transition-ON and four approaches for the power level classification. We also compared the performance of ML models trained with seismic-only, acoustic-only, and both types of data. Five-fold cross validations were implemented to assure a thorough evaluation of the model performances. Additionally, the results show the extreme boosting gradient algorithm worked best for the first model, whereas random forests performed best for the second model. Combining seismic and acoustic data leads to better performance than using a single type of data. Seismic data contributed more than acoustic data for both models. We reached an accuracy of 0.98 for reactor OFF and ON. The accuracies for the transition state and power levels are less optimal with a minimum accuracy of 0.66. However, our results suggest seismic and acoustic data contain useful information about the transition state as well as power levels. Seismic and acoustic data could be integrated with other observations to improve monitoring performance.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗