Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linear regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Predicting seismic amplitudes with machine learning

The accurate estimation of seismic wave amplitude is vital to precisely determine the yield, magnitude, and event discrimination possible for a given network – a critical element in nuclear explosion monitoring. This task is complicated by several factors, including but not limited to radiation pattern, scattering effects, and crustal variations, which can lead to the attenuation or amplification of amplitude along a given raypath. In this report, we explore the novel application of machine learning to the task of seismic amplitude estimation by training a simple Artificial Neural Network (ANN) on an S-wave amplitude dataset from Lai et al. (2019). Attributes from this dataset used as input to the ANN included event-station distances, station locations (latitude, longitude), event locations (latitude, longitude), event depths, event magnitudes, radiation patterns, signal-to noise ratio (SNR) measurements (average-amplitude, peak-to-trough, maximum peak), and signal periods. We find that the trained ANN predicts S-wave amplitudes with a modest tendency toward underestimating the actual values, as indicated by a linear regression between predicted and actual data (slope: 0.892, intercept: -0.651). These results suggest that an ANN can perform this task, with potential for significant improvements through improved datasets, architectures, and parameter tuning.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Comparative Analysis of HEATNETS for Geothermal Network Performance

Thermal energy networks (TENs), also known as 5th generation district energy systems, or more specifically geothermal networks when exchanging heat with geothermal boreholes, are an important technology for decarbonization. In these networks an ambient loop connects buildings and thermal sources, such as a borehole field, to exchange energy and maintain a desired loop temperature. Water-source heat pumps are used at the buildings to connect to the ambient or thermal loop to meet to the building heating and cooling loads and maintain comfort. A semi-transient, reduced-order technical model and techno-economic model, called HEATNETS, has been developed at NREL that captures the flow of energy around a TEN. In this work, a comparison of the HEATNETS technical model and a well-known coding platform used for modeling geothermal networks, TRNSYS, has been completed for a proposed geothermal network as a verification and validation process. Hourly data provided from the TRNSYS simulation included building loads, pumping power, heat pump power, temperature entering and leaving the borehole field, and mass flow rates. The hourly borehole temperatures were used to create a linear regression model utilized in HEATNETS to estimate the borehole field heat exchange. The building loads and mass flow rates were direct inputs to HEATNETS while the pumping power, heat pump power, borehole temperatures, and coefficients of performance were all simulated and calculated by HEATNETS, allowing for direct comparison of the thermal energy transfer HEATNETS considers the full process from design inputs to economic outputs and can provide modeling options for high-level initial system design and operational optimization. This study focuses on a validation of HEATNETS using results from TRNSYS. HEATNETS is not intended to replace other modeling tools, but this work demonstrates, via a comparison with an industry standard code, that HEATNETS can be a unique, high-level and rapid modeling tool for estimating the performance of a full geothermal network system.

15 GEOTHERMAL ENERGY↗

Machine Learning Analysis of Temperature-Strain Relationships for Structural Health Monitoring of Pipes: Self-powered wireless sensor system for health monitoring of liquid-sodium cooled fast reactors

This report presents machine learning (ML) analysis of temperature-strain relationships for structural health monitoring of nuclear reactor stainless steel (SS) pipes with the strain gauge sensor directly printed on the pipe with a 3D conformal aerosol jet printer. We investigate correlations for two sensor pairs installed on the same SS304 pipe: commercial K-type thermocouple with a printed gold strain gauge (TC3-SG3), and commercial K-type thermocouple with commercial Kyowa strain gauge (TC0-SG0). The temperature ranges for the sensor pairs TC0-SG0 and TC3-SG3 are 20.00°C to 266.37°C and 39.95°C to 219.28°C respectively. ML algorithms in this study include Linear Regression (baseline method), Ridge Regression, Lasso Regression, and Gradient Boosting. Performance evaluation metrics include Root Mean Square Error (RMSE), Mean Square Error (MSE), Mean Absolute Error (MAE), R 2 Score, and Explained Variance. Using advanced feature engineering techniques, we extracted 27 temperature-based features and 30 strategic inclusion features. The best performance was obtained with the Gradient Boosting method, which achieves prediction accuracy of R 2 = 0.9999 and RMSE = 7.69 μStrain for TC0-SG0, and R 2 = 0.9998 and RMSE = 18.03 μStrain for TC3-SG3. While the temperature-strain correlations are weaker for the gauge directly printed on the pipe than for the commercial strain gauge, deployment-ready performance exceeding industry standards is achieved for both sensor pairs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Prediction of Dielectric Constant in Series of Polymers by Quantitative Structure-Property Relationship (QSPR)

This work is devoted to the investigation of dielectric permittivity which is influenced by electronic, ionic, and dipolar polarization mechanisms, contributing to the material’s capacity to store electrical energy. In this study, an extended dataset of 86 polymers was analyzed, and two quantitative structure–property relationship (QSPR) models were developed to predict dielectric permittivity. From an initial set of 1273 descriptors, the most relevant ones were selected using a genetic algorithm, and machine learning models were built using the Gradient Boosting Regressor (GBR). In contrast to Multiple Linear Regression (MLR)- and Partial Least Squares (PLS)-based models, the gradient boosting models excel in handling nonlinear relationships and multicollinearity, iteratively optimizing decision trees to improve accuracy without overfitting. The developed GBR models showed high R2 coefficients of 0.938 and 0.822, for the training and test sets, respectively. An Accumulated Local Effect (ALE) technique was applied to assess the relationship between the selected descriptors—eight for the GB_A model and six for the GB_B model, and their impact on target property. ALE analysis revealed that descriptors such as TDB09m had a strong positive effect on permittivity, while MLOGP2 showed a negative effect. These results highlight the effectiveness of the GBR approach in predicting the dielectric properties of polymers, offering improved accuracy and interpretability.

Ascencio-Medina, Estefania↗

Wireless Patch Antenna Characterization for Live Health Monitoring Using Machine Learning

Temperature monitoring in extreme environments, such as coal-fired power plants, was addressed by designing and testing wireless patch antennas for use in machine learning-aided temperature estimation. The sensors were designed to monitor the temperature and health of boiler systems. Wireless interrogation of the sensor was performed using a Vector Network Analyzer (VNA) and a pair of interrogation antennas to capture resonance behavior under varying thermal and spatial conditions with sensitivities ranging from 0.052 to 0.20 $\frac{𝑀𝐻𝑧}{°C}$. Sensor calibration was conducted using a Long Short-Term Memory (LSTM) model, which leveraged temporal patterns to account for hysteresis effects. The calibration method demonstrated improved performance when combined with an LSTM model, achieving up to a 76% improvement in temperature estimation error when compared with Linear Regression (LR). The experiments highlighted an innovative solution for patch antenna-based non-contact temperature measurement, which addresses limitations with conventional methods such as RFID-based systems, infrared, and thermocouples.

20 FOSSIL-FUELED POWER PLANTS↗

Position-specific kinetic isotope effects for nitrous oxide: a new expansion of the Rayleigh model

Nitrous oxide (N 2 O) is a potent greenhouse gas and the most significant anthropogenic ozone-depleting substance currently being emitted. A major source of anthropogenic N 2 O emissions is the microbial conversion of fixed nitrogen species from fertilizers in agricultural soils. Thus, understanding the enzymatic mechanisms by which microbes produce N 2 O has environmental significance. Measurement of the 15 N/ 14 N isotope ratios of N 2 O produced by purified enzymes or axenic microbial cultures is a promising technique for studying N 2 O biosynthesis. Typically, N 2 O-producing enzymes combine nitrogen atoms from two identical substrate molecules (NO or NH 2 OH). Position-specific isotope analysis of the central (N α ) and outer (N β ) nitrogen atoms in N 2 O enables the determination of the individual kinetic isotope effects (KIEs) for N α and N β , providing mechanistic insight into the incorporation of each nitrogen atom. Previously, position-specific KIEs (and fractionation factors) were quantified using the Rayleigh distillation equation, i.e., via linear regression of δ 15 N α or δ 15 N β against [–f In f / (1 – f)], where f is the fraction of substrate remaining in a closed system. This approach, however, is inaccurate for N α and N β because it does not account for fractionation at N α affecting the isotopic composition of substrate available for incorporation into the β position (and vice versa). Therefore, we developed a new expansion of the Rayleigh model that includes specific terms for fractionation at the individual N 2 O nitrogen atoms. By applying this Expanded Rayleigh model to a variety of simulated N 2 O synthesis reactions with different combinations of normal, inverse, and/or no KIEs at N α and N β , we demonstrate that our new model is both accurate and robust. We also applied this new model to two previously published datasets describing N 2 O production from NH 2 OH oxidation in a methanotroph culture (Methylosinus trichosporium) and N 2 O production from NO by a purified Histoplasma capsulatum (fungal) P450 NOR, demonstrating that the Expanded Rayleigh model is a useful tool in calculating position-specific fractionation for N 2 O synthesis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On the predictability of turbulent fluxes from land: PLUMBER2 MIP experimental description and preliminary results

Accurate representation of the turbulent exchange of carbon, water, and heat between the land surface and the atmosphere is critical for modelling global energy, water, and carbon cycles in both future climate projections and weather forecasts. Evaluation of models' ability to do this is performed in a wide range of simulation environments, often without explicit consideration of the degree of observational constraint or uncertainty and typically without quantification of benchmark performance expectations. We describe a Model Intercomparison Project (MIP) that attempts to resolve these shortcomings, comparing the surface turbulent heat flux predictions of around 20 different land models provided with in situ meteorological forcing evaluated with measured surface fluxes using quality-controlled data from 170 eddy-covariance-based flux tower sites. Predictions from seven out-of-sample empirical models are used to quantify the information available to land models in their forcing data and so the potential for land model performance improvement. Sites with unusual behaviour, complicated processes, poor data quality, or uncommon flux magnitude are more difficult to predict for both mechanistic and empirical models, providing a means of fairer assessment of land model performance. When examining observational uncertainty, model performance does not appear to improve in low-turbulence periods or with energy-balance-corrected flux tower data, and indeed some results raise questions about whether the energy balance correction process itself is appropriate. In all cases the results are broadly consistent, with simple out-of-sample empirical models, including linear regression, comfortably outperforming mechanistic land models. In all but two cases, latent heat flux and net ecosystem exchange of CO 2 are better predicted by land models than sensible heat flux, despite it seeming to have fewer physical controlling processes. Land models that are implemented in Earth system models also appear to perform notably better than stand-alone ecosystem (including demographic) models, at least in terms of the fluxes examined here. The approach we outline enables isolation of the locations and conditions under which model developers can know that a land model can improve, allowing information pathways and discrete parameterisations in models to be identified and targeted for future model development.

54 ENVIRONMENTAL SCIENCES↗

The OCEAN ICE mooring compilation: a standardised, pan-Antarctic database of ocean hydrography and current time series

Continuous moored time series of temperature, salinity, pressure and current speed and direction are of great importance for understanding the continental shelf and under-ice-shelf dynamics and thermodynamics that govern water mass transformations and ice melting in and around Antarctic marginal seas. In these regions, icebergs and sea ice make ship-based mooring deployment and recovery challenging. Nevertheless, over decades, expeditions around the fringe of Antarctica sporadically deployed and recovered hundreds of moored instruments, including those facilitated through ice shelves boreholes. These datasets tend to be archived in a wide range of data centres, with, to our knowledge, no clear format standardisation. As a result, systematic analysis of historical mooring time series in the marginal seas is often challenging. Here we present the first version of a standardised pan-Antarctic moored hydrography and current time series compilation, with broad international contributions from data centres, research institutes and individual data owners. The mooring records in this compilation span over five decades, from the 1970s to the 2020s, providing an opportunity for a systematic study of the pan-Antarctic water mass transport and shelf connectivity. As a demonstration of the utility of this compilation, we present spectral analysis of the compiled current velocity time series, which unsurprisingly shows the dominating presence of tidal variability within most records. This component of the variability is fitted using multi-linear regression to tidal frequencies, and the tidal fit is removed from the original time series to leave de-tided variability. Given the limited record durations to months to years, de-tided variability is dominated by synoptic (3–10 d period), intraseasonal (10–80 d) and seasonal (∼6 months–1 year) signals. The spatial distribution of the kinetic energy integrated within frequency bands is presented and discussed within respective regional contexts, and future avenues of research are proposed. This data compilation is assembled under the endorsement of Ocean-Cryosphere Exchanges in ANtarctica: Impacts on Climate and the Earth System (OCEAN ICE) project (https://ocean-ice.eu/, last access: 23 October 2025) funded by the European Commission and UK Research and Innovation. It is available and regularly updated in NetCDF format with the SEANOE database at https://doi.org/10.17882/99922 (Zhou et al., 2024a).

54 ENVIRONMENTAL SCIENCES↗

Probing Signal-Based Inertia and Frequency Response Estimation for Power Systems with High Penetration of Inverter-Based Resources: Preprint

Power system inertia is the inherent capability of a power system to resist changes in its frequency during disturbances. Real-time inertia estimation technology has become more important due to the low-inertia issues caused by the increasing integration levels of inverter-based resources (IBRs) from renewable energy; however, existing inertia estimation methods hardly consider multiple frequency response controls that act within the same time frame as conventional inertial response, thus making measured inertia values vary under different testing conditions. To resolve this issue, this paper proposes a novel real-time estimation method to simultaneously estimate a power system's inertia constant and frequency response droop constant using a well-designed probing signal. First, we formulate the inertia and frequency response model of a power system with IBRs. Second, through the integration and manipulation of the developed model, we propose a multivariate linear regression- based estimation method that is resilient to measurement noise. Third, we design a probing signal that can be injected by IBRs to incite the required transients for estimation. Finally, we validate the proposed estimation method through comprehensive power- hardware-in-the-loop experiments using inverter hardware and a realistic island power system model. The results demonstrate that the proposed method can accurately estimate the inertia and droop value of the power system with grid-following IBRs and grid-forming IBRs with virtual synchronous machine control.

frequency response↗

Comparative Analysis of HEATNETS for Geothermal Network Performance: Preprint

Thermal energy networks (TENs), also known as 5th generation district energy systems, or more specifically geothermal networks when exchanging heat with geothermal boreholes, are an important technology for decarbonization. In these networks an ambient loop connects buildings and thermal sources, such as a borehole field, to exchange energy and maintain a desired loop temperature. Water-source heat pumps are used at the buildings to connect to the ambient or thermal loop to meet to the building heating and cooling loads and maintain comfort. A semi-transient, reduced-order technical model and techno-economic model, called HEATNETS, has been developed at NREL that captures the flow of energy around a TEN. In this work, a comparison of the HEATNETS technical model and a well-known coding platform used for modeling geothermal networks, TRNSYS, has been completed for a proposed geothermal network as a verification process. Hourly data provided from the TRNSYS simulation included building loads, pumping power, heat pump power, temperature entering and leaving the borehole field, and mass flow rates. The hourly borehole temperatures were used to create a linear regression model utilized in HEATNETS to estimate the borehole field heat exchange. The building loads and mass flow rates were direct inputs to HEATNETS while the pumping power, heat pump power, borehole temperatures, and coefficients of performance were all simulated and calculated by HEATNETS, allowing for direct comparison of the thermal energy transfer, rather than also comparing control systems responses. HEATNETS considers the full process from design inputs to economic outputs and can provide modeling options for high-level initial system design and operational optimization. This study shows that HEATNETS, while not intended to replace other modeling tools, can be a unique modeling tool for the performance of a full geothermal network system.

15 GEOTHERMAL ENERGY↗

Reduce-Order Modeling of Multigroup Neutron Cross Sections for High-Temperature Gas-cooled Reactors

Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which usually consists of a database of tabulated values, used to calculate the cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of micro cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. To address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multi-group cross section data across isotopes, reaction types and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs for have been trained for all isotopes in this work and systematic Griffin testing is ongoing at this moment to ensure the feasibility of this ROM technique for cross section predictions.

42 - ENGINEERING↗

Reduced-Order Modeling of Multigroup Neutron Cross Sections for High-Temperature Gas-cooled Reactors

Abstract – Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which consist of databases of tabulated values, used to calculate the neutron cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of microscopic cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. In order to address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multigroup cross section data across isotopes, reaction types, and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs have been trained for all isotopes in this work and systematic Griffin testing is ongoing to ensure the feasibility of this ROM technique for predicting cross section and reducing memory requirements without a significant sacrifice in computational performance.

42 - ENGINEERING↗

Advanced Cross Section Library Generation using Reduced Order Models

Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which consist of databases of tabulated values, used to calculate the neutron cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of microscopic cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. In order to address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multigroup cross section data across isotopes, reaction types, and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs have been trained for all isotopes in this work and systematic Griffin testing is ongoing to ensure the feasibility of this ROM technique for predicting cross section and reducing memory requirements without a significant sacrifice in computational performance.

42 - ENGINEERING↗

DESI Data Release 2 ELGs: Property-dependent subsamples, imaging systematics, and clustering

Using emission-line galaxies (ELGs) from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2, we evaluate a property-dependent correction to imaging systematics. We derive systematic weights following the same linear regression method used for other DESI tracers, but do so separately on ELG subsamples to provide a physically-informed alternative to the fiducial, neural-network-based approach. In doing so, we show that the deeper imaging in the Dark Energy Survey (DES) footprint leads to a higher overall number density but a lack of targets with extreme $g-r$ and $r-z$ colors. ELGs in the DES region also show a distinct redshift distribution when subsampled by position in the $g-r$ vs. $r-z$ plane. To address these effects, we implement a separate treatment of the DES footprint within the DESI catalog production pipeline, which is generally well-motivated and, in some cases, imperative for accurate clustering measurements. With DES treated separately, we find that property-dependent systematic weights further mitigate spurious clustering signal in $\sim$10% of subsamples, while the fiducial scheme remains optimal for the full sample.

Hagen, T. [Utah U.]↗

Unveiling the drivers contributing to global wheat yield shocks through quantile regression

Sudden reductions in crop yield (i.e., yield shocks) severely disrupt the food supply, intensify food insecurity, depress farmers' welfare, and worsen a country's economic conditions. Here, we study the spatiotemporal patterns of wheat yield shocks, quantified by the lower quantiles of yield fluctuations, in 86 countries over 30 years. Furthermore, we assess the relationships between shocks and their key ecological and socioeconomic drivers using quantile regression based on statistical (linear quantile mixed model) and machine learning (quantile random forest) models. Using a panel dataset that captures spatiotemporal patterns of yield shocks and possible drivers in 86 countries, we find that the severity of yield shocks has been increasing globally since 1997. Moreover, our cross-validation exercise shows that quantile random forest outperforms the linear quantile regression model. Despite this performance difference, both models consistently reveal that the severity of shocks is associated with higher weather stress, nitrogen fertilizer application rate, and gross domestic product (GDP) per capita (a typical indicator for economic and technological advancement in a country). While the unexpected negative association between more severe wheat yield shocks and higher fertilizer application rate and GDP per capita does not imply a direct causal effect, they indicate that the advancement in wheat production has been primarily on achieving higher yields and less on lowering the possibility and magnitude of sharp yield reductions. Hence, in the context of growing extreme weather stress, there is a critical need to enhance the technology and management practices that mitigate yield shocks to improve the resilience of the world food systems.

60 APPLIED LIFE SCIENCES↗

Protocol to detect dilution cycles in chemostat experiments and estimate growth rate slopes with linear modeling with R software chemostat_regression

Chemostat growth chambers measure optical density over time and require manual calculation of growth rates. Here, we present chemostat_regression, R software that enables users to automatically identify chemostat cycles and estimate growth rate using a linear regression approach. We describe steps for creating requisite software environment(s), formatting input data, executing the software via command line/RStudio/R-Shiny, interpreting results, assessing the validity of results, and modifying input parameters.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning Accelerated First-Principles Study of the Hydrodeoxygenation of Propanoic Acid

The complex reaction network of catalytic biomass conversions often involves hundreds of surface intermediates and thousands of reaction steps, greatly hindering the rational design of metal catalysts for these conversions. Here, we present a framework of machine learning (ML)-accelerated first-principles studies for the hydrodeoxygenation (HDO) of propanoic acid over transition metal surfaces. The microkinetic model (MKM) is initially parametrized by ML-predicted energies and iteratively improved by identifying the rate-determining species and steps (RDS), computing their energies by density functional theory (DFT), and reparameterizing the MKM until all the RDS are computed by DFT. The Gaussian process (GP) model performs significantly better than the linear ridge regression model for predicting both the adsorption free energies and transition state free energies. Parameterized with energies from the GP model, only 5–20% of the full reaction network has to be computed by DFT for the MKM to possess DFT-level accuracy for the TOF and dominant reaction pathway. While the linear ridge regression model performs worse than the GP model, its performance is greatly improved when only transition states are predicted by the regression model and adsorption energies are computed by DFT. Overall, we find that a high accuracy in adsorption free energies is more important for a reliable MKM than a high accuracy in TS free energies. Lastly, based on the GP model with GOH and GCHCHCO as catalyst descriptors, we build two-dimensional volcano plots in activity and selectivity that can help design promising alloy catalysts for HDO reactions of organic acids.

adsorption↗

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗