Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Logistic Regression in Clinical Studies

• A logistic regression model is used when the outcome of interest is binary. The term “logistic” refers to the underlying “logit” (log odds) function that is used to model the binary outcome. • Odds ratios are produced from a logistic regression model and have a useful interpretation. • Tips, tricks and concepts used to fit logistic regression models are similar to those used in linear regression models. • Modeling building that is knowledge-based rather than automatic is preferred in most applications of logistic regression. • A logistic regression model that is overparameterized (ie, too many variables for too few events) can result in odds ratios that are implausibly large and confidence intervals that are wide and uninterpretable. These types of “overfitted” models should be avoided. • Logistic regression models can be fit using most standard statistical software.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

2021 Monthly Rice Production in Chinese Coastal Provinces

This paper explores the dynamics of rice production in the Chinese provinces of Liaoning, Jilin, Heilongjiang, Shanghai, Jiangsu, and Zhejiang and seeks to predict monthly rice production in the months of April through October using precipitation and Normalized Difference Vegetation Index as the predictor variables available. We utilize ridge and lasso regression models to predict the rice yield. Results indicate that a lasso regression model with an R 2 value of 0.9991501 with an adjusted R 2 value of 0.9991502 and a ridge regression model with an R 2 value of 0.9865443 with an adjusted R 2 value of 0.9870637 are possible. The lasso regression model does not account for all predictor variables while the ridge regression model does. Both models could be expanded upon to include more observations.

Tubbs, Heidi↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

Functional Predictor Variables for the Leaching Potential of Arsenic and Selenium from Coal Fly Ash

The release of leachates from intact coal ash impoundments is a concern due to the enrichment and mobilization of toxic elements such as arsenic (As) and selenium (Se). This study aims to explore the intrinsic properties of coal fly ash that correlate with the relative leachability of As and Se. We performed leaching experiments with 52 fly ash samples collected from 15 different U.S. power plants and representing coal feedstocks from the three major domestic regions. We assessed the mobilization potential of As and Se in fly ash based on standardized leaching protocols and performed multivariate and lasso regression analyses to explore correlations of leachable As and Se contents with characteristics such as major element contents, loss on ignition, and pH. The results of regression models indicated that major elements (Fe, Ca, and Al) for a wide range of fly ashes can serve as predictor variables for the leaching potential of As but not for Se. LOI and pH were not important predictive variables in the models. Both regression approaches resulted in relatively strong fits for leachable As (correlation coefficient R 2 = 0.78 for both models) compared to models for leachable Se (R 2 = 0.49). Overall, these results suggest that correlation models combined with on-site elemental analysis with portable analyzers may enable a screening method for leachable As in coal ash.

01 COAL, LIGNITE, AND PEAT↗

Machine Learning Accelerated First-Principles Study of the Hydrodeoxygenation of Propanoic Acid

The complex reaction network of catalytic biomass conversions often involves hundreds of surface intermediates and thousands of reaction steps, greatly hindering the rational design of metal catalysts for these conversions. Here, we present a framework of machine learning (ML)-accelerated first-principles studies for the hydrodeoxygenation (HDO) of propanoic acid over transition metal surfaces. The microkinetic model (MKM) is initially parametrized by ML-predicted energies and iteratively improved by identifying the rate-determining species and steps (RDS), computing their energies by density functional theory (DFT), and reparameterizing the MKM until all the RDS are computed by DFT. The Gaussian process (GP) model performs significantly better than the linear ridge regression model for predicting both the adsorption free energies and transition state free energies. Parameterized with energies from the GP model, only 5–20% of the full reaction network has to be computed by DFT for the MKM to possess DFT-level accuracy for the TOF and dominant reaction pathway. While the linear ridge regression model performs worse than the GP model, its performance is greatly improved when only transition states are predicted by the regression model and adsorption energies are computed by DFT. Overall, we find that a high accuracy in adsorption free energies is more important for a reliable MKM than a high accuracy in TS free energies. Lastly, based on the GP model with GOH and GCHCHCO as catalyst descriptors, we build two-dimensional volcano plots in activity and selectivity that can help design promising alloy catalysts for HDO reactions of organic acids.

adsorption↗

Machine learning-enabled prediction of chemical durability of A 2 B 2 O 7 pyrochlore and fluorite

Pyrochlore-structure type and its derivative in a general formula A 2 B 2 O 7 (A = rare earth elements and actinides; B = Ti, Sn, Zr, Hf, Pb, Si, etc.) display excellent structural flexibility and rich crystal chemistry as promising nuclear waste form materials capable of immobilizing actinides and fission products. It is essential to understand these materials’ chemical durability and element release of radionuclides in order to evaluate their performance in near-field environment. However, it is a formidable grand technological challenge to experimentally perform durability testing across hundreds of thousands of possibilities resulting from their extreme compositional complexities due to cation substitutions at both A and B-sites. In this work, we demonstrate a machine learning approach to determine the key materials parameters and structural characteristics governing the leaching behaviors from a small set of selected compositions as model systems, enabling a science-based prediction of their chemical durability that can be extended to a wide range of chemical compositions. The combination of four key structural characteristics and materials parameters, including ionic radius size difference , ionic potential difference , electronegativity difference , and lattice parameter , creates features an optimized prediction of the chemical durability. Two machine learning models, linear regression and Kernel ridge regression models, are trained on the randomly-split training dataset derived from the experimentally-determined elemental release rates, and subsequently tested on the testing dataset. The predicted leaching rates from both machine learning models show an excellent agreement with the experimental data, demonstrating the feasibility of rapidly evaluating the material properties of new compositions. These results highlight the immense potential of synergizing informatics through machine learning-based models and well-controlled experiments of selected model systems to accelerate materials design and discovery with optimized compositions and performance of promising materials for effective nuclear waste management.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Retrospective Analysis of Backwater Habitat Availability Using Remote Sensing

Low-velocity channel-margin habitats serve as important nursery habitats for the endangered Colorado pikeminnow (Ptychocheilus lucius) in the middle Green River between Jensen and Ouray, Utah. These habitats, known as backwaters, are associated with emergent sandbars, and are shaped and reformed annually by peak flows. Our recent knowledge about backwater characteristics and dynamics that is summarized in the synthesis report (Grippo et al. 2015) was based on detailed annual survey data from a relatively small sample of backwaters that were collected from 2003 to 2014 and reach-wide evaluations of backwater surface area based on manual interpretation of aerial and satellite imagery. Methods that bridge the gap between the detailed surveys from a small number of backwaters and the reach-wide assessment of their surface area would enable an assessment of the availability of backwater habitats that meet the minimum depth requirements for suitable habitat for Colorado pikeminnow. In 2015 Argonne National Laboratory (Argonne) tested three regression models—linear, multiple, and partial least square (PLS) regression models—for estimating backwater depth using National Agriculture Imagery Program (NAIP) imagery collected in July 2006 that covered the Jensen-Ouray reach of the Green River (Hamada and LaGory 2016). The results suggested that a PLS regression model showed high correlation with the reference depth (R 2 = 0.69) and had the most unbiased and consistent estimate of backwater depth. The results also indicated that the PLS model would provide reasonable estimates of the amount of habitat providing a minimum suitable depth of 30 cm for young-of-the-year Colorado pikeminnow, even though absolute depth estimates may be uncertain for backwater areas deeper than approximately 40 cm. The study also provided insights regarding the amount and the selection of calibration and validation (cal-val) data needed for improving the accuracy of depth prediction.

54 ENVIRONMENTAL SCIENCES↗

Optimization of Geometric Perturbations on a Rod Moving Through a High Explosive Target

After completing a study to ensure the simulation results were converged, several high resolution 3D Smoothed Particle Hydrodynamic (SPH) simulations of copper rods impacting a high explosive (LX14) target were performed. This was then formulated into an optimization problem: I wanted to find the optimum shape and location of a perturbation on the rod that would maximize its erosion after it left the target. The shape of the perturbation was modeled as a 2D Gaussian bump and parameterized by its location along the rod axis (z 0 ) and amplitude (A). The final mass of the coherent part of the rod as it leaves the target was used as a metric to represent the erosion of the rod, and the optimization was formulated to maximize this metric with respect to the aforementioned design variables. Due to the expensive nature of the high-fidelity 3D SPH simulations, a surrogate model needed to be chosen so that many function calls to the optimizer would be feasible. Thus, a strategic full factorial sampling plan was chosen to build a dataset, which consisted of 24 high-fidelity simulations. Two surrogate models, a third order polynomial regression model and a Gaussian Process Model, were analyzed using a 14%/86% test/train holdout technique. The root mean square and R2 score of the test set was used to determine the best model, and the third order polynomial regression model was chosen as the surrogate model. Finally, the Nelder-Mead Simplex and Basin-hopping optimization algorithms were implemented, and it was found that the two algorithms gave slightly different optimum values. Nelder-Mead gave an optimum point of [z* 0 ;A*] = [9:9;0:4] and Basin-Hopping gave an optimum value of x* = [z* 0 ;A*] = [9:2;0:1].

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Optimization of artificial viscosity in production codes based on Gaussian Regression surrogate models

To accurately model flows with shock waves using staggered-grid Lagrangian hydrodynamics, artificial viscosity has to be introduced to convert kinetic energy into internal energy, thereby increasing the entropy across shocks. Determining the appropriate strength of the artificial viscosity is an art and strongly depends on the particular problem and experience of the researcher. The objective of this study is to pose the problem of finding the appropriate strength of artificial viscosity as an optimization problem and solve this problem using machine learning (ML) tools, specifically using surrogate models based on Gaussian Process regression and Bayesian analysis. We describe the optimization method and discuss various practical details of its implementation. The shock-containing problems for which we apply this method all have been implemented in the LANL code FLAG. First, we apply ML to find optimal values to isolated shock problems of different strengths. Second, we apply ML to optimize viscosity for a 1D propagating detonation problem based on Zel’dovich-von Neumann-Doring (ZND) detonation theory using a reactive burn model. We compare results for default (currently used values in FLAG) and optimized values of artificial viscosity for these problems demonstrating the potential for significant improvement in the accuracy of computations.

42 ENGINEERING↗

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

Benefit Analysis of CO 2 Delivery Options for Offshore Storage or Enhanced Oil Recovery

The analysis presented in this report evaluates the benefits of CO₂ offshore transport via pipeline or ship within the GOM. It takes a top-down framework to estimate the costs. First, this analysis designed a reduced-order model (ROM) based on the cash flows in the FECM/NETL CO₂ Transport Cost Model (also known as CO2_T_COM). The ROM takes capital expenses (CAPEX) and operating expenses (OPEX) to calculate the CO₂ breakeven price based on the cash flows. Second, this analysis developed regression models utilizing published data from other analyses to estimate CAPEX and OPEX. Since the ROM is a simplified cash flow calculation, it is easy to exchange the core regression models to estimate various costs. The ROM and regression models provided a framework that can be easily used by other researchers, decision-makers, operators, and regulators. The objective of this analysis is to assess the CO₂ breakeven cost range for pipeline and ship transport of captured CO₂ given the CO₂ source and storage reservoir located in the GOM.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Distribution Grid Modeling Using Smart Meter Data

The knowledge of distribution grid models, including topologies and line impedances, is essential for grid monitoring, control and protection. However, such information is often unavailable, incomplete or outdated. The increasing deployment of smart meters (SMs) provides a unique opportunity to tackle this issue. This paper proposes a two-stage framework for distribution grid modeling using SM data. In the first stage, the network topology is identified by reconstructing a weighted Laplacian matrix of distribution networks. In the second stage, a least absolute deviations (LAD) regression model is developed for estimating line impedance of a single branch based on the nonlinear (inverse) power flow model, wherein a conductor library is leveraged to narrow down the solution space. The LAD regression model is originally a mixed-integer nonlinear program whose continuous relaxation is still non-convex. Furthermore, we specially address its convex relaxation and discuss the exactness. The modified regression model is then embedded within a bottom-up sweep algorithm to achieve the identification across the network in a branch-wise manner. Numerical results on the IEEE 13-bus, 37-bus and 69-bus test feeders validate the effectiveness of the proposed methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A machine learning approach for efficient multi-dimensional integration

Many physics problems involve integration in multi-dimensional space whose analytic solution is not available. The integrals can be evaluated using numerical integration methods, but it requires a large computational cost in some cases, so an efficient algorithm plays an important role in solving the physics problems. We propose a novel numerical multi-dimensional integration algorithm using machine learning (ML). After training a ML regression model to mimic a target integrand, the regression model is used to evaluate an approximation of the integral. Then, the difference between the approximation and the true answer is calculated to correct the bias in the approximation of the integral induced by ML prediction errors. Because of the bias correction, the final estimate of the integral is unbiased and has a statistically correct error estimation. Three ML models of multi-layer perceptron, gradient boosting decision tree, and Gaussian process regression algorithms are investigated. The performance of the proposed algorithm is demonstrated on six different families of integrands that typically appear in physics problems at various dimensions and integrand difficulties. The results show that, for the same total number of integrand evaluations, the new algorithm provides integral estimates with more than an order of magnitude smaller uncertainties than those of the VEGAS algorithm in most of the test cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Prediction of Creep-Induced Strain Using a Symbolic Regression-Based Model

Material creep under high-temperature conditions limits the lifetime and safety of structural systems such as advanced nuclear reactors. Conventional creep testing is slow and often produces inconsistent results across nominally identical experiments, making lifetime prediction uncertain. Here, to address these challenges, this work develops a data-driven symbolic regression (SR) model that consolidates results from duplicate creep tests and predicts the remaining strain-time curve of an ongoing experiment. The method uses piece-wise multi-objective SR with physical constraints to generate analytic, interpretable functions describing transient creep strain. Applied to Inconel Alloy 617 data, the approach achieved relative mean absolute errors of 1.0–9.5%, providing closed-form predictions of strain evolution. These results demonstrate a first step toward reducing the duration and cost of long-term creep testing while retaining physically interpretable model forms.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗