Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Relationship of physiography and snow area to stream discharge

The author has identified the following significant results. A comparison of snowmelt runoff models shows that the accuracy of the Tangborn model and regression models is greater if the test data falls within the range of calibration than if the test data lies outside the range of calibration data. The regression models are significantly more accurate for forecasts of 60 days or more than for shorter prediction periods. The Tangborn model is more accurate for forecasts of 90 days or more than for shorter prediction periods. The Martinec model is more accurate for forecasts of one or two days than for periods of 3,5,10, or 15 days. Accuracy of the long-term models seems to be independent of forecast data. The sufficiency of the calibration data base is a function not only of the number of years of record but also of the accuracy with which the calibration years represent the total population of data years. Twelve years appears to be a sufficient length of record for each of the models considered, as long as the twelve years are representative of the population.

Mccuen, R. H.↗

Logistic Regression in Clinical Studies

• A logistic regression model is used when the outcome of interest is binary. The term “logistic” refers to the underlying “logit” (log odds) function that is used to model the binary outcome. • Odds ratios are produced from a logistic regression model and have a useful interpretation. • Tips, tricks and concepts used to fit logistic regression models are similar to those used in linear regression models. • Modeling building that is knowledge-based rather than automatic is preferred in most applications of logistic regression. • A logistic regression model that is overparameterized (ie, too many variables for too few events) can result in odds ratios that are implausibly large and confidence intervals that are wide and uninterpretable. These types of “overfitted” models should be avoided. • Logistic regression models can be fit using most standard statistical software.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Comparison of Iterative and Non-Iterative Strain-Gage Balance Load Calculation Methods

The accuracy of iterative and non-iterative strain-gage balance load calculation methods was compared using data from the calibration of a force balance. Two iterative and one non-iterative method were investigated. In addition, transformations were applied to balance loads in order to process the calibration data in both direct read and force balance format. NASA's regression model optimization tool BALFIT was used to generate optimized regression models of the calibration data for each of the three load calculation methods. This approach made sure that the selected regression models met strict statistical quality requirements. The comparison of the standard deviation of the load residuals showed that the first iterative method may be applied to data in both the direct read and force balance format. The second iterative method, on the other hand, implicitly assumes that the primary gage sensitivities of all balance gages exist. Therefore, the second iterative method only works if the given balance data is processed in force balance format. The calibration data set was also processed using the non-iterative method. Standard deviations of the load residuals for the three load calculation methods were compared. Overall, the standard deviations show very good agreement. The load prediction accuracies of the three methods appear to be compatible as long as regression models used to analyze the calibration data meet strict statistical quality requirements. Recent improvements of the regression model optimization tool BALFIT are also discussed in the paper.

Ulbrich, N.↗

2021 Monthly Rice Production in Chinese Coastal Provinces

This paper explores the dynamics of rice production in the Chinese provinces of Liaoning, Jilin, Heilongjiang, Shanghai, Jiangsu, and Zhejiang and seeks to predict monthly rice production in the months of April through October using precipitation and Normalized Difference Vegetation Index as the predictor variables available. We utilize ridge and lasso regression models to predict the rice yield. Results indicate that a lasso regression model with an R 2 value of 0.9991501 with an adjusted R 2 value of 0.9991502 and a ridge regression model with an R 2 value of 0.9865443 with an adjusted R 2 value of 0.9870637 are possible. The lasso regression model does not account for all predictor variables while the ridge regression model does. Both models could be expanded upon to include more observations.

Tubbs, Heidi↗

Neural Network and Regression Soft Model Extended for PAX-300 Aircraft Engine

In fiscal year 2001, the neural network and regression capabilities of NASA Glenn Research Center's COMETBOARDS design optimization testbed were extended to generate approximate models for the PAX-300 aircraft engine. The analytical model of the engine is defined through nine variables: the fan efficiency factor, the low pressure of the compressor, the high pressure of the compressor, the high pressure of the turbine, the low pressure of the turbine, the operating pressure, and three critical temperatures (T(sub 4), T(sub vane), and T(sub metal)). Numerical Propulsion System Simulation (NPSS) calculations of the specific fuel consumption (TSFC), as a function of the variables can become time consuming, and numerical instabilities can occur during these design calculations. "Soft" models can alleviate both deficiencies. These approximate models are generated from a set of high-fidelity input-output pairs obtained from the NPSS code and a design of the experiment strategy. A neural network and a regression model with 45 weight factors were trained for the input/output pairs. Then, the trained models were validated through a comparison with the original NPSS code. Comparisons of TSFC versus the operating pressure and of TSFC versus the three temperatures (T(sub 4), T(sub vane), and T(sub metal)) are depicted in the figures. The overall performance was satisfactory for both the regression and the neural network model. The regression model required fewer calculations than the neural network model, and it produced marginally superior results. Training the approximate methods is time consuming. Once trained, the approximate methods generated the solution with only a trivial computational effort, reducing the solution time from hours to less than a minute.

Patnaik, Surya N.↗

Winter Wheat Yield Assessment from Landsat 8 and Sentinel-2 Data: Incorporating Surface Reflectance, Through Phenological Fitting, into Regression Yield Models

A combination of Landsat 8 and Sentinel-2 offers a high frequency of observations (3–5 days) at moderate spatial resolution (10–30 m), which is essential for crop yield studies. Existing methods traditionally apply vegetation indices (VIs) that incorporate surface reflectances (SRs) in two or more spectral bands into a single variable, and rarely address the incorporation of SRs into empirical regression models of crop yield. In this work, we address these issues by normalizing satellite data (both VIs and SRs) derived from NASA’s Harmonized Landsat Sentinel-2 (HLS) product, through a phenological fitting. We apply a quadratic function to fit VIs or SRs against accumulated growing degree days (AGDDs), which affects the rate of crop development. The derived phenological metrics for VIs and SRs, namely peak, area under curve (AUC), and fitting coefficients from a quadratic function, were used to build empirical regression winter wheat models at a regional scale in Ukraine for three years, 2016–2018. The best results were achieved for the model with near infrared (NIR) and red spectral bands and derived AUC, constant, linear, and quadratic coefficients of the quadratic model. The best model yielded a root mean square error (RMSE) of 0.201 t/ha (5.4%) and coefficient of determination R2 = 0.73 on cross-validation.

phenological fitting↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

Development of a Non-Iterative Balance Load Prediction Algorithm for the NASA Ames Unitary Plan Wind Tunnel

A non-iterative load prediction algorithm for strain-gage balances was developed for the NASA Ames Unitary Plan Wind Tunnels that computes balance loads from the electrical outputs of the balance bridges and a set of state variables. A state variable could be, for example, a balance temperature difference or the bellows pressure of a flow-through balance. The algorithm directly uses regression models of the balance loads for the load prediction that were obtained by applying global regression analysis to balance calibration data. This choice greatly simplifies both implementation and use of the load prediction process for complex balance configurations as no load iteration needs to be performed. The regression model of a balance load is constructed by using terms from a total of nine term groups. Four term groups are derived from a Taylor Series expansion of the relationship between the load, gage outputs, and state variables. The remaining five term groups are defined by using absolute values of the gage outputs and state variables. Terms from these groups should only be included in the regression model if calibration data from a balance with known bi-directional outputs is analyzed. It is illustrated in detail how global regression analysis may be applied to obtain the coefficients of the chosen regression model of a load component assuming that no linear or massive near-linear dependencies between the regression model terms exist. Data from the machine calibration of a six-component force balance is used to illustrate both application and accuracy of the non-iterative load prediction process.

Ulbrich, Norbert M.↗

Functional Predictor Variables for the Leaching Potential of Arsenic and Selenium from Coal Fly Ash

The release of leachates from intact coal ash impoundments is a concern due to the enrichment and mobilization of toxic elements such as arsenic (As) and selenium (Se). This study aims to explore the intrinsic properties of coal fly ash that correlate with the relative leachability of As and Se. We performed leaching experiments with 52 fly ash samples collected from 15 different U.S. power plants and representing coal feedstocks from the three major domestic regions. We assessed the mobilization potential of As and Se in fly ash based on standardized leaching protocols and performed multivariate and lasso regression analyses to explore correlations of leachable As and Se contents with characteristics such as major element contents, loss on ignition, and pH. The results of regression models indicated that major elements (Fe, Ca, and Al) for a wide range of fly ashes can serve as predictor variables for the leaching potential of As but not for Se. LOI and pH were not important predictive variables in the models. Both regression approaches resulted in relatively strong fits for leachable As (correlation coefficient R 2 = 0.78 for both models) compared to models for leachable Se (R 2 = 0.49). Overall, these results suggest that correlation models combined with on-site elemental analysis with portable analyzers may enable a screening method for leachable As in coal ash.

01 COAL, LIGNITE, AND PEAT↗

Machine Learning Accelerated First-Principles Study of the Hydrodeoxygenation of Propanoic Acid

The complex reaction network of catalytic biomass conversions often involves hundreds of surface intermediates and thousands of reaction steps, greatly hindering the rational design of metal catalysts for these conversions. Here, we present a framework of machine learning (ML)-accelerated first-principles studies for the hydrodeoxygenation (HDO) of propanoic acid over transition metal surfaces. The microkinetic model (MKM) is initially parametrized by ML-predicted energies and iteratively improved by identifying the rate-determining species and steps (RDS), computing their energies by density functional theory (DFT), and reparameterizing the MKM until all the RDS are computed by DFT. The Gaussian process (GP) model performs significantly better than the linear ridge regression model for predicting both the adsorption free energies and transition state free energies. Parameterized with energies from the GP model, only 5–20% of the full reaction network has to be computed by DFT for the MKM to possess DFT-level accuracy for the TOF and dominant reaction pathway. While the linear ridge regression model performs worse than the GP model, its performance is greatly improved when only transition states are predicted by the regression model and adsorption energies are computed by DFT. Overall, we find that a high accuracy in adsorption free energies is more important for a reliable MKM than a high accuracy in TS free energies. Lastly, based on the GP model with GOH and GCHCHCO as catalyst descriptors, we build two-dimensional volcano plots in activity and selectivity that can help design promising alloy catalysts for HDO reactions of organic acids.

adsorption↗

Machine learning-enabled prediction of chemical durability of A 2 B 2 O 7 pyrochlore and fluorite

Pyrochlore-structure type and its derivative in a general formula A 2 B 2 O 7 (A = rare earth elements and actinides; B = Ti, Sn, Zr, Hf, Pb, Si, etc.) display excellent structural flexibility and rich crystal chemistry as promising nuclear waste form materials capable of immobilizing actinides and fission products. It is essential to understand these materials’ chemical durability and element release of radionuclides in order to evaluate their performance in near-field environment. However, it is a formidable grand technological challenge to experimentally perform durability testing across hundreds of thousands of possibilities resulting from their extreme compositional complexities due to cation substitutions at both A and B-sites. In this work, we demonstrate a machine learning approach to determine the key materials parameters and structural characteristics governing the leaching behaviors from a small set of selected compositions as model systems, enabling a science-based prediction of their chemical durability that can be extended to a wide range of chemical compositions. The combination of four key structural characteristics and materials parameters, including ionic radius size difference , ionic potential difference , electronegativity difference , and lattice parameter , creates features an optimized prediction of the chemical durability. Two machine learning models, linear regression and Kernel ridge regression models, are trained on the randomly-split training dataset derived from the experimentally-determined elemental release rates, and subsequently tested on the testing dataset. The predicted leaching rates from both machine learning models show an excellent agreement with the experimental data, demonstrating the feasibility of rapidly evaluating the material properties of new compositions. These results highlight the immense potential of synergizing informatics through machine learning-based models and well-controlled experiments of selected model systems to accelerate materials design and discovery with optimized compositions and performance of promising materials for effective nuclear waste management.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Reflectance of vegetation, soil, and water

The author has identified the following significant results. The Kubelka-Munk model, a regression model, and a combination of these models were used to extract plant, soil, and shadow reflectance components of vegetated surfaces. The combination model was superior to the others; it explained 86% of the variation in band 5 reflectance of corn and sorghum, and 90% of the variation in band 6 reflectance of cotton. A fractional shadow term substantially increased the proportion of the digital count sum of squares explained when plant parameters alone explained 85% or less of the variation. Overall recognition of 94 agricultural fields using simultaneously acquired aircraft and spacecraft MSS data was 61.8 and 62.8%, respectively; recognition of vegetable fields larger than 10 acres and taller than 25 cm, rose to 88.9 and 100% for aircraft and spacecraft, respectively. Agriculture and rangeland, were well discriminated for the entire county but level 2 categories of vegetables, citrus, and idle cropland, except for citrus, were not.

Wiegand, C. L.↗

Retrospective Analysis of Backwater Habitat Availability Using Remote Sensing

Low-velocity channel-margin habitats serve as important nursery habitats for the endangered Colorado pikeminnow (Ptychocheilus lucius) in the middle Green River between Jensen and Ouray, Utah. These habitats, known as backwaters, are associated with emergent sandbars, and are shaped and reformed annually by peak flows. Our recent knowledge about backwater characteristics and dynamics that is summarized in the synthesis report (Grippo et al. 2015) was based on detailed annual survey data from a relatively small sample of backwaters that were collected from 2003 to 2014 and reach-wide evaluations of backwater surface area based on manual interpretation of aerial and satellite imagery. Methods that bridge the gap between the detailed surveys from a small number of backwaters and the reach-wide assessment of their surface area would enable an assessment of the availability of backwater habitats that meet the minimum depth requirements for suitable habitat for Colorado pikeminnow. In 2015 Argonne National Laboratory (Argonne) tested three regression models—linear, multiple, and partial least square (PLS) regression models—for estimating backwater depth using National Agriculture Imagery Program (NAIP) imagery collected in July 2006 that covered the Jensen-Ouray reach of the Green River (Hamada and LaGory 2016). The results suggested that a PLS regression model showed high correlation with the reference depth (R 2 = 0.69) and had the most unbiased and consistent estimate of backwater depth. The results also indicated that the PLS model would provide reasonable estimates of the amount of habitat providing a minimum suitable depth of 30 cm for young-of-the-year Colorado pikeminnow, even though absolute depth estimates may be uncertain for backwater areas deeper than approximately 40 cm. The study also provided insights regarding the amount and the selection of calibration and validation (cal-val) data needed for improving the accuracy of depth prediction.

54 ENVIRONMENTAL SCIENCES↗

Optimization of Geometric Perturbations on a Rod Moving Through a High Explosive Target

After completing a study to ensure the simulation results were converged, several high resolution 3D Smoothed Particle Hydrodynamic (SPH) simulations of copper rods impacting a high explosive (LX14) target were performed. This was then formulated into an optimization problem: I wanted to find the optimum shape and location of a perturbation on the rod that would maximize its erosion after it left the target. The shape of the perturbation was modeled as a 2D Gaussian bump and parameterized by its location along the rod axis (z 0 ) and amplitude (A). The final mass of the coherent part of the rod as it leaves the target was used as a metric to represent the erosion of the rod, and the optimization was formulated to maximize this metric with respect to the aforementioned design variables. Due to the expensive nature of the high-fidelity 3D SPH simulations, a surrogate model needed to be chosen so that many function calls to the optimizer would be feasible. Thus, a strategic full factorial sampling plan was chosen to build a dataset, which consisted of 24 high-fidelity simulations. Two surrogate models, a third order polynomial regression model and a Gaussian Process Model, were analyzed using a 14%/86% test/train holdout technique. The root mean square and R2 score of the test set was used to determine the best model, and the third order polynomial regression model was chosen as the surrogate model. Finally, the Nelder-Mead Simplex and Basin-hopping optimization algorithms were implemented, and it was found that the two algorithms gave slightly different optimum values. Nelder-Mead gave an optimum point of [z* 0 ;A*] = [9:9;0:4] and Basin-Hopping gave an optimum value of x* = [z* 0 ;A*] = [9:2;0:1].

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗