Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Accelerated test modeling

Cycle life regression model, cycle life prediction model, and acceleration factors are discussed. A method was presented to: (1) select a mathematical model; (2) determine model coefficients using accelerated test data; (3) test model fit of the accelerated test data; and (4) predict normal packs.

Schwartz, D.↗

External Tank Liquid Hydrogen (LH2) Prepress Regression Analysis Independent Review Technical Consultation Report

The request to conduct an independent review of regression models, developed for determining the expected Launch Commit Criteria (LCC) External Tank (ET)-04 cycle count for the Space Shuttle ET tanking process, was submitted to the NASA Engineering and Safety Center NESC on September 20, 2005. The NESC team performed an independent review of regression models documented in Prepress Regression Analysis, Tom Clark and Angela Krenn, 10/27/05. This consultation consisted of a peer review by statistical experts of the proposed regression models provided in the Prepress Regression Analysis. This document is the consultation's final report.

Parsons, Vickie s.↗

Optimization of artificial viscosity in production codes based on Gaussian Regression surrogate models

To accurately model flows with shock waves using staggered-grid Lagrangian hydrodynamics, artificial viscosity has to be introduced to convert kinetic energy into internal energy, thereby increasing the entropy across shocks. Determining the appropriate strength of the artificial viscosity is an art and strongly depends on the particular problem and experience of the researcher. The objective of this study is to pose the problem of finding the appropriate strength of artificial viscosity as an optimization problem and solve this problem using machine learning (ML) tools, specifically using surrogate models based on Gaussian Process regression and Bayesian analysis. We describe the optimization method and discuss various practical details of its implementation. The shock-containing problems for which we apply this method all have been implemented in the LANL code FLAG. First, we apply ML to find optimal values to isolated shock problems of different strengths. Second, we apply ML to optimize viscosity for a 1D propagating detonation problem based on Zel’dovich-von Neumann-Doring (ZND) detonation theory using a reactive burn model. We compare results for default (currently used values in FLAG) and optimized values of artificial viscosity for these problems demonstrating the potential for significant improvement in the accuracy of computations.

42 ENGINEERING↗

Aerobic Fitness Does Not Contribute to Prediction of Orthostatic Intolerance

Several investigations have suggested that orthostatic tolerance may be inversely related to aerobic fitness (VO (sub 2max)). To test this hypothesis, 18 males (age 29 to 51 yr) underwent both treadmill VO(sub 2max) determination and graded lower body negative pressures (LBNP) exposure to tolerance. VO(2max) was measured during the last minute of a Bruce treadmill protocol. LBNP was terminated based on pre-syncopal symptoms and LBNP tolerance (peak LBNP) was expressed as the cumulative product of LBNP and time (torr-min). Changes in heart rate, stroke volume cardiac output, blood pressure and impedance rheographic indices of mid-thigh-leg initial accumulation were measured at rest and during the final minute of LBNP. For all 18 subjects, mean (plus or minus SE) fluid accumulation index and leg venous compliance index at peak LBNP were 139 plus or minus 3.9 plus or minus 0.4 ml-torr-min(exp -2) x 10(exp 3), respectively. Pearson product-moment correlations and step-wise linear regression were used to investigate relationships with peak LBNP. Variables associated with endurance training, such as VO(sub 2max) and percent body fat were not found to correlate significantly (P is less than 0.05) with peak LBNP and did not add sufficiently to the prediction of peak LBNP to be included in the step-wise regression model. The step-wise regression model included only fluid accumulation index leg venous compliance index, and blood volume and resulted in a squared multiple correlation coefficient of 0.978. These data do not support the hypothesis that orthostatic tolerance as measured by LBNP is lower in individuals with high aerobic fitness.

Convertino, Victor A.↗

Machine Learning–Augmented Laser-Induced Breakdown Spectroscopy for Spectral Discrimination of Iron Oxalates

Enhanced characterization and phase identification of post-PUREX Pu Oxalates (PuOXA) are pivotal for nonproliferation and pre-detonation nuclear forensics. Despite significant advances in the characterization of PuO 2 samples, little is known about the impact of both the chemical structure and oxidation states of PuOXA (i.e., Pu(III) and Pu(IV)) have on optical emission signatures. Here, we demonstrate the analytical capabilities of laser-induced breakdown spectroscopy (LIBS) applied to Fe(II) and Fe(III) oxalate samples as surrogates for PuOXA, highlighting the discriminating features in the LIBS emission spectra arising from differences in the oxidation states within mixed FeOXA samples. We report the enhancement of spectral feature selection using Principal Component Analysis (PCA), which enables the analytical superiority of machine learning algorithms such as Linear Discriminant Analysis (LDA), Quadratic Discriminant Analysis (QDA), Partial Least Squares Regression (PLSR), Support Vector Regression (SVR), and Random Forest Regression (RFR) over conventional univariate techniques for phase discrimination and chemometric analysis. Cluster analysis revealed how both matrix effects and laser ablation influence cluster separability by introducing spectral artifacts that misdirect the maximization of variance. PCA-selected emission lines were used in the regression models, demonstrating that both univariate and multivariate linear regression models (i.e., PLSR and SVR) can achieve acceptable performance, with machine learning models outperforming conventional calibration regressions. Furthermore, the application of non-linearly activated PCA-selected emission lines illustrates how simplifying the data while retaining captured variance enables the use of less complex and more computationally efficient models. Furthermore, this is particularly evident in the underperformance of RFR, which suffers from increased computational costs and overfitting owing to its high complexity.

Oxalates↗

Predicting lettuce canopy photosynthesis with statistical and neural network models

An artificial neural network (NN) and a statistical regression model were developed to predict canopy photosynthetic rates (Pn) for 'Waldman's Green' leaf lettuce (Latuca sativa L.). All data used to develop and test the models were collected for crop stands grown hydroponically and under controlled-environment conditions. In the NN and regression models, canopy Pn was predicted as a function of three independent variables: shootzone CO2 concentration (600 to 1500 micromoles mol-1), photosynthetic photon flux (PPF) (600 to 1100 micromoles m-2 s-1), and canopy age (10 to 20 days after planting). The models were used to determine the combinations of CO2 and PPF setpoints required each day to maintain maximum canopy Pn. The statistical model (a third-order polynomial) predicted Pn more accurately than the simple NN (a three-layer, fully connected net). Over an 11-day validation period, average percent difference between predicted and actual Pn was 12.3% and 24.6% for the statistical and NN models, respectively. Both models lost considerable accuracy when used to determine relatively long-range Pn predictions (> or = 6 days into the future).

Non-NASA Center↗

Multi-Variate LSTM Prediction of Alaska Magnetometer Chain Utilizing a Coupled Model Approach

During periods of rapidly changing geomagnetic conditions electric fields form within the Earth’s surface and induce currents known as geomagnetically induced currents(GICs), which interact with unprotected electrical systems our society relies on. In this study, we train multi-variate Long-Short Term Memory neural networks to predict magnitude of north-south component of the geomagnetic field (|BN|) at multiple ground magnetometer stations across Alaska provided by the SuperMAG database with a future goal of predicting geomagnetic field disturbances. Each neural network is driven by solar wind and interplanetary magnetic field inputs from the NASA OMNI database spanning from 2000–2015 and is fine tuned for each station to maximize the effectiveness in predicting |BN|. The neural networks are then compared against multivariate linear regression models driven with the same inputs at each station using Heidke skill scores with thresholds at the 50, 75, 85, and 99 percentiles for |BN|. The neural network models show significant increases over the linear regression models for |BN| thresholds. We also calculate the Heidke skill scores for d|BN|/dt by deriving d|BN|/dt from |BN| predictions. However, neural network models do not show clear outperformance compared to the linear regression models. To retain the sign information and thus predict BN instead of |BN|, a secondary so-called polarity model is utilized. The polarity model is run in tandem with the neural networks predicting geomagnetic field in a coupled model approach and results in a high correlation between predicted and observed values for all stations. We find this model a promising starting point for a machine learned geomagnetic field model to be expanded upon through increased output time history and fast turnaround times.

Matthew Blandin↗

Benefit Analysis of CO 2 Delivery Options for Offshore Storage or Enhanced Oil Recovery

The analysis presented in this report evaluates the benefits of CO₂ offshore transport via pipeline or ship within the GOM. It takes a top-down framework to estimate the costs. First, this analysis designed a reduced-order model (ROM) based on the cash flows in the FECM/NETL CO₂ Transport Cost Model (also known as CO2_T_COM). The ROM takes capital expenses (CAPEX) and operating expenses (OPEX) to calculate the CO₂ breakeven price based on the cash flows. Second, this analysis developed regression models utilizing published data from other analyses to estimate CAPEX and OPEX. Since the ROM is a simplified cash flow calculation, it is easy to exchange the core regression models to estimate various costs. The ROM and regression models provided a framework that can be easily used by other researchers, decision-makers, operators, and regulators. The objective of this analysis is to assess the CO₂ breakeven cost range for pipeline and ship transport of captured CO₂ given the CO₂ source and storage reservoir located in the GOM.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Distribution Grid Modeling Using Smart Meter Data

The knowledge of distribution grid models, including topologies and line impedances, is essential for grid monitoring, control and protection. However, such information is often unavailable, incomplete or outdated. The increasing deployment of smart meters (SMs) provides a unique opportunity to tackle this issue. This paper proposes a two-stage framework for distribution grid modeling using SM data. In the first stage, the network topology is identified by reconstructing a weighted Laplacian matrix of distribution networks. In the second stage, a least absolute deviations (LAD) regression model is developed for estimating line impedance of a single branch based on the nonlinear (inverse) power flow model, wherein a conductor library is leveraged to narrow down the solution space. The LAD regression model is originally a mixed-integer nonlinear program whose continuous relaxation is still non-convex. Furthermore, we specially address its convex relaxation and discuss the exactness. The modified regression model is then embedded within a bottom-up sweep algorithm to achieve the identification across the network in a branch-wise manner. Numerical results on the IEEE 13-bus, 37-bus and 69-bus test feeders validate the effectiveness of the proposed methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A machine learning approach for efficient multi-dimensional integration

Many physics problems involve integration in multi-dimensional space whose analytic solution is not available. The integrals can be evaluated using numerical integration methods, but it requires a large computational cost in some cases, so an efficient algorithm plays an important role in solving the physics problems. We propose a novel numerical multi-dimensional integration algorithm using machine learning (ML). After training a ML regression model to mimic a target integrand, the regression model is used to evaluate an approximation of the integral. Then, the difference between the approximation and the true answer is calculated to correct the bias in the approximation of the integral induced by ML prediction errors. Because of the bias correction, the final estimate of the integral is unbiased and has a statistically correct error estimation. Three ML models of multi-layer perceptron, gradient boosting decision tree, and Gaussian process regression algorithms are investigated. The performance of the proposed algorithm is demonstrated on six different families of integrands that typically appear in physics problems at various dimensions and integrand difficulties. The results show that, for the same total number of integrand evaluations, the new algorithm provides integral estimates with more than an order of magnitude smaller uncertainties than those of the VEGAS algorithm in most of the test cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Calibration and Data Analysis of the MC-130 Air Balance

Design, calibration, calibration analysis, and intended use of the MC-130 air balance are discussed. The MC-130 balance is an 8.0 inch diameter force balance that has two separate internal air flow systems and one external bellows system. The manual calibration of the balance consisted of a total of 1854 data points with both unpressurized and pressurized air flowing through the balance. A subset of 1160 data points was chosen for the calibration data analysis. The regression analysis of the subset was performed using two fundamentally different analysis approaches. First, the data analysis was performed using a recently developed extension of the Iterative Method. This approach fits gage outputs as a function of both applied balance loads and bellows pressures while still allowing the application of the iteration scheme that is used with the Iterative Method. Then, for comparison, the axial force was also analyzed using the Non-Iterative Method. This alternate approach directly fits loads as a function of measured gage outputs and bellows pressures and does not require a load iteration. The regression models used by both the extended Iterative and Non-Iterative Method were constructed such that they met a set of widely accepted statistical quality requirements. These requirements lead to reliable regression models and prevent overfitting of data because they ensure that no hidden near-linear dependencies between regression model terms exist and that only statistically significant terms are included. Finally, a comparison of the axial force residuals was performed. Overall, axial force estimates obtained from both methods show excellent agreement as the differences of the standard deviation of the axial force residuals are on the order of 0.001 % of the axial force capacity.

Booth, Dennis↗

Prediction of Creep-Induced Strain Using a Symbolic Regression-Based Model

Material creep under high-temperature conditions limits the lifetime and safety of structural systems such as advanced nuclear reactors. Conventional creep testing is slow and often produces inconsistent results across nominally identical experiments, making lifetime prediction uncertain. Here, to address these challenges, this work develops a data-driven symbolic regression (SR) model that consolidates results from duplicate creep tests and predicts the remaining strain-time curve of an ongoing experiment. The method uses piece-wise multi-objective SR with physical constraints to generate analytic, interpretable functions describing transient creep strain. Applied to Inconel Alloy 617 data, the approach achieved relative mean absolute errors of 1.0–9.5%, providing closed-form predictions of strain evolution. These results demonstrate a first step toward reducing the duration and cost of long-term creep testing while retaining physically interpretable model forms.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Quantifying mean, variability, and uncertainty in indoor radon exposure in Pennsylvania using random forest and quantile regression forest models

Radon is a naturally occurring radioactive gas that poses a serious health risk as the primary cause of lung cancer in non-smokers. Despite the well-known adverse association with health outcomes, current radon exposure assessments are limited to county-level or average-level estimates, which fail to capture regional variability. This study uses Machine Learning models, including Random Forest (RF) and Quantile Regression Forest (QRF), to estimate the indoor radon concentrations at the ZCTA (Zip code tabulation area)-level and characterize uncertainties in model estimates. Incorporating geological, meteorological, and building-specific data, the models aim to improve radon risk assessment by capturing mean exposure, variability, and extreme concentration levels. Processed radon test data (n = 718,111) were analyzed using average, variability, and quantile prediction methods. Models that estimate the average radon exposure at the ZCTA-level can yield promising model-fit results, but they do not capture the underlying variability of indoor radon exposure within a ZCTA. We utilize volatility analyses to identify characteristics indicative of high variability of indoor radon exposure. We also show that a QRF model can be used to estimate upper quantiles of residential radon exposure, thereby uncovering localized areas of elevated exposure that were not apparent in mean estimates. The results highlighted the need for a deep characterization of exposure risk and show that regions with moderate average exposure levels could still harbor extreme outliers with implications for evaluating health risks. Utilizing multiple radon exposure models allows for a deeper characterization of radon risk within a geographic area and can better identify high-risk areas. The results from this study provide a foundation for developing mitigation strategies and examining associations between radon exposure and health outcomes at fine scales. Future research should extend the geographic scope and incorporate additional environmental risk factors to establish a comprehensive framework for risk assessment.

Lee, Heechan [ORNL]↗

Fundamental microscopic properties as predictors of large-scale quantities of interest: Validation through grain boundary energy trends

Correlations between fundamental microscopic properties computable from first principles, which we term canonical properties, and complex large-scale quantities of interest (QoIs) provide an avenue to predictive materials discovery. Here, we propose that such correlations can be efficiently discovered through simulations utilizing approximate interatomic potentials (IPs), which serve as an ensemble of “synthetic materials”. As a proof of principle we build a regression model relating canonical properties to the symmetric tilt grain boundary (GB) energy curves in face-centered cubic crystals, characterized by the scaling factor in the universal lattice matching model of Runnels et al. (2016), which we take to be our QoI. Our analysis recovers known correlations of GB energy to other properties and discovers new ones. We also demonstrate, using available density functional theory (DFT) GB energy data, that the regression model constructed from IP data is consistent with DFT results, confirming the assumption that the IPs and DFT belong to same statistical pool and thereby validating the approach. Regression models constructed in this fashion can be used to predict large-scale QoIs based on first-principles data and provide a general method for training IPs for QoIs beyond the scope of first-principles calculations.

36 MATERIALS SCIENCE↗

Distributed Monitoring of the R(sup 2) Statistic for Linear Regression

The problem of monitoring a multivariate linear regression model is relevant in studying the evolving relationship between a set of input variables (features) and one or more dependent target variables. This problem becomes challenging for large scale data in a distributed computing environment when only a subset of instances is available at individual nodes and the local data changes frequently. Data centralization and periodic model recomputation can add high overhead to tasks like anomaly detection in such dynamic settings. Therefore, the goal is to develop techniques for monitoring and updating the model over the union of all nodes data in a communication-efficient fashion. Correctness guarantees on such techniques are also often highly desirable, especially in safety-critical application scenarios. In this paper we develop DReMo a distributed algorithm with very low resource overhead, for monitoring the quality of a regression model in terms of its coefficient of determination (R2 statistic). When the nodes collectively determine that R2 has dropped below a fixed threshold, the linear regression model is recomputed via a network-wide convergecast and the updated model is broadcast back to all nodes. We show empirically, using both synthetic and real data, that our proposed method is highly communication-efficient and scalable, and also provide theoretical guarantees on correctness.

Bhaduri, Kanishka↗

Salience Assignment for Multiple-Instance Data and Its Application to Crop Yield Prediction

An algorithm was developed to generate crop yield predictions from orbital remote sensing observations, by analyzing thousands of pixels per county and the associated historical crop yield data for those counties. The algorithm determines which pixels contain which crop. Since each known yield value is associated with thousands of individual pixels, this is a multiple instance learning problem. Because individual crop growth is related to the resulting yield, this relationship has been leveraged to identify pixels that are individually related to corn, wheat, cotton, and soybean yield. Those that have the strongest relationship to a given crop s yield values are most likely to contain fields with that crop. Remote sensing time series data (a new observation every 8 days) was examined for each pixel, which contains information for that pixel s growth curve, peak greenness, and other relevant features. An alternating-projection (AP) technique was used to first estimate the "salience" of each pixel, with respect to the given target (crop yield), and then those estimates were used to build a regression model that relates input data (remote sensing observations) to the target. This is achieved by constructing an exemplar for each crop in each county that is a weighted average of all the pixels within the county; the pixels are weighted according to the salience values. The new regression model estimate then informs the next estimate of the salience values. By iterating between these two steps, the algorithm converges to a stable estimate of both the salience of each pixel and the regression model. The salience values indicate which pixels are most relevant to each crop under consideration.

Wagstaff, Kiri L.↗

Global Variability in Sonic Boom Exposure due to Macroscopic Effects

Supersonic flight over land has been prohibited since 1973 due to the loudness of sonic booms. NASA is building the X-59 aircraft as part of its Quesst mission to demonstrate low-loudness shaped sonic booms, or “sonic thumps.” The Quesst mission will gather human perception data via a series of community noise surveys across the USA. The noise dose and perceptual response data will be provided to the International Civil Aviation Organization (ICAO) and the Federal Aviation Administration for use in determining potential future supersonic aircraft noise certification standards, effectively changing the prohibition from a speed limit to a noise limit. These noise regulations must be globally effective, as long travel distances see the largest benefit to supersonic flight. The state of the atmosphere through which a sonic boom travels affects the size of the region exposed to sound, the “carpet width” (CW), as well as the loudness. The focus of this dissertation is to understand and quantify the expected loudness and CW of sonic booms due to the macroscopic atmospheric effects around the world. A pair of large-scale propagation simulation studies were conducted using the NASA PCBoom code to compare predicted sonic boom loudness and CW statistics first across the USA and then across the world. For the USA study, near-field data of the X-59 in steady cruise was propagated at 4 cardinal headings at 138 locations through 5 years of Climate Forecast System Version 2 (CFSv2) atmospheric profiles. Results of a bootstrap forest predictor screening model indicated the importance of climate zone, latitude, ground elevation, season, and heading. It also noted the unimportance of time of day for predicting loudness and CW. The data is visualized in aggregate, and then broken out geographically, by season and heading, and by climate zone. Multiple linear regression models were fit to the data from the 138 locations so that estimates of the loudness and CW can be produced anywhere in the US. The results can aid in planning when and where to fly the X-59. For the global study, near-field data from three aircraft, the X-59 in a quiet and loud configuration, B-58, and Concorde, were propagated at four cardinal headings through data from three atmospheric models, the CFSv2, the Global Forecast System (GFS), and the ECMWF Reanalysis Version 5 (ERA5), at 100 global locations over 1 year. Results of a bootstrap forest predictor screening model indicated the importance of climate zone, ground elevation, season, and heading. Similar to the US study, the model indicated time of day was not an important predictor. The model also indicated that choice of weather model was not important, so the atmospheric model data are effectively interchangeable. The ERA5 model was chosen for use in an extension of the study to include 18 additional locations to ensure sampling of every climate zone. Loudness and CW results are shown in aggregate, and split geographically and by heading, season, and climate. Multiple linear regression models were fit to the data from the 118 locations so that estimates of loudness and CW can be produced around the world. N-waves and shaped booms did not have the same global variability. Koppen-Geiger climate zones were used as the climate zone definition for the global study. These are available as present-day and future climate projections. Making use of the multiple linear regression models, the future climate zones were input to estimate the effect of the changing climate on sonic boom loudness and CW. Results indicate that a changing climate would have little impact on the effectiveness of noise regulations.

X-59↗

Machine Learning-Driven Reliability Estimation of PV Inverters Considering Alert-Ambient Variability

Weather-induced spatio-temporal degradation limits outdoor PV inverter lifetime and reliability, necessitating advanced data analysis. This study employs a top-down, data-driven approach utilizing multiple machine learning (ML) algorithms to estimate inverter reliability in a 1.4 MW PV power plant, considering factors such as irradiance, humidity, temperature, time of day, and weather conditions. An extensive alert dataset from 17 identical inverters, including alert types, propagation, and frequency, reveals significant correlations with environmental factors and inverter output power, enabling the construction of a performance reliability model. Dual-stage supervised-ML models are evaluated for accuracy, with the ‘classification-regression’ model by an artificial neural network (ANN) tested on the averaged “Alert-Ambient” dataset, which is outperformed by ‘clustering-regression’ models using random forest (RF) and K-Nearest Neighbors (KNN) on individual inverter datasets. K-means clustering applies principal component analysis to reduce dimensions, achieving improved accuracy beyond the 80% achieved by ANN on the averaged dataset. Second-stage regression estimates inverter reliability with a mean square error of 0.0195 on the averaged dataset and as low as 0.002 on individual inverter datasets using RF. Furthermore, these findings highlight the method's suitability for estimating PV inverter output reliability under ambient conditions, essential for digital twin development and related applications.

14 SOLAR ENERGY↗