Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A regression model for the temporal development of soil pipes and associated gullies in the alluvial-fill valley of the Rio Puerco, central New Mexico

On Mars, the association of gullied escarpments and chaotic terrain is evidence for failure and scarp retreat of poorly consolidated materials. Some martian gullies have no surface outlets and may have drained through subterranean channels. Similar features, though on a much smaller scale, can be seen in alluvium along terrestrial river banks in semiarid regions, such as the Rio Puerco Valley of central New Mexico. Many of the escarpments along the Rio Puerco are developing through formation of collapse gullies, which drain through soil pipes. Gully development can be monitored on aerial photographs taken in 1935, 1962, and 1980. A regression model was developed to quantify gully evolution over a known time span. Soil pipes and their associated collapse gullies make recognizable signatures on the air photos. The areal extent of this signature can be normalized to the scarp length of each pipe-gully system, which makes comparisons between systems possible.

Condit, C. D.↗

Simulation and Regression Modeling of X-59 Low-Boom Carpets Across America

The NASA X-59 aircraft is predicted to produce a significantly quieter cruise sonic boom than traditional N-wave-producing aircraft. A propagation simulation study was undertaken to quantify loudness levels, exposure size, and variability of the X-59 low-boom carpet using realistic atmospheric profiles across the contiguous United States of America (CONUS). Near-field pressure data of the X-59 in supersonic cruise from NASA’s fully unstructured Navier–Stokes three-dimensional (known as FUN3D) computational fluid dynamics code were propagated using NASA’s PCBoom code, which solves an enhanced Burgers equation along acoustic rays. Atmospheric profiles from the National Oceanic and Atmospheric Administration’s Climate Forecast System Version 2 database were used for propagation at 138 locations across the CONUS. Carpets at each location were generated for aircraft headings in the four cardinal directions. Over one million X-59 carpets were generated in total. The effects of the heading, season, geography, and climate zone on boom levels and exposure size are presented. Multiple linear regression models were developed to estimate carpet width and loudness metrics across the CONUS. These results inform regulators and mission planners on expected variations in boom levels and carpet extent from atmospheric variations. Understanding potential carpet variability is important when planning community noise surveys using the X-59.

X-59↗

Simulation and Regression Modeling of Nasa'S X-59 Low-Boom Carpets Across America

NASA’s X-59 aircraft is predicted to produce a significantly quieter cruise sonic boom than traditional N-wave-producing aircraft. A propagation simulation study was undertaken to quantify loudness levels, exposure size, and variability of the X-59’s low-boom carpet using realistic atmospheric profiles across the contiguous United States of America (CONUS). Near-field pressure data of the X-59 in supersonic cruise from NASA’s fully unstructured Navier–Stokes three-dimensional (known as FUN3D) computational fluid dynamics code were propagated using NASA’s PCBoom code, which solves an enhanced Burgers equation along acoustic rays. Atmospheric profiles from the National Oceanic and Atmospheric Administration’s Climate Forecast System Version 2 database were used for propagation at 138 locations across the CONUS. Carpets at each location were generated for aircraft headings in the four cardinal directions. Over one million X-59 carpets were generated in total. The effects of the heading, season, geography, and climate zone on boom levels and exposure size are presented. Multiple linear regression models were developed to estimate carpet width and loudness metrics across the CONUS. These results inform regulators and mission planners on expected variations in boom levels and carpet extent from atmospheric variations. Understanding potential carpet variability is important when planning community noise surveys using the X-59.

X-59↗

Proxy quality control of biomass particles using thermogravimetric analysis and Gaussian process regression models

Abstract The temperature experienced by reactants during preparation in a reactor is a key component in determining the yield and homogeneity of usable chemical products such as biomass particles. Thermocouples with sensors can be used to monitor spatial temperature gradients within reactors but these sensors are often too expensive and/or invasive. The present work proposes a strategy to identify optimal machine learning models to infer the maximum effective temperature experienced by particles during oxidative biomass torrefaction using key thermochemical combustion parameters. The maximum rate of weight loss, the corresponding temperature, and fixed carbon content on a dry‐ash‐free basis are used as literature‐based predictor variables obtained from thermogravimetric analysis. The evaluation of 24 machine‐learning models using the standard tenfold cross‐validation method suggests that the exponential Gaussian process regression (GPR) model is the most effective, followed by other GPR models. These high‐performing GPR models were also utilized to predict the effective preparation temperature distribution of reactor‐produced biomass particles under eight conditions of varying residence time and air‐to‐biomass ratio. The effective preparation temperature and residence time of individual biomass particles were then encoded into the torrefaction severity factor and used to estimate the energy yield of the reactor output as a novel quality control method. © 2023 The Authors. Biofuels, Bioproducts and Biorefining published by Society of Industrial Chemistry and John Wiley & Sons Ltd.

09 BIOMASS FUELS↗

Regression models for vegetation radar-backscattering and radiometric emission

Simple regression estimation of radar backscatter and radiometric emission from vegetative terrain is proposed, based on the exact radiative transfer models. A vegetative canopy is modeled as a Rayleigh scattering layer above an irregular Kirchhoff surface. The rms errors between the exact and the estimated ones are found to be less than 5 percent for emission, and 1 dB for the backscattering case, in most practical uses. The proposed formulas are useful in quickly estimating backscattering and emission from the vegetative terrain.

Eom, H. J.↗

A crystal-plasticity-informed Gaussian Process Regression model to capture anisotropy in single crystal shape memory alloys

This work presents a machine learning (ML) framework that model the anisotropic actuation responses in a shape memory alloy. A Gaussian Process Regression (GPR) based ML model is trained on a set of different crystal orientations subjected to different actuation conditions. The training employed thermo-mechanical responses from a crystal-plasticity model that captures phase-transformation, stress-induced plasticity, and transformation-induced plasticity. Further, on training the GPR-ML model at fixed stress level for different orientations, it captured the thermo-mechanical responses accounting for the anisotropy, and predicted responses for new orientations with good accuracy. The GPR-ML model is able to capture the transformation temperature variations even when trained using multiple stress levels, and the transformation strain showed significant deviations. The developed GPR-ML model gave reasonable predictions for an unexplored sample set of orientations and loading conditions.

36 MATERIALS SCIENCE↗

A novel probabilistic regression model for electrical peak demand estimate of commercial and manufacturing buildings

Due to the high cost of electricity in commercial and industrial sectors, demand forecast models have gained increasing attention. However, there are two unresolved issues: (1) Models are not adaptable when exposed to previously unknown data (2) The value of regression methods vs. state-of-the-art machine learning models has not been made apparent before. This study’s goal is to develop probabilistic demand estimation models. Herein, we propose a probabilistic Bayesian regression framework that can not only estimate future demands with high accuracy but also be updated once new information is available. By applying the proposed algorithm to two real-world case studies (commercial and manufacturing), we show a 40.3% and 30.8% improvement in terms of mean absolute error for the two cases. Moreover, the proposed technique outperforms powerful machine learning approaches, including support vector machine by 10.39%, random forest by 6.17%, and multilayer perceptron by 9.14% in terms of mean absolute percentage error.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Comparison of regression modeling techniques for resource estimation

The development and validation of resource utilization models is an active area of software engineering research. Regression analysis is the principal tool employed in these studies. However, little attention was given to determining which of the various regression methods available is the most appropriate. The objective of the study is to compare three alternative regession procedures by examining the results of their application to one commonly accepted equations for resource estimation. The data studied was summarized, the resource estimation equation was described, the regression procedures were explained, and the results obtained from the proceures were compared.

Card, D. N.↗

Revisit the Scalability of Deep Auto-Regressive Models for Graph Generation

As a new promising approach to graph generations, deep auto-regressive graph generation has drawn increasing attention. It however has been commonly deemed as hard to scale up to work with large graphs. In existing studies, it is perceived that the consideration of the full non-local graph dependences is indispensable for this approach to work, which entails the needs for keeping the entire graph’s info in memory and hence the perceived “inherent” scalability limitation of the approach. This paper revisits the common perception. It proposes three ways to relax the dependences and conducts a series of empirical measurements. It concludes that the perceived “inherent” scalability limitation is a misperception; with the right design and implementation, deep auto-regressive graph generation can be applied to graphs much larger than the device memory. The rectified perception removes a fundamental barrier for this approach to meet practical needs.

Yang, Shuai↗

A prediction interval method for uncertainty quantification of regression models

This paper considers calculation of prediction intervals (PIs) by neural networks (NNs) for quantifying uncertainty in regression tasks, so as to provide fast, accurate and robust emulators to accelerate scientific simulations. We propose a novel method to learn lower and upper bounds of the PI using independent NNs without defining an exclusive loss. Our method requires no distributional assumption, does not introduce extra hyper-parameters, and can effectively identify out-of-distribution samples and quantify their uncertainty. We demonstrate advantages of our method using a benchmark problem and two real-world scientific applications.

Zhang, Pei↗

Design Sensitivity for a Subsonic Aircraft Predicted by Neural Network and Regression Models

A preliminary methodology was obtained for the design optimization of a subsonic aircraft by coupling NASA Langley Research Center s Flight Optimization System (FLOPS) with NASA Glenn Research Center s design optimization testbed (COMETBOARDS with regression and neural network analysis approximators). The aircraft modeled can carry 200 passengers at a cruise speed of Mach 0.85 over a range of 2500 n mi and can operate on standard 6000-ft takeoff and landing runways. The design simulation was extended to evaluate the optimal airframe and engine parameters for the subsonic aircraft to operate on nonstandard runways. Regression and neural network approximators were used to examine aircraft operation on runways ranging in length from 4500 to 7500 ft.

Hopkins, Dale A.↗

Graphical Gaussian Process Regression Model for Aqueous Solvation Free Energy Prediction of Organic Molecules in Redox Flow Battery

The solvation free energy of organic molecules is a critical parameter in determining emergent properties such as solubility, liquid-phase equilibrium constants, and pKa and redox potentials in an organic redox flow battery. In this work, we present a machine learning (ML) model that can learn and predict the aqueous solvation free energy of an organic molecule using Gaussian process regression method based on a new molecular graph kernel. To investigate the performance of the ML model on electrostatic interaction, the nonpolar interaction contribution of solvent and the conformational entropy of solute in solvation free energy, three data sets with implicit or explicit water solvent models, and contribution of conformational entropy of solute are tested. We demonstrate that our ML model can predict the solvation free energy of molecules at chemical accuracy with a mean absolute error of less than 1 kcal/mol for subsets of the QM9 dataset and the Freesolv database. To solve the general data scarcity problem for a graph-based ML model, we propose a dimension reduction algorithm based on the distance between molecular graphs, which can be used to examine the diversity of the molecular data set. It provides a promising way to build a minimum training set to improve prediction for certain test sets where the space of molecular structures is predetermined.

25 ENERGY STORAGE↗

Linear Regression Model for Predictive Service Provider Selection

The increasing number of satellites in orbit has led to a growing reliance on third-party service providers for data transfer between Earth and space. Traditional approaches to managing satellite communications require human intervention, which becomes more burdensome with the escalating number of satellites. This research addresses the need for an efficient and automated system to optimize service provider selection for NASA space communication. Previous research has utilized human-operated approaches for service provider management. Our study fills a gap by developing a cognitive algorithm that automates and optimizes the selection process based on various parameters, such as data volume, priority, quality of service and cost. This novel solution reduces user burden, facilitates service management, and contributes to the development of cognitive spaceflight missions, ultimately supporting NASA’s research into Cognitive Communications technology. The algorithm design consists of three major steps: modeling data, developing a Link Selection Algorithm (LSA) based on a grading system, and applying machine learning using linear regression. The LSA evaluates providers based on user-defined constraints, considering factors such as delivery time, cost, and quality of service. We define a suitability metric which allows our algorithm to make a recommendation to a user regarding which commercial service providers to select. The addition of Linear Regression predicts the future suitability value. Our main findings demonstrate that the resulting algorithm can autonomously manage connections between satellites and providers, maximizing communication channel efficiency. This research has significant implications, as it not only addresses a pressing issue in satellite communication management but also advances the field of cognitive spaceflight missions.

Linear regression↗

An ℓ 0 ℓ 2 -norm regularized regression model for construction of robust cluster expansions in multicomponent systems

In this work we introduce ℓ 0 ℓ 2 -norm regularization and hierarchy constraints into linear regression for the construction of cluster expansions to describe configurational disorder in materials. The approach is implemented through mixed integer quadratic programming (MIQP). The ℓ 2 -norm regularization is used to suppress intrinsic data noise, while the ℓ 0 -norm is used to penalize the number of nonzero elements in the solution. The hierarchy relation between clusters imposes relevant physics and is naturally included by the MIQP paradigm. As such, sparseness and cluster hierarchy can be well optimized to obtain a robust, converged set of effective cluster interactions with improved physical meaning. We demonstrate the effectiveness of ℓ 0 ℓ 2 -norm regularization in two high-component disordered rocksalt cathode material systems, where we compare the cross-validation, convergence speed, and the reproduction of phase diagrams, voltage profiles, and Li-occupancy energies with those of the conventional ℓ 1 -norm regularized cluster expansion models.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Interpretable, extensible linear and symbolic regression models for charge density prediction using a hierarchy of many-body correlation descriptors

Here, density functional theory (DFT) is routinely used to make electronic structure predictions for high-throughput screening of materials and molecules for technologically relevant areas, like the identification of better catalysts, electronic materials, and drug discovery. However, the DFT formalism is limited by (a) its poor (quadratic-to-quartic) scaling, and (b) the need to perform repeated eigenvalue computations of the electronic Hamiltonian as part of its self-consistent field (SCF) iteration procedure to obtain the converged ground state electron density, ρ (r). Approaches that directly predict ρ (r) of a structure with high accuracy can accelerate conventional SCF calculations and can also be used in linearly scaling methods such as orbital-free DFT. To this end, we present a procedure to predict the ground state electron density of molecular and periodic three-dimensional systems directly from the atomic structure with a particular emphasis on physical interpretability. In our framework, ρ (r) is modeled using many-body correlation descriptors that accurately capture the effects of local atomic arrangements in the neighborhood of a grid point. Our use of a linear regression scheme to fit to charge density data enables transparent analysis of the relative contributions of various types of local atomic correlations. By systematically including increasingly complex correlations, our model is shown to accurately predict ρ (r) for a variety of chemically and electronically diverse systems — amorphous Ge, Al(001) slab, crystalline Ga 2 O 3 , molecular benzene, and polyethylene. We then demonstrate a symbolic regression-based protocol to construct easily computable, interpretable features from lower-order correlations that significantly improves our electron density predictions with effectively no increase in the computational cost.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Conditional Karhunen–Loève regression model with Basis Adaptation for high-dimensional problems: Uncertainty quantification and inverse modeling

Here, we propose a methodology for improving the accuracy of surrogate models of the observable response of physical systems as a function of the systems’ spatially heterogeneous parameter fields, with applications to uncertainty quantification and parameter estimation in high-dimensional problems. Practitioners often formulate finite-dimensional representations of spatially heterogeneous parameter fields using truncated unconditional Karhunen–Loève expansions (KLEs) for a certain choice of unconditional covariance kernel and construct surrogate models of the observable response with respect to the KLE coefficients. When direct measurements of the parameter fields are available, we propose improving the accuracy of these surrogate models by representing the parameter fields via conditional Karhunen-Loève expansions (CKLEs). CKLEs are constructed by conditioning the covariance kernel of the unconditional expansion on the direct measurements of the parameter field via Gaussian process regression, and then truncating the corresponding KLE. We apply the proposed methodology to constructing surrogate models via the Basis Adaptation (BA) method of the stationary hydraulic head response, measured at spatially discrete observation locations, of a groundwater flow model of the Hanford Site, as a function of the 1000-dimensional representation of the model’s log-transmissivity field. We find that BA surrogate models of the hydraulic head based on CKLEs are more accurate than BA surrogate models based on unconditional expansions for forward uncertainty quantification tasks. Furthermore, we find that inverse estimates of the hydraulic transmissivity field computed using CKLE-based BA surrogate models are more accurate than those computed using unconditional BA surrogate models.

97 MATHEMATICS AND COMPUTING↗

Case study: Accounting for response measurement error in fitting a regression model

This article presents a case study motivated by a plot of data that suggested an emerging trend that the authors were faced with explaining. When the measurement error of the data is accounted for, it turns out there was no real trend. Furthermore, this article shows how to use a Bayesian modeling approach to account for the measurement error.

42 ENGINEERING↗

Accurate Prediction of Algal Biomass Lipid, Protein, and Carbohydrate Composition with Machine Learning Regression Modelling of Near-IR Spectra

During large scale algal biomass cultivation, it is difficult to reliably control relative composition to target levels. Rapid determination of chemical composition is feasible by using near infrared (NIR) spectral data. We sought to build and improve on reliable high-throughput screening prediction method based on partial least squares regression (PLSR) by the application of artificial neural networks (ANN) and associated optimization strategies. The algal biomass sample set was designed and created in an iterative process of culturing in physiologically diverse conditions at the GAI field site, followed by compositional analyses at NREL. The workflow allowed us to identify gaps in compositional space for informing the subsequent cultivation and sampling efforts and generated a high quality set of 210 unique samples with chemical analysis results, spectral scanning data, and cultivation metadata. We observed a significant improvement in the performance of carbohydrate content predictions using an optimized ANN model compared to PLSR, with > 16% reduction in mean absolute percent error (MAPE) when tested on the same set of reserved data. The optimized ANN models for FAME and protein prediction performed exceptionally well with 5.99% and 5.09% MAPE, respectively. Application of these methods to detection and quantification of minor biomass constituents that are relevant to certain product streams has shown positive preliminary results, opening the possibility for extensions to the outputs of this powerful data type. All models are accompanied by prediction uncertainties and unsupervised spectral outlier detection to alert an operator to unreliable spectral data. These tools can be deployed for rapid determination of algal culture status, and cultivation and biomass quality improvement.

algal biofuels↗