Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Exploring biofiber properties and their influence on biocomposite tensile properties

Biofibers serve as effective reinforcements for neat polylactic acid (PLA) in biocomposites, offering an attractive opportunity to decarbonize the manufacturing sector of the United States by displacing fossil-based reinforcement fibers such as carbon fibers. Also, biofiber production can stimulate economic growth in rural economies, fueling sustainable development. PLA resins are commonly compounded with biofibers to create biocomposites suitable for additive manufacturing. PLA-biofiber composites often exhibit better overall material properties than neat (pure) PLA, but the associations between biofiber properties and the material properties of their biocomposites remain largely unexplored. Hence, this research delves into a comprehensive exploration of diverse biofibers, scrutinizing their physical and chemical attributes, including size, shape, ash content and biochemical composition. The study meticulously analyzes the flow properties of each biofiber and elucidates the ultimate tensile strengths and Young's modulus of corresponding biocomposite samples. Noteworthy correlations between biofiber and biocomposite tensile properties are uncovered, shedding light on critical interrelationships. The study introduces an approach employing regression models to predict the ultimate tensile strength and Young's modulus of biocomposites. These models, validated with a cross-validation technique, exhibit remarkable predictive accuracy, particularly in estimating ultimate tensile strength. © 2024 Oak Ridge National Laboratory managed by UT-Battelle, LLC and The Author(s). Polymer International published by John Wiley & Sons Ltd on behalf of Society of Chemical Industry.

36 MATERIALS SCIENCE↗

Early season prediction of within-field crop yield variability by assimilating CubeSat data into a crop model

Accurate early season predictions of crop yield at the within-field scale can be used to address a range of crop production, management, and precision agricultural challenges. While the remote sensing of within-field insights has been a research goal for many years, it is only recently that observations with the required spatio-temporal resolutions, together with efficient assimilation methods to integrate these into modeling frameworks, have become available to advance yield prediction efforts. Here we explore a yield prediction approach that combines daily high-resolution CubeSat imagery with the APSIM crop model. The approach employs APSIM to train a linear regression that relates simulated yield to simulated leaf area index (LAI). That relationship is then used to identify the optimal regression date at which the LAI provides the best prediction of yield: in this case, approximately 14 weeks prior to harvest. Instead of applying the regression on satellite imagery that is coincident, or closest to, the regression date, our method implements a particle filter that integrates CubeSat-based LAI into APSIM to provide end-of-season high-resolution (3 m) yield maps weeks before the optimal regression date. The approach is demonstrated on a rainfed maize field located in Nebraska, USA, where suitable collections of both imagery and in-situ data were available for assessment. The procedure does not require in-field data to calibrate the regression model, with results showing that even with a single assimilation step, it is possible to provide yield estimates with good accuracy up to 21 days before the optimal regression date. Yield spatial variability was reproduced reasonably well, with a strong correlation to independently collected measurements (R 2 = 0.73 and rRMSE = 12%). When the field averaged yield was compared, our approach reduced yield prediction error from 1 Mg/ha (control case based on a calibrated APSIM model), to 0.5 Mg/ha (using satellite imagery alone), and then to 0.2 Mg/ha (results with assimilation up to three weeks prior to the optimal regression date). Such a capacity to provide spatially explicit yield predictions early in the season has considerable potential to enhance digital agricultural goals and improve end-of-season yield predictions.

54 ENVIRONMENTAL SCIENCES↗

Predicting initial trans-membrane pressure across cycles in the ultrafiltration process using random forest

With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.

direct potable reuse↗

Advanced Laboratory and Field Arrays: Evaluating Sampling Techniques for MHK Biological Monitoring (Task 6)

The overall goal of the task was to identify cost effective biological sampling techniques for MHK environmental monitoring. Protected, demersal, and pelagic fish and selected nektonic invertebrates were surveyed using capture and remote sensing techniques at the PacWave sites. The performance of capture and remote sensing monitoring techniques were evaluated and generic nekton monitoring indices were developed for MHK technologies and sites. In parallel to data collections at the PacWave site, the ability of regression models to characterize, detect, and predict change in acoustic data was evaluated using acoustic data collected in Admiralty Inlet, WA, during a biological monitoring study of the proposed SnoPud tidal turbine project site. Expected outcomes of these efforts included an evaluation of instrumentation and techniques used to monitor biological variability; identification of data streams that can be used to detect and quantify change; and sampling requirements to ensure detection of change in monitored variables.

13 HYDRO ENERGY↗

Learning to Count Grave Sites for Cemetery Observation Models With Satellite Imagery

Understanding how people occupy open spaces is important for research in support of population modeling, policy, national security, emergency response, and sustainability. For the past decade, there has been an increase in research toward capturing and reporting population dynamics and patterns of life at the building level and in some open public spaces such as cemeteries and parks. This is done through observation models developed from local sociocultural information acquired at various spatiotemporal scales to inform night, day, and episodic population occupancy estimates (people/1000 sq ft). Sociocultural information for cemeteries and parks is scarcely available and often collected manually. The process is not only marred by inconsistencies but is laborious and time consuming. In this study, we leverage convolutional neural networks (CNNs) and satellite imagery to derive grave site counts as proxy variables to support scalable and accurate sociocultural data required in a population observation model. Through a hybrid workflow (weak localization plus regression model), we characterize a large scale automation process to counting of grave sites. We evaluate and demonstrate the efficacy of proposed workflow using out-of-data set large satellite imagery and establish its broader impact on cemetery observation models.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

In Situ Inference for Earth System Predictability

An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

In Situ Inference for Earth System Predictability

Focal Area: Focal Area 3: Insight gleaned from complex simulated data using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI. Science Challenge: An understanding of future evolution in precipitation extremes is critical to numerous DOE mission questions. Extreme events are by nature short time-scale events that are difficult to diagnose in available model data. Accurate modeling of extreme events necessarily requires high spatial resolution at the storm scale locally. However, the environment in which storms grow is dependent on global, remote, processes. These complex spatiotemporal relationships are impossible to diagnose at resolutions required to accurately model storms responsible for extreme precipitation. At exascale, climate simulations will produce results at fine enough resolution to investigate these relationships. However, the resulting data from these simulations will be far too large to save for post-simulation analysis. We advocate for fitting statistical models inside the simulations as they run, a context known as in situ, which will facilitate scientific investigations using the full fine-scale data stream. Figure 1 shows an example of the type of model we could consider, a Bayesian hierarchical spatial regression model. Precipitation extremes at each grid cell are modeled using extreme value distributions. Since extremes are rare, fitting models to individual grid cells can result in high variance and poor estimates. Instead, the model can be made more robust by smoothing the parameters of the extreme value model across space. Additionally, the parameters themselves can be functionally linked to other variables elsewhere in the simulation. Thus, we can use the fine-scale data to build more robust models for extremes that link extreme behavior to other climate patterns.

54 ENVIRONMENTAL SCIENCES↗

High-Resolution PM2.5 Concentrations Estimation Based on Stacked Ensemble Learning Model Using Multi-Source Satellite TOA Data

Nepal has experienced severe fine particulate matter (PM2.5) pollution in recent years. However, few studies have focused on the distribution of PM2.5 and its variations in Nepal. Although many researchers have developed PM2.5 estimation models, these models have mainly focused on the kilometer scale, which cannot provide accurate spatial distribution of PM2.5 pollution. Based on Gaofen-1/6 and Landsat-8/9 satellite data, we developed a stacked ensemble learning model (named XGBLL) combined with meteorological data, ground PM2.5 concentrations, ground elevation, and population data. The model includes two layers: a XGBoost and Light GBM model in the first layer, and a linear regression model in the second layer. The accuracy of XGBLL model is better than that of a single model, and the fusion of multi-source satellite remote sensing data effectively improves the spatial coverage of PM2.5 concentrations. Besides, the spatial distribution of the daily mean PM2.5 concentrations in the Kathmandu region under different air conditions was analyzed. The validation results showed that the monthly averaged dataset was accurate (R2 = 0.80 and root mean square error = 7.07). In addition, compared to previous satellite PM2.5 datasets in Nepal, the dataset produced in this study achieved superior accuracy and spatial resolution.

Environmental Sciences & Ecology↗

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data↗

Nonlinear Logistic Regression Mixture Experiment Modeling for Binary Data Using Dimensionally Reduced Components

This article presents and illustrates an approach using nonlinear logistic regression for modeling binary response data from a mixture experiment when the components can be partitioned into groups used to form dimensionally-reduced pseudocomponents (DRPs). A DRP is a linear combination of the components in a group, where the linear combinations over all groups are normalized so that the DRP proportions sum to unity. Nonlinear logistic regression is required because, after normalization of the linear combinations, the model expressed in terms of the DRPs is nonlinear in the parameters that specify the linear combinations. A method for obtaining nonparametric tolerance limits on the probability of a “success” for the binary response variable using a bootstrap approach is also presented. Having three DRPs enables viewing data and modeling results on a ternary plot even though there may be many more than three mixture components. A real database, involving whether or not nepheline crystals form in simulated nuclear waste glass after cooling, is used to illustrate the nonlinear logistic regression modeling and nonparametric tolerance limit approaches when there are three DRPs.

Waste glass, Nepheline, Mixture experiment, Logist↗

Supervised Learning for Distribution Secondary Systems Modeling: Improving Solar Interconnection Processes

The current interconnection process and hosting capacity analysis for distributed energy resources (DERs), such as photovoltaics (PV) and battery energy storage systems, are based on analyzing grid network constraints (voltage and thermal) using only medium-voltage distribution network models. This is because most utilities do not have secondary low-voltage system models that connect service transformers and residential customers. This is important because in many cases the main impact of interconnecting DERs could occur on the low-voltage distribution systems. This paper proposes a supervised learning method to approximate local secondary models to improve the interconnection process. The proposed supervised learning method includes a decision tree model that predicts the secondary topology and a logistic regression model that predicts conductor types. The case studies demonstrate the benefits of including secondary low-voltage circuits in the interconnection process. We report the proposed modeling methodology is readily scalable and thus can reduce the cost and effort of PV interconnection for the industry and stakeholders.

14 SOLAR ENERGY↗

Shock Hugoniot calculations using on-the-fly machine learned force fields with ab initio accuracy

We present a framework for computing the shock Hugoniot using on-the-fly machine learned force field (MLFF) molecular dynamics simulations. In particular, we employ an MLFF model based on the kernel method and Bayesian linear regression to compute the free energy, atomic forces, and pressure, in conjunction with a linear regression model between the internal and free energies to compute the internal energy, with all training data generated from Kohn–Sham density functional theory (DFT). We verify the accuracy of the formalism by comparing the Hugoniot for carbon with recent Kohn–Sham DFT results in the literature. In so doing, we demonstrate that Kohn–Sham calculations for the Hugoniot can be accelerated by up to two orders of magnitude, while retaining ab initio accuracy. We apply this framework to calculate the Hugoniots of 14 materials in the FPEOS database, comprising 9 single elements and 5 compounds, between temperatures of 10 kK and 2 MK. We find good agreement with first principles results in the literature while providing tighter error bars. In addition, we confirm that the inter-element interaction in compounds decreases with temperature.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Concurrent Inter-Model Spread of Boreal Winter Westerly Jet Meridional Positions Between the Northern and Southern Hemispheres in CMIP6 Models

Here, this study investigates the inter-model spread of climatological extratropical westerly jets in boreal winter, using the historical simulation of 52 Coupled Model Intercomparison Project phase 6 (CMIP6) models from 1851 to 2014. The results show that there is a substantial spread in the latitude of the upper-tropospheric westerly jet across models, characterised by large inter-model standard deviations to both the poleward and equatorward sides of the jet axis, although the multi-model ensemble mean (MME) performs well in simulating meridional position of westerly jets. Furthermore, we detect the consistency of inter-model jet position spread between the Northern and Southern Hemispheres, based on the inter-model empirical orthogonal function (EOF) decomposition and correlation of regional-averaged zonal winds. Specifically, the models that simulate the westerly jets poleward/equatorward relative to the MME position in one hemisphere also tend to simulate the jets poleward/equatorward in the other hemisphere. Accordingly, we define a global jet spread index to depict the concurrence of jet shift in the two hemispheres. The results of inter-model regression analyses based on this index indicate that the models positioning the jets poleward than the MME tend to simulate a wider Hadley Cell, a poleward-shifted Ferrel Cell in the Southern Hemisphere, enhanced precipitation in the subtropics and suppressed precipitation in the tropics, and warmer sea surface temperatures in the subtropics and mid-latitudes. The present results suggest that improving the simulation of jet positions in climate models requires a comprehensive consideration of thermal states in the tropics and subtropics/mid latitudes.

54 ENVIRONMENTAL SCIENCES↗

Opportunities for Using the Industrial Assessment Center Database for Industrial Water Use Analysis

The manufacturing sector accounted for approximately 5–6% of total U.S. water use in 2015. Of that amount, 75–80% is self supplied withdrawal from surface-water and groundwater sources and the remainder is from public water supplies. Although manufacturing facilities commonly locate in water-scarce areas, water scarcity still poses a great risk to the manufacturing sector. Reliable water is necessary for any facility that relies on it for process and comfort cooling, cleaning, employee use, and steam generation. One of these barriers to water efficiency is the lack of reliable data on overall U.S. industrial water use—how it is used and the quantities required for each sector. If a facility cannot be easily compared with a facility of similar size and sector, knowing if it is effectively using water conservation best practices is difficult. One potential source of industrial water use data is the U.S. Department of Energy (DOE)–sponsored Industrial Assessment Centers (IACs). IACs are university-based organizations that provide free audits to small- and medium-sized manufacturing facilities to identify productivity improvement and waste and energy reduction opportunities. The IACs also maintain a database of all the audits conducted, which currently holds more than 19,267 assessments and 145,000 recommendations (as of July 24, 2020). This database also contains energy utility (electricity, natural gas, and other fuels) and water utility data, making it a potential data source for industrial water use. This report attempts to create regression models to predict a small- or medium-sized industrial facility’s annual water use or cost based on its industrial subsector and several possible relevant variables. Using data collected by IAC assessments, models for several industrial subsectors were generated via stepwise regression techniques to determine which variables (annual sales, number of employees, facility/plant area, annual production hours, and a water stress metric) are relevant.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nonlinear sparse Bayesian learning for physics-based models

This paper addresses the issue of overfitting while calibrating unknown parameters of over-parameterized physics-based models with noisy and incomplete observations. Here, a semi-analytical Bayesian framework of nonlinear sparse Bayesian learning (NSBL) is proposed to identify sparsity among model parameters during Bayesian inversion. NSBL offers significant advantages over machine learning algorithm of sparse Bayesian learning (SBL) for physics-based models, such as 1) the likelihood function or the posterior parameter distribution is not required to be Gaussian, and 2) prior parameter knowledge is incorporated into sparse learning (i.e. not all parameters are treated as questionable). NSBL employs the concept of automatic relevance determination (ARD) to facilitate sparsity among questionable parameters through parameterized prior distributions. The analytical tractability of NSBL is enabled by employing Gaussian ARD priors and by building a Gaussian mixture-model approximation of the posterior parameter distribution that excludes the contribution of ARD priors. Subsequently, type-II maximum likelihood is executed using Newton's method whereby the evidence and its gradient and Hessian information are computed in a semi-analytical fashion. We show numerically and analytically that SBL is a special case of NSBL for linear regression models. Subsequently, a linear regression example involving multimodality in both parameter posterior pdf and model evidence is considered to demonstrate the performance of NSBL in cases where SBL is inapplicable. Next, NSBL is applied to identify sparsity among the damping coefficients of a mass-spring-damper model of a shear building frame. These numerical studies demonstrate the robustness and efficiency of NSBL in alleviating overfitting during Bayesian inversion of nonlinear physics-based models.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Deep learning approaches for instantaneous laser absorptance prediction in additive manufacturing

Abstract The quantification of absorbed light is essential for understanding laser-material interactions and melt pool dynamics in order to minimize defects in additively manufactured metal components. The geometry of a vapor depression formed during laser melting is closely related to laser energy absorption. This relationship has been observed by the state-of-the-art in situ high-speed synchrotron X-ray visualization and integrating sphere radiometry. These two techniques create a temporally resolved dataset consisting of vapor depression images and corresponding laser absorptance. In this work, we propose two different approaches to predict instantaneous laser absorptance. The end-to-end approach uses deep convolutional neural networks to learn implicit features of X-ray images automatically and predict the laser energy absorptance. The two-stage approach uses a semantic segmentation model to engineer geometric features and predict absorptance using classical regression models. While having distinct advantages, both approaches achieved a consistently low mean absolute error of less than 3.3%.

Chemistry↗

Simple Heat Transfer Model for Film Cooling Applications

This report describes the development of a simple engineering model for film cooling. This model is used to derive a relationship between local wall temperature variations and key cooling performance parameters like local heat transfer coefficients and film effectiveness. This relation and method new and different from previously published models. The scope of this report includes the derivation of regression model equations for a flat plate with and without film cooling. The model equation for a flat plate without film cooling can be used to estimate local heat transfer coefficients using surface temperatures measured from infrared thermography. The model equation for the flat plate with film cooling can be used to estimate film cooling effectiveness, $η_f$, and heat transfer augmentation from the film cooling jet(s).

42 ENGINEERING↗

Array-Based Machine Learning for Functional Group Detection in Electron Ionization Mass Spectrometry

Mass spectrometry is a ubiquitous technique capable of complex chemical analysis. The fragmentation patterns that appear in mass spectrometry are an excellent target for artificial intelligence methods to automate and expedite the analysis of data to identify targets such as functional groups. To develop this approach, we trained models on electron ionization (a reproducible hard fragmentation technique) mass spectra so that not only the final model accuracies but also the reasoning behind model assignments could be evaluated. The convolutional neural network (CNN) models were trained on 2D images of the spectra using transfer learning of Inception V3, and the logistic regression models were trained using array-based data and Scikit Learn implementation in Python. Our training dataset consisted of 21,166 mass spectra from the United States’ National Institute of Standards and Technology (NIST) Webbook. The data was used to train models to identify functional groups, both specific (e.g., amines, esters) and generalized classifications (aromatics, oxygen-containing functional groups, and nitrogen-containing functional groups). We found that the highest final accuracies on identifying new data were observed using logistic regression rather than transfer learning on CNN models. It was also determined that the mass range most beneficial for functional group analysis is 0–100 m/z. We also found success in correctly identifying functional groups of example molecules selected from both the NIST database and experimental data. Beyond functional group analysis, we also have developed a methodology to identify impactful fragments for the accurate detection of the models’ targets. The results demonstrate a potential pathway for analyzing and screening substantial amounts of mass spectral data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗