Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Predicting the evolution of biomass bulk density through feedstock preprocessing: Discrete element modeling, regression analysis, and pilot-scale validation

Bulk density is an important material property of biomass feedstocks, influencing handling, storage, transport costs, and conversion efficiency. In this study, predictive regression models for loose and tapped bulk densities of Alamo and Cave-in-Rock switchgrass are developed using a comprehensive dataset generated via calibrated bonded-sphere discrete element method (DEM) simulations. Here, a key contribution of this study is the use of a DEM-based approach, which correlates density with moisture content and particle size distribution parameters and enables analysis across a continuous particle size range, overcoming limitations of purely experimental data. For comparison, regression models are also developed using only experimental data from pilot-scale runs at the Biomass Feedstock National User Facility at Idaho National Laboratory. Validation against pilot-scale data showed reasonable prediction accuracy for both model types, particularly for smaller particle sizes (post-secondary grinding). While the experimental model showed slightly better performance matching the validation data in some cases, the DEM-based model benefits from a much larger dataset, reduced predictor multicollinearity, and continuous parameter coverage, highlighting the utility of validated simulation models for developing robust predictive tools for biomass preprocessing applications.

09 - BIOMASS FUELS↗

Interpreting Write Performance of Supercomputer I/O Systems with Regression Models

This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10× on some samples for both of the target systems.

Xie, Bing↗

Identifying microbial drivers in biological phenotypes with a Bayesian network regression model

Abstract In Bayesian Network Regression models, networks are considered the predictors of continuous responses. These models have been successfully used in brain research to identify regions in the brain that are associated with specific human traits, yet their potential to elucidate microbial drivers in biological phenotypes for microbiome research remains unknown. In particular, microbial networks are challenging due to their high dimension and high sparsity compared to brain networks. Furthermore, unlike in brain connectome research, in microbiome research, it is usually expected that the presence of microbes has an effect on the response (main effects), not just the interactions. Here, we develop the first thorough investigation of whether Bayesian Network Regression models are suitable for microbial datasets on a variety of synthetic and real data under diverse biological scenarios. We test whether the Bayesian Network Regression model that accounts only for interaction effects (edges in the network) is able to identify key drivers (microbes) in phenotypic variability. We show that this model is indeed able to identify influential nodes and edges in the microbial networks that drive changes in the phenotype for most biological settings, but we also identify scenarios where this method performs poorly which allows us to provide practical advice for domain scientists aiming to apply these tools to their datasets. BNR models provide a framework for microbiome researchers to identify connections between microbes and measured phenotypes. We allow the use of this statistical model by providing an easy‐to‐use implementation which is publicly available Julia package at https://github.com/solislemuslab/BayesianNetworkRegression.jl .

59 BASIC BIOLOGICAL SCIENCES↗

Regression models using shapes of functions as predictors

Functional variables are often used as predictors in regression problems. A commonly used parametric approach, called scalar-on-function regression, uses the $\mathbb L^2$ inner product to map functional predictors into scalar responses. This method can perform poorly when predictor functions contain undesired phase variability, causing phases to have disproportionately large influence on the response variable. One past solution has been to perform phase–amplitude separation (as a pre-processing step) and then use only the amplitudes in the regression model. In this paper, we propose a more integrated approach, termed elastic functional regression model (EFRM), where phase-separation is performed inside the regression model, rather than as a pre-processing step. This approach generalizes the notion of phase in functional data, and is based on the norm-preserving time warping of predictors. Due to its invariance properties, this representation provides robustness to predictor phase variability and results in improved predictions of the response variable over traditional models. We demonstrate this framework using a number of datasets involving gait signals, NMR data, and stock market prices.

97 MATHEMATICS AND COMPUTING↗

A Robust Segmented Mixed Effect Regression Model for Baseline Electricity Consumption Forecasting

Renewable energy production has been surging around the world in recent years. To mitigate the increasing uncertainty and intermittency of the renewable generation, proactive demand response algorithms and programs are proposed and developed to further improve the utilization of load flexibility and increase the efficiency of power system operation. One of the biggest challenges to efficient control and operation of demand response resources is how to forecast the baseline electricity consumption and estimate the load impact from demand response resources accurately. In this paper, we propose a mixed effect segmented regression model and a new robust estimate for forecasting the baseline electricity consumption in Southern California, USA, by combining the ideas of random effect regression model, segmented regression model, and the least trimmed squares estimate. Since the log-likelihood of the considered model is not differentiable at breakpoints, we propose a new backfitting algorithm to estimate the unknown parameters. The estimation performance of the new estimation procedure has been demonstrated with both simulation studies and the real data application for the electric load baseline forecasting in Southern California.

42 ENGINEERING↗

Comparing Designed Training Sets to Optimize Multivariate Regression Models for Pr, Nd, and Nitric Acid Using Spectrophotometry

Chemometric regression models were developed for the quantification of praseodymium (Pr, 0–1000 µg/mL), neodymium (Nd, 0–1000 µg/mL), and nitric acid (HNO 3 , 0.1–5 M) using spectrophotometry. Designed calibration sets were composed of 20 samples each: 10 model points and 10 lack-of-fit (LOF) points. The D-optimal designs effectively minimized the number of samples required to build models, and each design resulted in similar prediction performance, suggesting that statistical design of experiments can provide a reliable framework for selecting training set samples in three-variable systems. Partial least squares regression (PLSR) models were validated against a one-factor-at-a-time validation set composed of 125 samples (three variables, five levels). The top PLS-1 models resulted in average percent root mean square error of prediction error values of 3.5%, 1.7%, and 1.2% for Pr(III), Nd(III), and HNO 3 , respectively. Power set augmentations of the model and LOF samples were investigated to optimize the number of training set samples. PLSR models built using just required model points (10) had similar predictive capabilities as models including the LOF points (20) but with fewer samples. The number of validation samples was also varied systematically to learn how many samples are needed to validate regression models. This work addresses long-standing questions in the field of chemometrics to help make this approach amenable to the near-real-time quantification of hazardous species in remote settings.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Improving Prediction of Peroxide Value of Edible Oils Using Regularized Regression Models

We present four unique prediction techniques, combined with multiple data pre-processing methods, utilizing a wide range of both oil types and oil peroxide values (PV) as well as incorporating natural aging for peroxide creation. Samples were PV assayed using a standard starch titration method, AOCS Method Cd 8-53, and used as a verified reference method for PV determination. Near-infrared (NIR) spectra were collected from each sample in two unique optical pathlengths (OPLs), 2 and 24 mm, then fused into a third distinct set. All three sets were used in partial least squares (PLS) regression, ridge regression, LASSO regression, and elastic net regression model calculation. While no individual regression model was established as the best, global models for each regression type and pre-processing method show good agreement between all regression types when performed in their optimal scenarios. Furthermore, small spectral window size boxcar averaging shows prediction accuracy improvements for edible oil PVs. Best-performing models for each regression type are: PLS regression, 25 point boxcar window fused OPL spectral information RMSEP = 2.50; ridge regression, 5 point boxcar window, 24 mm OPL, RMSEP = 2.20; LASSO raw spectral information, 24 mm OPL, RMSEP = 1.80; and elastic net, 10 point boxcar window, 24 mm OPL, RMSEP = 1.91. The results show promising advancements in the development of a full global model for PV determination of edible oils.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Meta‐Analysis and Regression Modeling of the Impacts of Four Indoor Environmental Quality Metrics on Office Performance

Awareness of how buildings interact with the occupant experience—especially human performance—is becoming more prevalent, as seen by increasing interest and investment in healthy built environments. However, there is a need to synthesize the wide array of existing indoor environmental assessment and performance research in a way that can translate directly to building design and operation. Existing research in this area typically focuses on a single isolated metric and has not focused on making the results utilizable by building practitioners. The aim of this research is to investigate existing office performance literature through meta‐analyses and produce regression models for four indoor environmental quality (IEQ) metrics to support critical decision‐making for building operation and renovation. To reach this aim, a literature review was conducted to identify studies that measure the impact of changing ventilation rate, temperature, horizontal illuminance, and noise level in offices on occupant task performance. This repository of field and laboratory studies was analyzed to visualize the trends between the selected IEQ metrics and task performance. The temperature, ventilation rate, and horizontal illuminance regression models showed clear improvement potential when modifying indoor conditions toward the defined high‐performance range, while the regression model for noise level was inconclusive. The discussion notes the importance of designing holistically for all components of these IEQ categories to utilize the results, for example, good filtration on outdoor air for quantifying ventilation impact and uniform overhead lighting with low contrast for quantifying horizontal illuminance impact. The novelty of this work is in considering multiple facets of the indoor environment under a single, unified analysis schema and producing IEQ‐based performance gains that can directly inform cost‐benefit analyses of building design and renovation.

60 APPLIED LIFE SCIENCES↗

Parton labeling without matching: unveiling emergent labelling capabilities in regression models

Parton labeling methods are widely used when reconstructing collider events with top quarks or other massive particles. State-of-the-art techniques are based on machine learning and require training data with events that have been matched using simulations with truth information. In nature, there is no unique matching between partons and final state objects due to the properties of the strong force and due to acceptance effects. We propose a new approach to parton labeling that circumvents these challenges by recycling regression models. The final state objects that are most relevant for a regression model to predict the properties of a particular top quark are assigned to said parent particle without having any parton-matched training data. This approach is demonstrated using simulated events with top quarks and outperforms the widely used $χ$ 2 method.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Classification and regression models of audio and vibration signals for machine state monitoring in precision machining systems

Here we present a data-driven method for monitoring machine status in manufacturing processes. Audio and vibration data from precision machining are used for inference in two operating scenarios: (a) variable machine health states (anomaly detection); and (b) settings of machine operation (state estimation). Audio and vibration signals are first processed through Fast Fourier Transform and Principal Component Analysis to extract transformed and informative features. These features are then used in the training of classification and regression models for machine state monitoring. Specifically, three classifiers (K-nearest neighbors, convolutional neural networks and support vector machines) and two regressors (support vector regression and neural network regression) were explored, in terms of their accuracy in machine state prediction. It is shown that the audio and vibration signals are sufficiently rich in information about the machine that 100% state classification accuracy could be accomplished. Data fusion was also explored, showing overall superior accuracy of data-driven regression models.

42 ENGINEERING↗

Machine Learning-Based Regression Models for Ironmaking Blast Furnace Automation

Computational fluid dynamics (CFD)-based simulation has been the traditional way to model complex industrial systems and processes. One very large and complex industrial system that has benefited from CFD-based simulations is the steel blast furnace system. The problem with the CFD-based simulation approach is that it tends to be very slow for generating data. The CFD-only approach may not be fast enough for use in real-time decisionmaking. To address this issue, in this work, the authors propose the use of machine learning techniques to train and test models based on data generated via CFD simulation. Regression models based on neural networks are compared with tree-boosting models. In particular, several areas (tuyere, raceway, and shaft) of the blast furnace are modeled using these approaches. The results of the model training and testing are presented and discussed. The obtained R 2 metrics are, in general, very high. The results appear promising and may help to improve the efficiency of operator and process engineer decisionmaking when running a blast furnace.

97 MATHEMATICS AND COMPUTING↗

Estimation of conditional cumulative incidence functions under generalized semiparametric regression models with missing covariates, with application to analysis of biomarker correlates in vaccine trials

Herein, this article presents generalized semiparametric regression models for conditional cumulative incidence functions with competing risks data when covariates are missing by sampling design or happenstance. A doubly robust augmented inverse probability weighted (AIPW) complete-case approach to estimation and inference is investigated. This approach modifies IPW complete-case estimating equations by exploiting the key features in the relationship between the missing covariates and the phase-one data to improve efficiency. An iterative numerical procedure is derived to solve the nonlinear estimating equations. The asymptotic properties of the proposed estimators are established. A simulation study examining the finite-sample performances of the proposed estimators shows that the AIPW estimators are more efficient than the IPW estimators. The developed method is applied to the RV144 HIV-1 vaccine efficacy trial to investigate vaccine-induced IgG binding antibodies to HIV-1 as correlates of acquisition of HIV-1 infection while taking account of whether the HIV-1 sequences are near or far from the HIV-1 sequences represented in the vaccine construct.

97 MATHEMATICS AND COMPUTING↗

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

Corresponding Standard Reference Material Data used in Partial Least Squares Regression Models for Sugar Composition Estimates in Biomass in: Economic Impact of Yield and Composition Variation in Bioenergy Crops: Populus trichocarpa

Corresponding Standard Reference Material Data used in Partial Least Squares Regression Models for Sugar Composition Estimates in Biomass in: Economic Impact of Yield and Composition Variation in Bioenergy Crops: Populus trichocarpa (for corresponding manuscript: DOI: 10.1002/bbb.2148) PDF Files: Images of 1H NMR spectra for neutralized 2-stage acid hydrolysates of 4 NIST Standard Reference Material biomass samples (Monterey Pine 8493, Sugarcane Bagasse 8491, Wheat Straw 8494, and Eastern Cottonwood/Poplar 8492) and 2 Center for Bioenergy Innovation reference biomass samples (Poplar - Populus trichocarpa and Switchgrass - Panicum Virgatum). Suppression of the water peak was achieved using a NOESY-1D with presaturation, a recycle delay of 5 s, and a total of 64 scans. Spectra were acquired at 298 K and processed with automatic phase correction, baseline correction, and chemical shift referencing to TSP-d4. Images show all 1H data from 10 to 1ppm with inset spectra of region of interest (4.0 to 3.1 ppm). Text Files: Spectra for neutralized 2-stage acid hydrolysates of 4 NIST Standard Reference Material biomass samples (Monterey Pine 8493, Sugarcane Bagasse 8491, Wheat Straw 8494, and Eastern Cottonwood/Poplar 8492) and 2 Center for Bioenergy Innovation reference biomass samples (Poplar - Populus trichocarpa and Switchgrass - Panicum Virgatum) were converted into text files for plotting. Files contain 8192 points of raw spectral data from 12.23 to -2.78 ppm. The text file contains 4 columns of data and includes: Point number, Intensity, Hz, and ppm. Xcel Spreadsheet: HPLC measured monomeric sugar concentrations and bucketed 1H NMR data used to build monomeric sugar composition prediction models. Sugar composition in biomass determined from HPLC analyses are given in mg sugar/mg of biomass. Spectral bucketing was performed using Bruker’s AMIX software. Spectra were divided into 0.005 ppm buckets in the region of 3.10– 4.15 ppm for a total of 210 buckets. Headers for the bucketed data are the chemical shift in ppm of the center of the bucket. Bucketed data was used to build partial least squares models for subsequent predictions in The Unscrambler v. 10.5(CAMO A/S, Trondheim, Norway). The formation of methanol during hydrolysis interferes with the quantitative NMR analysis of sugars, so the methanol peak centered at 3.37 ppm and spanning four buckets (3.2925 – 3.2775 ppm) was set to zero for all spectra.

09 BIOMASS FUELS↗

Proxy quality control of biomass particles using thermogravimetric analysis and Gaussian process regression models

Abstract The temperature experienced by reactants during preparation in a reactor is a key component in determining the yield and homogeneity of usable chemical products such as biomass particles. Thermocouples with sensors can be used to monitor spatial temperature gradients within reactors but these sensors are often too expensive and/or invasive. The present work proposes a strategy to identify optimal machine learning models to infer the maximum effective temperature experienced by particles during oxidative biomass torrefaction using key thermochemical combustion parameters. The maximum rate of weight loss, the corresponding temperature, and fixed carbon content on a dry‐ash‐free basis are used as literature‐based predictor variables obtained from thermogravimetric analysis. The evaluation of 24 machine‐learning models using the standard tenfold cross‐validation method suggests that the exponential Gaussian process regression (GPR) model is the most effective, followed by other GPR models. These high‐performing GPR models were also utilized to predict the effective preparation temperature distribution of reactor‐produced biomass particles under eight conditions of varying residence time and air‐to‐biomass ratio. The effective preparation temperature and residence time of individual biomass particles were then encoded into the torrefaction severity factor and used to estimate the energy yield of the reactor output as a novel quality control method. © 2023 The Authors. Biofuels, Bioproducts and Biorefining published by Society of Industrial Chemistry and John Wiley & Sons Ltd.

09 BIOMASS FUELS↗

A crystal-plasticity-informed Gaussian Process Regression model to capture anisotropy in single crystal shape memory alloys

This work presents a machine learning (ML) framework that model the anisotropic actuation responses in a shape memory alloy. A Gaussian Process Regression (GPR) based ML model is trained on a set of different crystal orientations subjected to different actuation conditions. The training employed thermo-mechanical responses from a crystal-plasticity model that captures phase-transformation, stress-induced plasticity, and transformation-induced plasticity. Further, on training the GPR-ML model at fixed stress level for different orientations, it captured the thermo-mechanical responses accounting for the anisotropy, and predicted responses for new orientations with good accuracy. The GPR-ML model is able to capture the transformation temperature variations even when trained using multiple stress levels, and the transformation strain showed significant deviations. The developed GPR-ML model gave reasonable predictions for an unexplored sample set of orientations and loading conditions.

36 MATERIALS SCIENCE↗

A novel probabilistic regression model for electrical peak demand estimate of commercial and manufacturing buildings

Due to the high cost of electricity in commercial and industrial sectors, demand forecast models have gained increasing attention. However, there are two unresolved issues: (1) Models are not adaptable when exposed to previously unknown data (2) The value of regression methods vs. state-of-the-art machine learning models has not been made apparent before. This study’s goal is to develop probabilistic demand estimation models. Herein, we propose a probabilistic Bayesian regression framework that can not only estimate future demands with high accuracy but also be updated once new information is available. By applying the proposed algorithm to two real-world case studies (commercial and manufacturing), we show a 40.3% and 30.8% improvement in terms of mean absolute error for the two cases. Moreover, the proposed technique outperforms powerful machine learning approaches, including support vector machine by 10.39%, random forest by 6.17%, and multilayer perceptron by 9.14% in terms of mean absolute percentage error.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗