Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Linear Multivariable Regression Models for Prediction of Eddy Dissipation Rate from Available Meteorological Data

Linear multivariable regression models for predicting day and night Eddy Dissipation Rate (EDR) from available meteorological data sources are defined and validated. Model definition is based on a combination of 1997-2000 Dallas/Fort Worth (DFW) data sources, EDR from Aircraft Vortex Spacing System (AVOSS) deployment data, and regression variables primarily from corresponding Automated Surface Observation System (ASOS) data. Model validation is accomplished through EDR predictions on a similar combination of 1994-1995 Memphis (MEM) AVOSS and ASOS data. Model forms include an intercept plus a single term of fixed optimal power for each of these regression variables; 30-minute forward averaged mean and variance of near-surface wind speed and temperature, variance of wind direction, and a discrete cloud cover metric. Distinct day and night models, regressing on EDR and the natural log of EDR respectively, yield best performance and avoid model discontinuity over day/night data boundaries.

MCKissick, Burnell T.↗

The Use of Absolute-Value Terms in Regression Modeling of Multi-Piece Force Balances

Different aspects of the use of absolute-value terms in regression models of the electrical outputs of multi-piece force balance calibration data are discussed. First, characteristics of a variety of regression model term combinations with absolute-value terms are reviewed that are currently used in the aerospace testing community to fit the gage outputs of a balance. Then, a semi-empirical test is presented that quantifies bidirectional characteristics of the balance bridge outputs. Several diagnostic methods are discussed to assess the severity of near-linear dependencies between regressors of models with absolute-value terms. In particular, connections between the linear, absolute-value, quadratic, signed quadratic, and cubic terms are studied in greater detail. Data from an automated calibration of NASAs MK29B force balance are used to illustrate the most important observations and results. Rules of thumb that variance-inflation factors be less than 10 must be relaxed when using absolute-value terms to describe bidirectional balances.

calibration analysis↗

Meta‐Analysis and Regression Modeling of the Impacts of Four Indoor Environmental Quality Metrics on Office Performance

Awareness of how buildings interact with the occupant experience—especially human performance—is becoming more prevalent, as seen by increasing interest and investment in healthy built environments. However, there is a need to synthesize the wide array of existing indoor environmental assessment and performance research in a way that can translate directly to building design and operation. Existing research in this area typically focuses on a single isolated metric and has not focused on making the results utilizable by building practitioners. The aim of this research is to investigate existing office performance literature through meta‐analyses and produce regression models for four indoor environmental quality (IEQ) metrics to support critical decision‐making for building operation and renovation. To reach this aim, a literature review was conducted to identify studies that measure the impact of changing ventilation rate, temperature, horizontal illuminance, and noise level in offices on occupant task performance. This repository of field and laboratory studies was analyzed to visualize the trends between the selected IEQ metrics and task performance. The temperature, ventilation rate, and horizontal illuminance regression models showed clear improvement potential when modifying indoor conditions toward the defined high‐performance range, while the regression model for noise level was inconclusive. The discussion notes the importance of designing holistically for all components of these IEQ categories to utilize the results, for example, good filtration on outdoor air for quantifying ventilation impact and uniform overhead lighting with low contrast for quantifying horizontal illuminance impact. The novelty of this work is in considering multiple facets of the indoor environment under a single, unified analysis schema and producing IEQ‐based performance gains that can directly inform cost‐benefit analyses of building design and renovation.

60 APPLIED LIFE SCIENCES↗

Regression model estimation of early season crop proportions: North Dakota, some preliminary results

To estimate crop proportions early in the season, an approach is proposed based on: use of a regression-based prediction equation to obtain an a priori estimate for specific major crop groups; modification of this estimate using current-year LANDSAT and weather data; and a breakdown of the major crop groups into specific crops by regression models. Results from the development and evaluation of appropriate regression models for the first portion of the proposed approach are presented. The results show that the model predicts 1980 crop proportions very well at both county and crop reporting district levels. In terms of planted acreage, the model underpredicted 9.1 percent of the 1980 published data on planted acreage at the county level. It predicted almost exactly the 1980 published data on planted acreage at the crop reporting district level and overpredicted the planted acreage by just 0.92 percent.

Lin, K. K.↗

Parton labeling without matching: unveiling emergent labelling capabilities in regression models

Parton labeling methods are widely used when reconstructing collider events with top quarks or other massive particles. State-of-the-art techniques are based on machine learning and require training data with events that have been matched using simulations with truth information. In nature, there is no unique matching between partons and final state objects due to the properties of the strong force and due to acceptance effects. We propose a new approach to parton labeling that circumvents these challenges by recycling regression models. The final state objects that are most relevant for a regression model to predict the properties of a particular top quark are assigned to said parent particle without having any parton-matched training data. This approach is demonstrated using simulated events with top quarks and outperforms the widely used $χ$ 2 method.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Classification and regression models of audio and vibration signals for machine state monitoring in precision machining systems

Here we present a data-driven method for monitoring machine status in manufacturing processes. Audio and vibration data from precision machining are used for inference in two operating scenarios: (a) variable machine health states (anomaly detection); and (b) settings of machine operation (state estimation). Audio and vibration signals are first processed through Fast Fourier Transform and Principal Component Analysis to extract transformed and informative features. These features are then used in the training of classification and regression models for machine state monitoring. Specifically, three classifiers (K-nearest neighbors, convolutional neural networks and support vector machines) and two regressors (support vector regression and neural network regression) were explored, in terms of their accuracy in machine state prediction. It is shown that the audio and vibration signals are sufficiently rich in information about the machine that 100% state classification accuracy could be accomplished. Data fusion was also explored, showing overall superior accuracy of data-driven regression models.

42 ENGINEERING↗

Error analysis of leaf area estimates made from allometric regression models

Biological net productivity, measured in terms of the change in biomass with time, affects global productivity and the quality of life through biochemical and hydrological cycles and by its effect on the overall energy balance. Estimating leaf area for large ecosystems is one of the more important means of monitoring this productivity. For a particular forest plot, the leaf area is often estimated by a two-stage process. In the first stage, known as dimension analysis, a small number of trees are felled so that their areas can be measured as accurately as possible. These leaf areas are then related to non-destructive, easily-measured features such as bole diameter and tree height, by using a regression model. In the second stage, the non-destructive features are measured for all or for a sample of trees in the plots and then used as input into the regression model to estimate the total leaf area. Because both stages of the estimation process are subject to error, it is difficult to evaluate the accuracy of the final plot leaf area estimates. This paper illustrates how a complete error analysis can be made, using an example from a study made on aspen trees in northern Minnesota. The study was a joint effort by NASA and the University of California at Santa Barbara known as COVER (Characterization of Vegetation with Remote Sensing).

Feiveson, A. H.↗

Machine Learning-Based Regression Models for Ironmaking Blast Furnace Automation

Computational fluid dynamics (CFD)-based simulation has been the traditional way to model complex industrial systems and processes. One very large and complex industrial system that has benefited from CFD-based simulations is the steel blast furnace system. The problem with the CFD-based simulation approach is that it tends to be very slow for generating data. The CFD-only approach may not be fast enough for use in real-time decisionmaking. To address this issue, in this work, the authors propose the use of machine learning techniques to train and test models based on data generated via CFD simulation. Regression models based on neural networks are compared with tree-boosting models. In particular, several areas (tuyere, raceway, and shaft) of the blast furnace are modeled using these approaches. The results of the model training and testing are presented and discussed. The obtained R 2 metrics are, in general, very high. The results appear promising and may help to improve the efficiency of operator and process engineer decisionmaking when running a blast furnace.

97 MATHEMATICS AND COMPUTING↗

Estimation of conditional cumulative incidence functions under generalized semiparametric regression models with missing covariates, with application to analysis of biomarker correlates in vaccine trials

Herein, this article presents generalized semiparametric regression models for conditional cumulative incidence functions with competing risks data when covariates are missing by sampling design or happenstance. A doubly robust augmented inverse probability weighted (AIPW) complete-case approach to estimation and inference is investigated. This approach modifies IPW complete-case estimating equations by exploiting the key features in the relationship between the missing covariates and the phase-one data to improve efficiency. An iterative numerical procedure is derived to solve the nonlinear estimating equations. The asymptotic properties of the proposed estimators are established. A simulation study examining the finite-sample performances of the proposed estimators shows that the AIPW estimators are more efficient than the IPW estimators. The developed method is applied to the RV144 HIV-1 vaccine efficacy trial to investigate vaccine-induced IgG binding antibodies to HIV-1 as correlates of acquisition of HIV-1 infection while taking account of whether the HIV-1 sequences are near or far from the HIV-1 sequences represented in the vaccine construct.

97 MATHEMATICS AND COMPUTING↗

Statistical methods for efficient design of community surveys of response to noise: Random coefficients regression models

Research studies of residents' responses to noise consist of interviews with samples of individuals who are drawn from a number of different compact study areas. The statistical techniques developed provide a basis for those sample design decisions. These techniques are suitable for a wide range of sample survey applications. A sample may consist of a random sample of residents selected from a sample of compact study areas, or in a more complex design, of a sample of residents selected from a sample of larger areas (e.g., cities). The techniques may be applied to estimates of the effects on annoyance of noise level, numbers of noise events, the time-of-day of the events, ambient noise levels, or other factors. Methods are provided for determining, in advance, how accurately these effects can be estimated for different sample sizes and study designs. Using a simple cost function, they also provide for optimum allocation of the sample across the stages of the design for estimating these effects. These techniques are developed via a regression model in which the regression coefficients are assumed to be random, with components of variance associated with the various stages of a multi-stage sample design.

Tomberlin, T. J.↗

A componential model of human interaction with graphs: 1. Linear regression modeling

Task analyses served as the basis for developing the Mixed Arithmetic-Perceptual (MA-P) model, which proposes (1) that people interacting with common graphs to answer common questions apply a set of component processes-searching for indicators, encoding the value of indicators, performing arithmetic operations on the values, making spatial comparisons among indicators, and repsonding; and (2) that the type of graph and user's task determine the combination and order of the components applied (i.e., the processing steps). Two experiments investigated the prediction that response time will be linearly related to the number of processing steps according to the MA-P model. Subjects used line graphs, scatter plots, and stacked bar graphs to answer comparison questions and questions requiring arithmetic calculations. A one-parameter version of the model (with equal weights for all components) and a two-parameter version (with different weights for arithmetic and nonarithmetic processes) accounted for 76%-85% of individual subjects' variance in response time and 61%-68% of the variance taken across all subjects. The discussion addresses possible modifications in the MA-P model, alternative models, and design implications from the MA-P model.

Gillan, Douglas J.↗

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

Corresponding Standard Reference Material Data used in Partial Least Squares Regression Models for Sugar Composition Estimates in Biomass in: Economic Impact of Yield and Composition Variation in Bioenergy Crops: Populus trichocarpa

Corresponding Standard Reference Material Data used in Partial Least Squares Regression Models for Sugar Composition Estimates in Biomass in: Economic Impact of Yield and Composition Variation in Bioenergy Crops: Populus trichocarpa (for corresponding manuscript: DOI: 10.1002/bbb.2148) PDF Files: Images of 1H NMR spectra for neutralized 2-stage acid hydrolysates of 4 NIST Standard Reference Material biomass samples (Monterey Pine 8493, Sugarcane Bagasse 8491, Wheat Straw 8494, and Eastern Cottonwood/Poplar 8492) and 2 Center for Bioenergy Innovation reference biomass samples (Poplar - Populus trichocarpa and Switchgrass - Panicum Virgatum). Suppression of the water peak was achieved using a NOESY-1D with presaturation, a recycle delay of 5 s, and a total of 64 scans. Spectra were acquired at 298 K and processed with automatic phase correction, baseline correction, and chemical shift referencing to TSP-d4. Images show all 1H data from 10 to 1ppm with inset spectra of region of interest (4.0 to 3.1 ppm). Text Files: Spectra for neutralized 2-stage acid hydrolysates of 4 NIST Standard Reference Material biomass samples (Monterey Pine 8493, Sugarcane Bagasse 8491, Wheat Straw 8494, and Eastern Cottonwood/Poplar 8492) and 2 Center for Bioenergy Innovation reference biomass samples (Poplar - Populus trichocarpa and Switchgrass - Panicum Virgatum) were converted into text files for plotting. Files contain 8192 points of raw spectral data from 12.23 to -2.78 ppm. The text file contains 4 columns of data and includes: Point number, Intensity, Hz, and ppm. Xcel Spreadsheet: HPLC measured monomeric sugar concentrations and bucketed 1H NMR data used to build monomeric sugar composition prediction models. Sugar composition in biomass determined from HPLC analyses are given in mg sugar/mg of biomass. Spectral bucketing was performed using Bruker’s AMIX software. Spectra were divided into 0.005 ppm buckets in the region of 3.10– 4.15 ppm for a total of 210 buckets. Headers for the bucketed data are the chemical shift in ppm of the center of the bucket. Bucketed data was used to build partial least squares models for subsequent predictions in The Unscrambler v. 10.5(CAMO A/S, Trondheim, Norway). The formation of methanol during hydrolysis interferes with the quantitative NMR analysis of sugars, so the methanol peak centered at 3.37 ppm and spanning four buckets (3.2925 – 3.2775 ppm) was set to zero for all spectra.

09 BIOMASS FUELS↗

Updated Trends of the Stratospheric Ozone Vertical Distribution in the 60°S-60°N Latitude Range Based on the LOTUS Regression Model

This study presents an updated evaluation of stratospheric ozone profile trends in the 60°S - 60°N latitude range over the 2000 - 2020 period using an updated version of the Long-term Ozone Trends and Uncertainties in the Stratosphere (LOTUS) regression model that was used to evaluate such trends up to 2016 for the last WMO Ozone Assessment (2018). In addition to the derivation of detailed trends as a function of latitude and vertical coordinates, the regressions are performed with the data sets averaged over broad latitude bands, i.e., 60°S–35°S, 20°S–20°N and 35°N–60°N. The same methodology as in the last Assessment is applied to combine trends in these broad latitude bands in order to compare the results with the previous studies. Longitudinally resolved merged satellite records are also considered in order to provide a better comparison with trends retrieved from ground-based records, e.g., lidar, ozone sondes, Umkehr, microwave and Fourier Transform Infrared (FTIR) spectrometers at selected stations where long-term time series are available. The study includes a comparison with trends derived from the REF-C2 simulations of the Chemistry Climate Model Initiative (CCMI-1). This work confirms past results showing an ozone increase in the upper stratosphere, which is now significant in the three broad latitude bands. The increase is largest in the northern and southern hemisphere midlatitudes, with ~2.2%/decade at ~2.1 hPa, and ~2.1%/decade at ~3.2 hPa respectively, compared to ~1.6%/decade at ~2.6 hPa in the tropics. New trend signals have emerged from the records, such as a significant decrease of ozone in the tropics around 35 hPa and a non-significant increase of ozone in the southern midlatitudes at about 20 hPa. Non-significant negative ozone trends are derived in the lowermost stratosphere, with the most pronounced trends in the tropics. While a very good agreement is obtained between trends from merged satellite records and the CCMI-1 REF-C2 simulation in the upper stratosphere, observed negative trends in the lower stratosphere are not reproduced by models at southern and, in particular, at northern midlatitudes, where models report an ozone increase. However, the lower stratospheric trend uncertainties are quite large, for both measured and modelled trends. Finally, 2000-2020 stratospheric ozone trends derived from the ground-based and longitudinally resolved satellite records are in reasonable agreement over the European Alpine and tropical regions, while at the Lauder station in the southern hemisphere mid-latitudes they show some differences.

Stratospheric ozone trends↗

LACIE: Yield-weather regression models for the Canadian prairies

Most of the variability in wheat production is due to weather fluctuations. Climatic differences within the region account for a large portion of the variability in yields for different parts of the region. Separate regression models were developed for each of the areas indicated.

Source record↗

Dynamic and Regression Modeling of Ocean Variability in the Tide-Gauge Record at Seasonal and Longer Periods

Comparison of monthly mean tide-gauge time series to corresponding model time series based on a static inverted barometer (IB) for pressure-driven fluctuations and a ocean general circulation model (OM) reveals that the combined model successfully reproduces seasonal and interannual changes in relative sea level at many stations. Removal of the OM and IB from the tide-gauge record produces residual time series with a mean global variance reduction of 53%. The OM is mis-scaled for certain regions, and 68% of the residual time series contain a significant seasonal variability after removal of the OM and IB from the tide-gauge data. Including OM admittance parameters and seasonal coefficients in a regression model for each station, with IB also removed, produces residual time series with mean global variance reduction of 71%. Examination of the regional improvement in variance caused by scaling the OM, including seasonal terms, or both, indicates weakness in the model at predicting sea-level variation for constricted ocean regions. The model is particularly effective at reproducing sea-level variation for stations in North America, Europe, and Japan. The RMS residual for many stations in these areas is 25-35 mm. The production of "cleaner" tide-gauge time series, with oceanographic variability removed, is important for future analysis of nonsecular and regionally differing sea-level variations. Understanding the ocean model's strengths and weaknesses will allow for future improvements of the model.

Hill, Emma M.↗