Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “regression models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Remote Sensing-Based Estimation of Advanced Perennial Grass Biomass Yields for Bioenergy

A sustainable bioeconomy would require growing high-yielding bioenergy crops on marginal agricultural areas with minimal inputs. To determine the cost competitiveness and environmental sustainability of such production systems, reliably estimating biomass yield is critical. However, because marginal areas are often small and spread across the landscape, yield estimation using traditional approaches is costly and time-consuming. This paper demonstrates the (1) initial investigation of optical remote sensing for predicting perennial bioenergy grass yields at harvest using a linear regression model with the green normalized difference vegetation index (GNDVI) derived from Sentinel-2 imagery and (2) evaluation of the model’s performance using data from five U.S. Midwest field sites. The linear regression model using midsummer GNDVI predicted yields at harvest with R2 as high as 0.879 and a mean absolute error and root mean squared error as low as 0.539 Mg/ha and 0.616 Mg/ha, respectively, except for the establishment year. Perennial bioenergy grass yields may be predicted 152 days before the harvest date on average, except for the establishment year. The green spectral band showed a greater contribution for predicting yields than the red band, which is indicative of increased chlorophyll content during the early growing season. Although additional testing is warranted, this study showed a great promise for a remote sensing approach for forecasting perennial bioenergy grass yields to support critical economic and logistical decisions of bioeconomy stakeholders.

09 BIOMASS FUELS↗

Prescribed fires, smoke exposure, and hospital utilization among heart failure patients

Abstract Background Prescribed fires often have ecological benefits, but their environmental health risks have been infrequently studied. We investigated associations between residing near a prescribed fire, wildfire smoke exposure, and heart failure (HF) patients’ hospital utilization. Methods We used electronic health records from January 2014 to December 2016 in a North Carolina hospital-based cohort to determine HF diagnoses, primary residence, and hospital utilization. Using a cross-sectional study design, we associated the prescribed fire occurrences within 1, 2, and 5 km of the patients’ primary residence with the number of hospital visits and 7- and 30-day readmissions. To compare prescribed fire associations with those observed for wildfire smoke, we also associated zip code-level smoke density data designed to capture wildfire smoke emissions with hospital utilization amongst HF patients. Quasi-Poisson regression models were used for the number of hospital visits, while zero-inflated Poisson regression models were used for readmissions. All models were adjusted for age, sex, race, and neighborhood socioeconomic status and included an offset for follow-up time. The results are the percent change and the 95% confidence interval (CI). Results Associations between prescribed fire occurrences and hospital visits were generally null, with the few associations observed being with prescribed fires within 5 and 2 km of the primary residence in the negative direction but not the more restrictive 1 km radius. However, exposure to medium or heavy smoke (primarily from wildfires) at the zip code level was associated with both 7-day (8.5% increase; 95% CI = 1.5%, 16.0%) and 30-day readmissions (5.4%; 95% CI = 2.3%, 8.5%), and to a lesser degree, hospital visits (1.5%; 95% CI: 0.0%, 3.0%) matching previous studies. Conclusions Area-level smoke exposure driven by wildfires is positively associated with hospital utilization but not proximity to prescribed fires.

Raab, Henry↗

Predicting Damages to Remainder Parcels in Right-of-Way Acquisitions for Expanding Transportation Infrastructure: Using a Truncated Finite-Mixture Model

Right-of-way acquisition is a critical component of transportation infrastructure development. Transportation infrastructure projects cannot proceed without proper right-of-way acquisition or may face significant delays. State Departments of Transportation frequently acquire parcels of land for roadway expansion projects. A majority of these acquisitions can be partial takings, referring to a portion of a parcel that is acquired. The remainder of the property usually suffers economic changes due to the partial acquisition, which can be calculated as damage percentages. The damage percentage represents the extent to which the remaining land or property value has been diminished due to the acquisition. It reflects the remaining property value percentage that may have been lost or compromised due to the acquisition. Here, this study aims to provide a robust model to estimate damage percentages to the remainder parcels that may help state Departments of Transportation appraisers make early predictions about the damages in cases involving partial takings. The research uses 509 appraisal reports from the Tennessee Department of Transportation to identify the key parcel attributes that influence the percentage of damages. Three regression models are developed: a linear regression model, a finite-mixture model (FMM), and a truncated FMM with two latent classes. The modeling results show that the truncated FMM with two classes outperforms the other models. To validate the models, actual sales data is collected and analyzed for 59 properties, and the results suggest that the model predictions are fairly accurate. A predictive tool is developed based on the models to help appraisers anticipate right-of-way damages under different scenarios and can provide early predictions about the damages.

42 ENGINEERING↗

RxnRover/amlro

AMLRO (Active Machine Learning Reaction Optimizer) is an open-source framework designed to accelerate chemical reaction optimization using active learning with classical machine learning regression models. AMLRO integrates space-filling sampling strategies (e.g., Sobol and Latin Hypercube sampling) with iterative model training, prediction, and experiment selection to efficiently navigate complex reaction spaces. The platform supports multiple regression models, flexible multi-objective definitions, and user-defined parameter bounds, enabling data-efficient optimization from small initial datasets. AMLRO is designed for ease of use by experimentalists and can operate as a standalone decision-support tool or be integrated into closed-loop automated experimentation workflows.

Kulathunga, Dulitha Prasanna [Iowa State Universit↗

Restrictive spirometry pattern among construction trade workers

Spirometry-based studies of occupational lung disease have mostly focused on obstructive or mixed obstructive/restrictive outcomes. We wanted to determine if restrictive spirometry pattern (RSP) is associated with occupation and increased mortality. Study participants included 18,145 workers with demographic and smoking data and repeatable spirometry. The mortality analysis cohort included 15,445 workers with known vital status and cause of death through December 31, 2016. Stratified analyses explored RSP prevalence by demographic and clinical variables and trade. Log-binomial regression models explored RSP risk factors while controlling for important confounders such as smoking, obesity, and comorbidities. Cox regression models explored mortality risk by spirometry category. Prevalence of RSP was very high (28.6%). Mortality hazard ratios for RSP were 1.50 for all causes, 1.86 for cardiovascular diseases, 2.31 for respiratory diseases, and 1.66 for lung cancer. All construction trades except painters, machinists, and roofers had significantly elevated risk for RSP compared to our internal reference group. RSP was significantly associated with both parenchymal and pleural changes seen by chest X-ray. Construction trade workers are at significantly increased risk for RSP independent of obesity. Individuals with RSP are at increased risk for all-cause mortality as well as mortality attributable to respiratory diseases, cardiovascular diseases, and lung cancer. RSP deserves greater attention in occupational medicine and epidemiology.

60 APPLIED LIFE SCIENCES↗

Tea consumption and risk of bladder cancer in the Bladder Cancer Epidemiology and Nutritional Determinants (BLEND) Study: Pooled analysis of 12 international cohort studies

Tea has been shown to be associated with reduced risk of several diseases including cardiovascular diseases, stroke, metabolic syndrome, and obesity. However, the results on the relationship between tea consumption and bladder cancer are conflicting. This research aimed to assess the association between tea consumption and risk of bladder cancer using a pooled analysis of prospective cohort data. Individual data from 532,949 participants in 12 cohort studies, were pooled for analyses. Cox regression models stratified by study centre was used to estimate hazard ratios (HR) and corresponding 95% CIs. Fractional polynomial regression models were used to examine the dose–response relationship. A higher level of tea consumption was associated with lower risk of bladder cancer incidence (compared with no tea consumption: HR = 0.87, 95% C.I. = 0.77–0.98 for low consumption; HR = 0.86, 95% C.I. = 0.77–0.96 for moderate consumption; HR = 0.84, 95% C.I. = 0.75–0.95 for high consumption). When stratified by sex and smoking status, this reduced risk was statistically significant among men and current and former smokers. In addition, dose–response analyses showed a lower bladder cancer risk with increment of 100 ml of tea consumption per day (HR-increment = 0.97; 95% CI = 0.96–0.98). A similar inverse association was found among males, current and former smokers while never smokers and females showed non-significant results, suggesting potential sex-dependent effect. Higher consumption of tea is associated with reduced risk of bladder cancer with potential interaction with sex and smoking status. Further studies are needed to clarify the mechanisms for a protective effect of tea (e.g. inhibition of the survival and proliferation of cancer cells and anti-inflammatory mechanisms) and its interaction with smoking and sex.

60 APPLIED LIFE SCIENCES↗

Data-Centric Development of Lignin Structure–Solubility Relationships in Deep Eutectic Solvents Using Molecular Simulations

Lignin is a natural source of aromatic chemicals with significant potential as an abundant, renewable feedstock for value-added products. Deep eutectic solvents (DES)–solvents composed of a hydrogen bond donor (HBD) and acceptor (HBA) in varying ratios–have emerged as a highly tunable class of solvents for lignin solubilization. However, the variety of possible DES compositions and limited molecular-scale understanding of lignin solubility makes solvent selection a challenge without laborious trial-and-error experimentation. To address these challenges, we use classical molecular dynamics (MD) simulations to study the interactions of lignin model compounds with various DES–water systems. Quantitative parameters (descriptors) were calculated by postprocessing the MD results and used to train a regression model that predicts experimentally determined solubilities of lignin model compounds. This approach revealed that the most important descriptors of solubility are the system temperature, solute hydrophilicity, and metrics quantifying hydrogen bonding. Maximizing the interactions between solute–HBD (hydrophobic group), water–HBD (hydrophilic group), and water–HBA molecules led to the highest model compound solubility. Our results support a hydrotropic mechanism in which extensive DES–water hydrogen bonding and favorable HBD interactions with the solute promote high solubility. We applied the regression model derived using model compounds to predict the solubility of representative lignin oligomers. The model predicted lignin oligomers’ solubilities in good agreement with experiments, indicating that the simulations of model compounds can be extended to predict the solubility of larger lignin compounds across a range of solvent compositions and temperatures. Furthermore, these findings provide new molecular-scale insight into lignin solubilization mechanisms and a new method for computationally screening potential solvent systems for lignin valorization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluating Cell Temperature Models and the Effect of Wind Speed in PV System Capacity Testing: Preprint

Capacity testing is a routine procedure for assessing a photovoltaic system's performance relative to expectations. The most common test method involves fitting a regression model that predicts system output power using operating weather conditions including wind speed. Structural modifications to the regression model to incorporate wind in different ways improved the model's ability to fit measured system performance, but the observed improvements were small and unlikely to change the result of a capacity test. However, the results showed that the choice of reporting wind speed and inclusion or exclusion of wind speed in the performance model used as the test benchmark can significantly change the test result.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Applying machine learning and quantum chemistry to predict the glass transition temperatures of polymers

Glass transition temperature (T g ) is important for understanding the physical and mechanical properties of a polymer material because it relates to the thermal energy required to transition between a hard glassy state and a soft rubbery one. Over the years, various models have been developed for predicting this thermal property from molecular structure to aid in designing novel polymers in selected classes. This work builds on those efforts by utilizing both machine learning (ML) and quantum chemistry (QC) techniques to develop models that can predict T g values from the molecular structure under different data availability scenarios and for a wide variety of polymer types. For the ML model, a graph convolutional network (GCN) was used to map topological polymer features; this model was trained against a dataset of more than 7500 T g values and resulted in a root mean square error (RMSE) of 38.1 °C. The QC-based regression model was trained on 83 T g values and produced an RMSE of 34.5 °C. In conclusion, this work demonstrated that while both model techniques produce accurate predictions and are suitable for different data availability scenarios, the QC-based regression model offered a more interpretable model framework with significantly less training data.

36 MATERIALS SCIENCE↗

Analyzing count data with measurement error

In this article, we analyze observed count data such as the number of defects in a steel product where the observed counts are the true counts measured with errors. We account for the measurement error by using a measurement error model based on a latent lognormal (LLN) distribution. We consider making inference about a single population (e.g., from samples of a production lot) and a regression model (e.g., from runs of a designed experiment), where the measurement system properties are known, that is, the parameters of the LLN distribution are known. Then, we consider simultaneous inference for the single population and regression model as well as the measurement system. We demonstrate the proposed methodology with both simulated and real observed counts.

42 ENGINEERING↗

Machine learning prediction of self-diffusion in Lennard-Jones fluids

In this work, different machine learning (ML) methods were explored for the prediction of self-diffusion in Lennard-Jones (LJ) fluids. Using a database of diffusion constants obtained from the molecular dynamics simulation literature, multiple Random Forest (RF) and Artificial Neural Net (ANN) regression models were developed and characterized. The role and improved performance of feature engineering coupled to the RF model development was also addressed. The performance of these different ML models was evaluated by comparing the prediction error to an existing empirical relationship used to describe LJ fluid diffusion. It was found that the ANN regression models provided superior prediction of diffusion in comparison to the existing empirical relationships.

74 ATOMIC AND MOLECULAR PHYSICS↗

Assessing the Energy Resilience of Office Buildings: Development and Testing of a Simplified Metric for Real Estate Stakeholders

Increasing concern over higher frequency extreme weather events is driving a push towards a more resilient built environment. In recent years there has been growing interest in understanding how to evaluate, measure, and improve building energy resilience, i.e., the ability of a building to provide energy-related services in the event of a local or regional power outage. In addition to human health and safety, many stakeholders are keenly interested in the ability of a building to allow continuity of operations and minimize business disruption. Office buildings are subject to significant economic losses when building operations are disrupted due to a power outage. We propose “occupant hours lost” (OHL) as a means to measure the business productivity lost as the result of a power outage in office buildings. OHL is determined based on indoor conditions in each space for each hour during a power outage, and then aggregated spatially and temporally to determine the whole building OHL. We used quasi-Monte Carlo parametric energy simulations to demonstrate how the OHL metric varies due to different building characteristics across different climate zones and seasons. The simulation dataset was then used to develop simple regression models for assessing the impact of ten key building characteristics on OHL. The most impactful were window-to-wall ratio and window characteristics. The regression models show promise as a simple means to assess and screen for resilience using basic building characteristics, especially for non-critical facilities where it may not be viable to conduct detailed engineering analysis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Temporal deposition of copper and zinc in the sediments of metal removal constructed wetlands

The objective of this study was to explore the effects of time, seasons, and total carbon (TC) on Copper (Cu) and Zinc (Zn) deposition in the surface sediments. This study was performed at the H-02 constructed wetland on the Savannah River Site (Aiken, SC, USA). Covering both warm (April-September) and cool (October-March) seasons, several sediment cores were collected twice a year from the H-02 constructed wetland cells from 2007 to 2013. Total concentrations of Cu and Zn were measured in the sediments. Concentrations of Cu and Zn (mean ± standard deviation) in the surface sediments over 7 years of operation increased from 6.0 ± 2.8 and 14.6 ± 4.5 mg kg -1 to 139.6 ± 87.7 and 279.3 ± 202.9 mg kg -1 dry weight, respectively. The linear regression model explained the behavior and the variability of Cu deposition in the sediments. On the other hand, using the generalized least squares extension with the linear regression model allowed for unequal variance and thus produced a model that explained the variance properly, and as a result, was more successful in explaining the pattern of Zn deposition. Total carbon significantly affected both Cu ( p = 0.047) and Zn ( p < 0.001). Time effect on Cu deposition was statistically significant ( p = 0.013), whereas Zn was significantly affected by the season ( p = 0.009).

54 ENVIRONMENTAL SCIENCES↗

Predictive modeling of Néel temperature in austenitic alloys using CALPHAD and data analytics

The Néel temperature is a crucial yet often overlooked parameter in calculating the stacking fault energy (SFE) of austenitic alloys. Several empirical equations have been proposed to estimate the Néel temperature of austenitic alloys, which are then used to calculate the SFE and explain deformation mechanisms. However, these empirical equations, typically derived using linear regression algorithms, are often simplistic and may fail to capture the complex interactions among multiple alloying elements that influence the Néel temperature. Moreover, their applicability is usually limited to specific compositional ranges. In this study, we propose a CALPHAD based approach and develop a surrogate decision tree based regression model capable of capturing the interactions among multiple alloying elements to predict the Néel temperature. Predictions from both the CALPHAD approach and the regression model show close agreement with experimental measurements reported in the literature. In conclusion, the implications of accurate Néel temperature predictions on the calculated SFE and deformation mechanisms are also discussed.

36 MATERIALS SCIENCE↗

The Effect of Updraft Entrainment on Convective Cell Deepening in Realistic Large-Eddy Simulations

Entrainment of surrounding cooler and drier air into convective updrafts is one of the key processes that influence deep convection initiation and growth. Numerous studies have investigated the effect of entrainment on isolated convective cloud growth in idealized simulations, but the importance of this effect in realistic conditions with many interacting convective clouds remains uncertain. We examine the impact of entrainment on the depth reached by convective clouds in realistic large-eddy simulations (LES) over central Argentina during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign. Cloudy updrafts and their associated properties are assigned to convective cells tracked with radar reflectivity signatures. Several thousand convective cells are tracked over two high convective available potential energy (CAPE) and two low CAPE cases that support cells of varying depths and intensities. Entrainment is calculated explicitly as the fluxes of air into the outer surface of each cloudy updraft. Single-predictor logistic regression models are used to determine the relative importance of updraft, near-updraft, and preconvective initiation atmospheric conditions in predicting whether convective cells become deep. We then build a multiple-predictor regression model pairing important updraft and meteorological metrics with fractional entrainment rate. The probability of cells transitioning to deep convection is most sensitive to ambient 600-hPa relative humidity (42% of total metric contribution to cloud depth predictability), followed by low-level CAPE (28%), cloud-base updraft width (19%), and fractional entrainment (11%). Thus, the initial width of the updraft along with potential buoyancy and its dilution through the midtroposphere collectively determine whether deep convection will result from shallower clouds.

54 ENVIRONMENTAL SCIENCES↗

pnnl/simple-building-calculator

The Simple Building Calculator is a web application that estimates annual energy use using regression models that have been fit to simulation results for the DOE Commercial Prototype Building Models. The tool is a single-page application that will allow a user to rapidly analyze the impact of energy efficiency measures and design options on building performance. As it uses linear regression models rather than more complicated machine learning models or physics based simulation, the tool can deploy easily and run client-side.

Xu, Weili↗

Computationally efficient and error aware surrogate construction for numerical solutions of subsurface flow through porous media

Limiting the injection rate to restrict the pressure below a threshold at a critical location can be an important goal of simulations that model the subsurface pressure between injection and extraction wells. The pressure is approximated by the solution of Darcy’s partial differential equation for a given permeability field. The subsurface permeability is modeled as a random field since it is known only up to statistical properties. This induces uncertainty in the computed pressure. Solving the partial differential equation for an ensemble of random permeability simulations enables estimating a probability distribution for the pressure at the critical location. These simulations are computationally expensive, and practitioners often need rapid online guidance for real-time pressure management. An ensemble of numerical partial differential equation solutions is used to construct a Gaussian process regression model that can quickly predict the pressure at the critical location as a function of the extraction rate and permeability realization. The Gaussian process surrogate analyzes the ensemble of numerical pressure solutions at the critical location as noisy observations of the true pressure solution, enabling robust inference using the conditional Gaussian process distribution. Our first novel contribution is to identify a sampling methodology for the random environment and matching kernel technology for which fitting the Gaussian process regression model scales as O ( n log n ) instead of the typical O ( n 3 ) rate in the number of samples n used to fit the surrogate. The surrogate model allows almost instantaneous predictions for the pressure at the critical location as a function of the extraction rate and permeability realization. Our second contribution is a novel algorithm to calibrate the uncertainty in the surrogate model to the discrepancy between the true pressure solution of Darcy’s equation and the numerical solution. Finally, although our method is derived for building a surrogate for the solution of Darcy’s equation with a random permeability field, the framework broadly applies to solutions of other partial differential equations with random coefficients.

54 ENVIRONMENTAL SCIENCES↗