Engineering PapersSearch

SEARCH · Engineering Papers

Results for “logistic regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Exploring the environmental drivers of human blastomycosis cases in the Midwestern United States

Blastomycosis is a fungal infection endemic to the eastern United States (US) and Canada caused by the inhalation of the fungi Blastomyces spp. Currently, the environmental drivers of disease dynamics are poorly understood. The goal of our work was to explore what environmental conditions are associated with the annual presence of blastomycosis cases, and therefore are potentially explanatory of the ecological niche of Blastomyces. We examined the relationships between reported cases of blastomycosis in three Midwestern US states (Michigan, Minnesota, and Wisconsin) from 2007–2017 in relation to eleven hypothesized environmental conditions, including climate, stream and soil mineral content, and land cover variables. Then, we fit logistic regression models to explore the relationships between the environmental variables and yearly blastomycosis case occurrence. Mean soil moisture, stream sediment mercury content, percent of water within the county, and woody wetlands land cover were all positively associated with the presence of annual cases, with woody wetlands having the most consistent signal across the three states. We also found significant differences in the likelihood of case presence between US states that were not explained by the variables in our model, suggesting state-level differences in case reporting and disease awareness. Our results provide a perspective on potential biological hypotheses to further test regarding environmental controls on the life cycle and ecological niche of Blastomyces.

54 ENVIRONMENTAL SCIENCES

Open Specy 1.0: Automated (Hyper)spectroscopy for Microplastics

Microplastic spectral analysis is one of the most time-consuming processes in studying microplastic pollution, often requiring days per sample. Researchers are transitioning to automated batch and hyperspectral image analysis techniques to enhance efficiency. Open Specy, initially aimed at manual single-spectrum analysis, has now integrated automated methods. This updated version, Open Specy 1.0, introduces several new features, including two algorithms for automated processing (smoothing and particle compression), an extensive library containing over 40,000 open-source Raman and FTIR spectra, and two machine learning classifiers (logistic regression and k medoids) developed from this library. Furthermore, it includes a revamped user interface, an R package, and a benchmark data set for testing future advancements in automated techniques. Researchers evaluated various configurations for hyperspectral smoothing, particle identification, compression, and splitting, to achieve combined recovery rates between 50 and 150% particle counts, identities, and sizes with a coefficient of variation (CV) of less than 40% (the accredited standard). Mean absorbance times the standard deviation provided a consistent particle identification. Hyperspectral smoothing led to a 96% combined recovery rate and reduced variability (CV = 38%) compared to the 86% recovery (CV = 83%) of nonsmoothed controls. Additionally, compressing spectra for particles was significantly faster (>3x) and showed similar accuracy but with reduced variability than processing each pixel individually. Key challenges persist in automating spectral analysis, particularly in refining particle splitting algorithms, and improving identification routines to minimize false positives and negatives. In conclusion, new methods in sample preparation for better stabilization and dispersion of particles could overcome some of these issues.

13 HYDRO ENERGY

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES

Lethality is Local, but Survival is Systemic: Temporal and Multi-Organ Responses to Chlorine Gas Exposure in a Murine Model

Chlorine gas (Cl2) is a highly toxic chemical associated with both localized lung injury and systemic health effects. While pulmonary damage has been well characterized, the systemic inflammatory and metabolic responses remain poorly understood. We aimed to define the temporal and multi-organ responses to Cl2 exposure in a murine model, with a focus on identifying spatiotemporal inflammation and its impact on survival and lethality. SKH1 mice were exposed for 10 min to varying concentrations of Cl2 (94.4–810 ppm, representative of non-lethal, LD10, and LD50 doses) and monitored for respiratory function, perfusion, and acidosis using organ-specific imaging. At multiple time points (40 min, 6 h, 24 h, and 7 d), we measured phosphoproteins, cytokines, chemokines, growth factors, and metabolic hormones in the lungs, heart, cortex, and plasma. Statistical modeling and logistic regression were used to identify biomarkers associated with lethality and survival. We found that lung injury was the primary cause of potential lethality, particularly via early phosphoprotein signaling disruptions. However, survival correlated with early systemic coordination of inflammatory and metabolic signals across organs. Perfusion and acidosis imaging were strongly associated with chemokine and hormone responses. Key survival-associated plasma biomarkers included decreased insulin, increased ghrelin, and decreased eotaxin. While potential lethality from Cl2 exposure is locally driven by pulmonary injury, survival depends on systemic, multi-organ responses that occur rapidly post-exposure. Within this model, our findings identify a potential therapeutic window to enhance survival and suggest candidate biomarkers that may be explored translationally for both triage and treatment of chlorine-related incidents.

chlorine gas

Independent and interactive effects of wet bulb globe temperature and air pollution exposures on suicide mortality

Background: Individual components of the ambient environment, such as temperature and air pollution, exist as part of a complex mixture and have been associated with suicide; however, their interactive effects remain poorly understood. This study examined the independent and interactive effects of wet bulb globe temperature (WBGT), nitrogen dioxide (NO 2 ), and fine particulate matter (PM 2.5 ) on suicide mortality. Methods: We identified 7,551 suicide cases in Utah, USA, from 2000 to 2016 and assigned exposure to daily maximum WBGT (sourced from the European Center for Medium-Range Weather Forecasts) and PM 2.5 and NO 2 concentrations (sourced from a national spatiotemporal ensemble model) using decedent’s residential address at the time of death. A case-crossover design with conditional logistic regression was used to estimate the independent and interactive effects of WBGT max , PM 2.5 , and NO 2 on suicide. For exposure windows, we considered single days preceding suicide (lag 0 to 6) and their averages across preceding days (lag 0–1, 0–3, and 0–6). Analyses were stratified by season. Results: We identified a significant association between WBGT max and suicide across all seasons (odds ratio [OR] = 1.05, 95% confidence interval [CI]: 1.01, 1.10; per 5 °C increase on lag 0–3 days). The associations were stronger in the warm season (March 22 to September 21), with ORs and 95% CIs ranging from 1.08 (1.02, 1.15) to 1.20 (1.10, 1.30) per 5 °C increase depending on the lag periods. We observed synergistic interactions between WBGT max and PM 2.5 and NO 2 in the warm season, associated with higher odds of suicide. The associations of WBGT max with suicide were most pronounced at high NO 2 levels. Conclusions: We found evidence of synergistic interactions between WBGT max and PM 2.5 and NO 2 on suicide in the warm season, emphasizing the need for considering the combined effects of heat stress and air pollution in suicide prevention strategies.

Fine particulate matter

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification

Investigating the Determinants of Household Capabilities Burden During Power Outages: The Case of Winter Storm Uri

Existing research primarily uses census data to identify the vulnerability of communities to hazards. These vulnerability indices provide aggregated data and are not hazard-specific nor well-validated with post-event data. In contrast, our study uses household survey data (n=1065) to understand which Texan households suffered the greatest loss of their capabilities due to power outages and other utility service disruptions during Winter Storm Uri. Inspired by the Capabilities Approach, our measures of burden include the number of household capability types disrupted during the outages (e.g., cooking, heating, refrigeration), the severity of impact for each disrupted capability, and the additional time and financial costs of coping with these disruptions. We perform a clustering analysis, and find two distinct groups in our data, consisting of ‘lesser burden' and ‘heavier burden' households. Results indicate that the households experiencing the heaviest capabilities burden were most likely to experience longer power outages and the loss of water services. They were also more likely to have a Hispanic-Latino household member, lack access to a generator, live in a rented home, have larger households with more young children, fewer adults over 65, lower household incomes, been impacted by the COVID-19 pandemic, and more family characteristics that made life harder. We also fit a logistic regression model to assess the role of outage, household, and community characteristics in predicting differences in capabilities burden. Our results offer insights into enumerating the consequences of utility service disruptions on households, which can inform more targeted and equitable resilience strategies.

24 POWER TRANSMISSION AND DISTRIBUTION

Unlocking nighttime mobility: Land use and accessibility in public transit for night commuters

Night commuters are integral to urban transportation systems. Essential services such as healthcare and manufacturing rely on workers who travel at night, and reliable mobility options are crucial for them. A gap exists in understanding how land use and accessibility influence public transportation use among night commuters. This study addresses this gap by using public data to explore land use and accessibility factors that affect night commuters' public transportation use in New York State. We investigated (1) the demographic characteristics of night commuters; (2) the influence of land use and accessibility on nighttime public transportation use; and (3) potential improvements to increase public transportation use and their impact. We combined data from the National Household Travel Survey with the Smart Location Database to link home locations with land use characteristics. Using logistic regression, we found that although females are generally less likely to be night commuters, they are more likely to use public transportation. Longer commute distances are associated with higher use of public transportation. Increasing job density along fixed-guideway transit routes and improving overall job accessibility via public transportation significantly enhances public transportation use among night commuters. In conclusion, this research provides actionable insights for public transportation agencies and urban planners to support night commuters, improving access and encouraging nighttime employment.

Job accessibility

Confidence-weighted integration of human and machine judgments for superior decision-making

Large language models (LLMs) can surpass humans in certain forecasting tasks. What role does this leave for humans in the overall decision process? One possibility is that humans, despite performing worse than LLMs, can still add value when teamed with them. A human and machine team can surpass each individual teammate when team members’ confidence is well calibrated and team members diverge in which tasks they find difficult (i.e., calibration and diversity are needed). We simplified and extended a Bayesian approach to combining judgments using a logistic regression framework that integrates confidence-weighted judgments for any number of team members. Using this straightforward method, we demonstrated its effectiveness in both image classification and neuroscience forecasting tasks. Combining human judgments with one or more machines consistently improved overall team performance. Our hope is that this simple and effective strategy for integrating the judgments of humans and machines will lead to productive collaborations.

97 MATHEMATICS AND COMPUTING

Impact of Extreme Heat on Emergency Department Admissions for Childhood and Adult Asthma: An Evaluation of Earth Observations and Heat Wave Definitions

Extreme heat has been associated with adverse health outcomes, yet its impact on asthma exacerbations remains understudied. This is, in part, due to data limitations: research that relies on weather station records and aggregated health statistics cannot resolve fine-scale differences in heat impacts. This study investigates the association between heat wave definitions and summertime asthma-related emergency department visits in Baltimore, Maryland from 2016 to 2022, including 819 adult and 695 pediatric exacerbations. Using geocoded electronic health records and air temperature measurements at several spatial resolutions, we applied a case-crossover design with conditional logistic regressions at the census block group and tract levels. We found strong associations between asthma exacerbations and nighttime heat wave definitions based on relative thresholds of minimum temperatures when census block group or tract level temperature estimates were used. These relationships were significant for both age groups and showed elevated risks in socially vulnerable areas. In contrast, heat wave definitions derived from the city's primary National Weather Service synoptic weather station show associations between asthma and daytime heat extremes, suggesting that the character of the heat hazard depends on the scale at which it is defined. The extreme heat event definition used by Baltimore City's Code Red system showed no significant association with exacerbations. These findings highlight the importance of data resolution in shaping health inferences related to extreme heat in urban environments. Further, this study demonstrates that, regardless of spatial scale, extreme heat is associated with asthma exacerbations in both age groups.

Corpuz, B. [Johns Hopkins University, Baltimore, M

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Chemistry

Derivative-free stochastic optimization via adaptive sampling strategies

In this paper, we present a novel derivative-free framework for solving unconstrained stochastic optimization problems. Many problems in fields ranging from simulation optimization to reinforcement learning to quantum computing involve settings where only stochastic function values are obtained via a zeroth-order oracle, which has no available gradient information and necessitates the usage of derivative-free optimization methodologies. Our approach includes estimating gradients using stochastic function evaluations and integrating adaptive sampling techniques to control the accuracy in these stochastic approximations. Our framework encapsulates several gradient estimation techniques, including standard finite-difference, Gaussian smoothing, sphere smoothing, randomized coordinate finite-difference, and randomized subspace finite-difference methods. We provide theoretical convergence guarantees for our framework and analyze the worst-case iteration and sample complexities associated with each gradient estimation method. Finally, we demonstrate the empirical performance of the methods on logistic regression and nonlinear least squares problems.

Adaptive sampling

STag. II. Classification of Serendipitous Supernovae Observed by Galaxy Redshift Surveys

With the number of supernovae observed expected to drastically increase thanks to large-scale surveys like the Dark Energy Spectroscopic Instrument (DESI), it is necessary that the tools we use to classify these objects keep up with this increase. We previously created Supernova Tagging and Classification (STag) to address this problem by employing machine learning techniques alongside logistic regression in order to assign “tags” to spectra based on spectral features. STag II is a continuation of this work, which now makes use of model supernova spectra combined with real DESI spectra in order to train STag to better deal with realistic data. Furthermore, we also make use of the rlap score as a trustworthiness cut, making for a more robust and accurate supernova classifier than before.

Astrostatistics techniques

Heat exposure and maternal stress: evidence from the GRAPHS pregnancy cohort in Ghana

Heat exposure has been linked to psychosocial stress, an established antecedent of perinatal depression; however, evidence on heat-related stress during pregnancy in sub-Saharan Africa remains limited. We analyzed psychosocial stress scores and covariate data from the Ghana Randomized Air Pollution and Health Study, linking daily maximum and minimum shaded wet bulb globe temperature (WBGT) metrics to participants’ stress scores derived from the Crisis in Family Systems-Revised Life Events Questionnaire. We evaluated associations using ordinal logistic regression of pregnancy-average and trimester-average exposures and distributed lag non-linear models (DLNMs) to assess time-varying associations across gestation. Higher average maximum WBGT exposure across pregnancy was associated with increased odds of higher psychosocial stress; each 1 °C increase in maximum WBGT was associated with 64% higher odds of belonging to a higher stress category (OR = 1.64; 95% CI = 1.17–2.31; p = 0.0040). In trimester-average models, higher first-trimester maximum WBGT was also associated with higher stress (OR = 1.44; 95% CI = 1.15–1.81; p = 0.0014). DLNMs suggested that relatively cooler daily maximum WBGT values (25th percentile) were associated with decreased odds of stress in early pregnancy, whereas extreme daily maximum WBGT values (99th percentile) showed a pattern consistent with increased odds of stress in mid-to-late gestation. These findings highlight gestational windows in which heat exposure may influence stress, emphasizing the need for further research into underlying mechanisms and effective interventions to protect maternal mental health in heat-vulnerable settings.

White, Lewis [Columbia University] (ORCID:00090005

Association of short-term ambient environmental exposures with suicide and drug overdose deaths among U.S. Veterans

This study examined associations between short-term ambient environmental exposures and suicide (n = 3210) and overdose mortality (n = 4293; 2226 opioid-related) among U.S. Veterans from 2018 to 2019. Daily exposure to 24-hour maximum temperature, average atmospheric pressure, average fine particulate matter (PM 2.5 ), 1-hour maximum nitrogen dioxide (NO 2 ), and 8-hour maximum ozone (O 3 ) was assessed at the decedent's county of residence on the day of death and up to 6 days prior. A national bi-directional, time-stratified case-crossover design was applied. Conditional logistic regression models estimated associations between each exposure and suicide or overdose deaths, overall, and stratified by season, region, elevation, and urbanicity. Over lag days 0-1, an interquartile range increase in maximum temperature was associated with increased suicide (19%) and overdose (27%) mortality, with stronger summer effects for suicide (55%) and overdose (71%). In winter, interquartile range increases in atmospheric pressure, PM 2.5 , and NO 2 were associated with 104%, 15%, and 19% increases in suicide mortality. Maximum temperature was associated with a 22% increase in suicide risk in metropolitan areas and 57% in the Western U.S., while NO 2 was associated with a 26% increase in overdose mortality in nonmetropolitan areas. Findings suggest environmental stressors contribute to suicide and overdose mortality among Veterans, supporting environmentally informed prevention efforts.

air pollution

A Physics-Aligned Multi-Domain Machine Learning Framework for Time-Localised Diagnosis of Power Electronics Faults

This paper presents a physics-aligned framework for fault diagnosis in multi-phase power-electronic systems using cycle-synchronous windowing and multi-domain features derived from Fourier, wavelet, and Hilbert–Huang representations. While both logistic regression and multilayer perceptron (MLP) models achieve perfect performance under standard evaluation, blind unseen testing reveals a critical failure in a baseline MLP. This is shown to arise from model selection based on validation accuracy. Using validation-loss-based selection restores correct unseen performance and improves confidence. Feature ablation shows that Fourier and wavelet features dominate, while computational analysis indicates that feature extraction, particularly HHT, governs runtime.

Kumar, Praveen [ORNL] (ORCID:0000000291877857)

Enhancing Automotive Intrusion Detection Through Multi-Modal Fusion: A CAN FD-LiDAR Approach

As vehicles become smarter and more autonomous, they increasingly depend on advanced sensors and communication technologies to operate securely. However, such growing dependence on technology—whether it’s CAN (Controller Area Network) for internal communication or LiDAR (Light Detection and Ranging) for sensing the world around them—also expands the attack surface for the types of cyber attacks. Traditional intrusion detection systems (IDS) typically monitor these systems in isolation, limiting their ability to detect sophisticated, crosssystem attacks. To address this, we propose a multi-modal fusion approach that combines real-world CAN FD signals (from the HCRL dataset) with LiDAR features (from the nuScenes dataset) to enhance attack detection. Our method employs a twostage ensemble approach. Calibrated XGBoost and LightGBM models initially process CAN FD (Fuzzing Data) and LiDAR data independently, detecting timing anomalies and space abnormalities. They are subsequently logarithmically combined with a logistic regression meta-model along with 17 engineered features capturing cross-modal behavior, prediction conflicts, and nonlinear interactions. This approach achieves an AUC of 0.87 and an F1-score of 0.82, surpassing single-modality baselines and early fusion methods, at merely 2 ms inference latency. Compared with deep learning competitors, it is 3 times more efficient, providing a lightweight, interpretable, and real time solution to automotive cybersecurity.

97 MATHEMATICS AND COMPUTING

Predictive analytics to direct clinical attention to complex patients with elevated suicide risk: enhancement of the Veterans Health Administration REACH VET model

Suicide is a major public health concern, particularly among Veterans. The U.S. Department of Veterans Affairs Veterans Health Administration (VHA) employs the Recovery Engagement and Coordination for Health–Veterans Enhanced Treatment (REACH VET) model to prioritise high-risk patients for targeted clinical attention. REACH VET 1.0 (RV 1.0) was developed on 2008–2011 data. To reflect changes in clinical practice and populations, VHA updated it to REACH VET 2.0 (RV 2.0). This study describes its development and validation. RV 2.0 used longitudinal data from 7,248,170 VHA patients (4,967 suicide deaths) in 2018–2019, with 650 time-varying demographic, clinical and area-level predictors derived from a 2-year lookback (2016–2019). An ensemble of Elastic-Net logistic regression models was trained on 2018 data and evaluated monthly at the population level in 2019, focusing on the top 0.1% intervention risk tier. Analyses assessed model discrimination, suicide detection, risk concentration, subgroup consistency (sex, age and race/ethnicity) and performance relative to RV 1.0 using the same percentile-based risk strata. RV 2.0 outperformed RV 1.0 across all risk strata, with better discrimination (C-statistic 0.76 vs 0.69) and consistent performance across demographic subgroups. Within the top 0.1% of predicted risk, RV 2.0 identified more deaths, higher suicide rates and greater mortality risk concentration both when averaged across the 12 monthly 2019 test sets (5.6 vs 3.6; 83.6 vs 53.7 per 100,000 person-years; 21.0 vs 14.1) and when annualised for 2019 (67 vs 43; 2.7% vs 1.7%; 1,003 vs 644 per 100,000 person-years; 26.7 vs 17.1). RV 2.0 improves suicide risk stratification among Veterans, demonstrating better performance and consistent prediction across subgroups and highlighting the need for regular model updates and evaluation.

Peluso, Alina [Oak Ridge National Laboratory (ORNL