Engineering PapersSearch

SEARCH · Engineering Papers

Results for “logistic regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

33 records · Page 2

Random forest models accurately classify synthetic opioids using high-dimensionality mass spectrometry datasets

Detection of novel threat agents presents several challenges, a principle one being the development of untargeted methods to screen an increasing number of threat chemicals whose exact structures are unknown. With the use of Machine Learning (ML) tools, we can guide the development of analytical methods for broad-spectrum detection of unbounded threat chemical families in complex mixtures. Toward this goal, we used nominal mass and high-resolution mass spectrometry data for hundreds of synthetic opioids and non-opioid compounds. We tested two ML techniques, logistic regression and random forest, to develop models towards a practical, implementable method for opioid detection. We found that of these tested ML methods, random forest models resulted in the highest validation accuracy (95+%) for both nominal mass and high-resolution classification of opioids versus non-opioids, with low false positive and false negative rates. The RF models were then used to successfully predict the classification of 10 compounds—five opioids and five non-opioids not part of the training and validation analysis. This application of ML is a critical step towards the development of field-deployable nominal mass spectrometers with ML-driven analyses for classification of emergent threats.

Arasteh, Kourosh [Lawrence Livermore National Labo

The Effect of Updraft Entrainment on Convective Cell Deepening in Realistic Large-Eddy Simulations

Entrainment of surrounding cooler and drier air into convective updrafts is one of the key processes that influence deep convection initiation and growth. Numerous studies have investigated the effect of entrainment on isolated convective cloud growth in idealized simulations, but the importance of this effect in realistic conditions with many interacting convective clouds remains uncertain. We examine the impact of entrainment on the depth reached by convective clouds in realistic large-eddy simulations (LES) over central Argentina during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign. Cloudy updrafts and their associated properties are assigned to convective cells tracked with radar reflectivity signatures. Several thousand convective cells are tracked over two high convective available potential energy (CAPE) and two low CAPE cases that support cells of varying depths and intensities. Entrainment is calculated explicitly as the fluxes of air into the outer surface of each cloudy updraft. Single-predictor logistic regression models are used to determine the relative importance of updraft, near-updraft, and preconvective initiation atmospheric conditions in predicting whether convective cells become deep. We then build a multiple-predictor regression model pairing important updraft and meteorological metrics with fractional entrainment rate. The probability of cells transitioning to deep convection is most sensitive to ambient 600-hPa relative humidity (42% of total metric contribution to cloud depth predictability), followed by low-level CAPE (28%), cloud-base updraft width (19%), and fractional entrainment (11%). Thus, the initial width of the updraft along with potential buoyancy and its dilution through the midtroposphere collectively determine whether deep convection will result from shallower clouds.

54 ENVIRONMENTAL SCIENCES

The impact of kidney function on Alzheimer’s disease blood biomarkers: implications for predicting amyloid-β positivity

Impaired kidney function has a potential confounding effect on blood biomarker levels, including biomarkers for Alzheimer’s disease (AD). Given the imminent use of certain blood biomarkers in the routine diagnostic work-up of patients with suspected AD, knowledge on the potential impact of comorbidities on the utility of blood biomarkers is important. We aimed to evaluate the association between kidney function, assessed through estimated glomerular filtration rate (eGFR) calculated from plasma creatinine and AD blood biomarkers, as well as their influence over predicting Aβ-positivity. We included 242 participants from the Translational Biomarkers in Aging and Dementia (TRIAD) cohort, comprising cognitively unimpaired individuals (CU; n = 124), mild cognitive impairment (MCI; n = 58), AD dementia (n = 34), and non-AD dementia (n = 26) patients all characterized by [ 18 F] AZD-4694. Plasma samples were analyzed for Aβ42, Aβ40, glial fibrillary acidic protein (GFAP), neurofilament light chain (NfL), tau phosphorylated at threonine 181 (p-tau181), 217 (p-tau217), 231 (p-tau231) and N-terminal containing tau fragments (NTA-tau) using Simoa technology. Kidney function was assessed by eGFR in mL/min/1.73 m 2 , based on plasma creatinine levels, age, and sex. Participants were also stratified according to their eGFR-indexed stages of chronic kidney disease (CKD). We evaluated the association between eGFR and blood biomarker levels with linear models and assessed whether eGFR provided added predictive value to determine Aβ-positivity with logistic regression models. Biomarker concentrations were highest in individuals with CKD stage 3, followed by stages 2 and 1, but differences were only significant for NfL, Aβ42, and Aβ40 (not Aβ42/Aβ40). All investigated biomarkers showed significant associations with eGFR except plasma NTA-tau, with stronger relationships observed for Aβ40 and NfL. However, after adjusting for either age, sex or Aβ-PET SUVr, the association with eGFR was no longer significant for all biomarkers except Aβ40, Aβ42, NfL, and GFAP. When evaluating whether accounting for kidney function could lead to improved prediction of Aβ-positivity, we observed no improvements in model fit (Akaike Information Criterion, AIC) or in discriminative performance (AUC) by adding eGFR to a base model including each plasma biomarker, age, and sex. While covariates like age and sex improved model fit, eGFR contributed minimally, and there were no significant differences in clinical discrimination based on AUC values. We found that kidney function seems to be associated with AD blood biomarker concentrations. However, these associations did not remain significant after adjusting for age and sex, except for Aβ40, Aβ42, NfL, and GFAP. While covariates such as age and sex improved prediction of Aβ-positivity, including eGFR in the models did not lead to improved prediction for any biomarker. Our findings indicate that renal function, within the normal to mild impairment range, does not seem to have a clinically relevant impact when using highly accurate blood biomarkers, such as p-tau217, in a biomarker-supported diagnosis.

60 APPLIED LIFE SCIENCES

Variability in Performance of a Machine Learning Seismicity Catalog: Central Italy, 2016–2017

Machine learning (ML) catalogs contain many more earthquakes than routine catalogs, but their performance in phase picking and earthquake detection has not been fully evaluated. We develop station‐level detection probabilities using logistic regression and combine them across a seismic network to compute spatial magnitude‐of‐completeness fields. We apply this approach to two catalogs from the 2016–2017 Central Italy sequence that were constructed from the same seismic network, one routine and one ML‐based. At the station level, the ML picker increases detection sensitivity by identifying smaller magnitude events and detecting earthquakes at greater distances. Spatially, the magnitude of completeness decreases substantially, with median values shifting from 1.6 to 0.5 for P waves and from 1.7 to 0.5 for S waves. However, the ML catalog also shows greater variability in station‐level performance than the routine catalog. These results demonstrate that ML‐based improvements in detectability are widespread but spatially nonuniform, highlighting their benefits, their limitations, and the potential for further improvements.

15 GEOTHERMAL ENERGY

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS

The relationship between below average cognitive ability at age 5 years and the child’s experience of school at age 9

Background At age 5, while only embarking on their educational journey, substantial differences in children’s cognitive ability will already exist. The aim of this study was to examine the causal association between below average cognitive ability at age 5 years and child-reported experience of school and self-concept, and teacher-reported class engagement and emotional-behavioural function at age 9 years. Methods This longitudinal cohort study used data from 7,392 children in the Growing Up in Ireland Infant Cohort, who had completed the Picture Similarities and Naming Vocabulary subtests of the British Abilities Scales at age 5. Principal components analysis was used to produce a composite general cognitive ability score for each child. Children with a general cognitive ability score more than 1 standard deviation (SD) below the mean at age 5 were categorised as ‘Below Average Cognitive Ability’ (BACA), and those scoring above this as ‘Typical Cognitive Development’ (TCD). The outcomes of interest, measured at age 9, were child-reported experience of school, child’s self-concept, teacher-reported class engagement, and teacher-reported emotional behavioural function. Binary and multinomial logistic regression models were used to examine the association between BACA and these outcomes. Results Compared to those with TCD, those with BACA had significantly higher odds of never liking school [Adjusted odds ratio (AOR) 1.82, 95% CI 1.37–2.43, p < 0.001], of being picked on (AOR 1.27, 95% CI 1.09–1.48) and of picking on others (AOR 1.53, 95% CI 1.27–1.84). They had significantly higher odds of experiencing low self-concept (AOR 1.20, 95% CI 1.02–1.42) and emotional-behavioural difficulties (AOR 1.34, 95% CI 1.10–1.63, p = 0.003). Compared to those with TCD, children with BACA had significantly higher odds of hardly ever or never being interested, motivated and excited to learn (AOR 2.29, 95% CI 1.70–3.10). Conclusion Children with BACA at school-entry had significantly higher odds of reporting a negative school experience and low self-concept at age 9. They had significantly higher odds of having teacher-reported poor class engagement and problematic emotional-behavioural function at age 9. The findings of this study suggest BACA has a causal role in these adverse outcomes. Early childhood policy and intervention design should be cognisant of the important role of cognitive ability in school and childhood outcomes.

Bowe, Andrea K.

Mining Product Reviews for Important Product Features of Refurbished iPhones

Problem: Remanufacturers want to increase consumer interest in refurbished products, which motivates the need to understand which product features are important to buyers of refurbished products such as mobile phones. Research Questions: This study addresses two questions. First, which product features are most important for buyers of refurbished iPhones? Second, how do those preferences differ from the preferences of buyers of new iPhones? Methods: Online reviews of iPhones are obtained and converted into a document–term matrix. Using this text model, three subsets of features are identified using statistical analysis of frequency of mention: most frequent, average, and least frequent. A logistic regression (LR) model is then used to identify which features are most predictive of whether a review is for a new or refurbished phone. Results: Buyers of refurbished phones mention battery health, screen/display, shell condition, and brand significantly more often than other features. Directly contrasting reviews of refurbished versus new phones shows that shell condition, brand, speaker, and charger are found to be the most predictive product features indicated in reviews for refurbished phones. Of those, the shell condition is significantly more predictive than the others. Implications: The results identify product features that remanufacturers of iPhones can emphasize to increase customer demand.

Anisi, Atefeh

Do Piperonyl Butoxide Long-Lasting Insecticide Treated Nets Provide Additional Protection Against Malaria Infections Compared with Conventional Nets in an Operational Setting in Western Kenya?

Malaria control in sub-Saharan Africa has stagnated despite widespread adoption of control measures such as long-lasting insecticidal nets (LLINs). Progress has stalled, in part, because of pyrethroid insecticide resistance, driving the need for retooling to increase the effectiveness of bed nets. Consequently, LLINs have been treated with the chemical synergist piperonyl butoxide (PBO). Piperonyl butoxide LLINs have been shown to be efficacious in controlled settings; however, their effectiveness in real-world settings warrants investigation. In Bungoma County, Western Kenya, a cohort of 768 participants was followed from June 2017 to December 2023 via active and passive surveillance. Household visits were conducted monthly, during which LLIN use for nets distributed in 2017 and 2021 was recorded, and symptomatic malaria cases were identified using rapid diagnostic tests (RDTs). The comparative effectiveness of PBO versus conventional LLINs was assessed in terms of malaria infections. A multilevel logistic regression model was fit with monthly RDT results as the dependent variable. The study results indicate that PBO LLINs provide greater protection against malaria at the individual level than conventional LLINs (odds ratio: 0.70; 95% CI: 0.47–1.03), although the findings were not statistically significant. The added protection against malaria infections provided by PBO LLINs compared with conventional LLINs observed in the current study aligns with findings from most previous studies, although this finding was not statistically significant. In areas with documented pyrethroid resistance, the use of LLINs with an added synergist, such as PBO, can provide additional protection against malaria infections (compared with pyrethroid-only LLINs) and should be considered for scaled-up scenarios despite the additional cost.

60 APPLIED LIFE SCIENCES

Edge ML for CAN bus intrusion detection in AVs

Autonomous Vehicles (AVs) are revolutionizing transportation, but their reliance on interconnected cyber-physical systems exposes them to unprecedented cybersecurity risks. This study addresses the critical challenge of detecting real-time cyber intrusions in self-driving vehicles by leveraging a dataset from the Udacity self-driving car project. We simulate four high-impact attack vectors, Denial of Service (DoS), spoofing, replay, and fuzzy attacks, by injecting noise into spatial features (e.g., bounding box coordinates) to replicate adversarial scenarios. We develop and evaluate two lightweight neural network architectures (NN-1 and NN-2) alongside a logistic regression baseline (LG-1) for intrusion detection. The models achieve exceptional performance, with NN-2 attaining an AUC score of 93.15% and 93.15% accuracy, demonstrating their suitability for edge deployment in AV environments. Through explainable AI techniques, we uncover unique forensic fingerprints of each attack type, such as spatial corruption in fuzzy attacks and temporal anomalies in replay attacks, offering actionable insights for feature engineering and proactive defense. Visual analytics, including confusion matrices, ROC curves, and feature importance plots, validate the models' robustness and interpretability. This research sets a new benchmark for AV cybersecurity, delivering a scalable, field-ready toolkit for Original Equipment Manufacturers (OEMs) and policymakers. By aligning intrusion fingerprints with SAE J3061 automotive security standards, we provide a pathway for integrating machine learning into safety-critical AV systems. Our findings underscore the urgent need for security-by-design AI, ensuring that AVs not only drive autonomously but also defend autonomously. This work bridges the gap between theoretical cybersecurity and life-preserving engineering, offering a leap toward safer, more secure autonomous transportation.

97 MATHEMATICS AND COMPUTING

Machine Learning for Predicting Team Functioning in HERA Missions

Team functioning is integral to success in future long term space exploration missions. Proactively detecting declines in team functioning can mitigate conflict and ensure mission success. This project developed a speech-based artificial intelligence (AI) system that unobtrusively predicts degradation in team functioning, including performance and cohesion, in the Human Exploration Research Analog (HERA) Campaigns 4 and 5. The AI system conducted automated analysis of the prosodic (tone of voice) and linguistic (language content) components of speech, modeling interpersonal dynamics at both the turn-taking and day-wide levels. We investigated team functioning via observing structured interactions (i.e., multi-mission space exploration vehicle-extra vehicular activity [MMSEV-EVA], team interaction battery [TIB]) and unstructured interactions before the MMSEV-EVA task. We developed machine learning models to predict team functioning (objective task accuracy, self reported team efficacy and self reported team cohesion) by analyzing OpenSmile acoustic features, linguistic descriptors extracted via the linguistic inquiry and word count (LIWC) dictionary, and semantic embeddings. In the TIB, static models using logistic regression and random forests were not able to predict task accuracy, but predicted team efficacy and cohesion during both the decision making and relational tasks to a moderate level (60-70%). Majority voting on the individual turns to predict day long team efficacy further increased accuracies (70-80%). Finally, long short-term memory (LSTM) models showed the best performance across all variables (80-91%), including task performance. In the MMSEV-EVA, static models achieved an accuracy of 60% with majority voting, which increased to 80% through the incorporation of mission day as a variable, accounting for the learning effect. A key finding across both tasks was the "team-dependent" nature of these interactions; models achieved much higher accuracy when trained on prior days of the same team's data rather than attempting to generalize across entirely different teams, with even 1-2 days of prior data per team achieving 5-15% improvement over team-independent models. In addition, the incorporation of pre-task data from the same team also improves model performance, e.g., incorporating data from the decision-making task of the TIB, which preceded the relational task, improved the prediction of team efficacy and cohesion during the latter. We compared model performance when trained on machine-generated data compared to data that had been further corrected by human annotators. Overall, models trained on human-corrected data exhibited a modest improvement in performance, particularly when acoustic features were used. We found no significant correlation between word error rate (WER) and model accuracy (r(55) = -0.08, p = 0.51), but model’s accuracy was significantly higher for medium/high quality transcription (0.74 (SD = 0.48)) compared to the low-quality group (0.64 (SD = 0.36)) (t(63)=2.82, p = 0.006). Based on these, several design recommendation emerge, that could inform Standards at NASA. Models predicting team functioning should incorporate at least one to two days of historical interaction data, include brief pre-task discussions, and explicitly model temporal learning effects, especially for longer operational tasks. Minimum quality standards for automated speech-processing pipelines are needed, given the performance gains observed with manually corrected acoustic data. Finally, systems should leverage both acoustic features and language embeddings in complementary ways, with modality choices and fusion strategies tailored to mission context, task demands, and data quality requirements.

Shrivatsa Mishra

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy

Quantifying Operational Drivers of Multimodal Biometric Verification in Aerial Surveillance

Multimodal biometric verification is increasingly applied across operational contexts ranging from close-range security cameras and building-mounted surveillance to long-range ground sensors and unmanned aerial system (UAS) imagery. Variations in acquisition conditions—such as image resolution, viewing geometry, and motion artifacts—pose significant challenges for cross-domain algorithmic generalization. This study evaluates two independent multimodal biometric verification systems developed under the Intelligence Advanced Research Projects Activity (IARPA) Biometric Recognition and Identification at Altitude and Range (BRIAR) program, comparing performance on close-range and aerial datasets. Close-range video served as a baseline to quantify the decline in verification performance on aerial footage. The dataset included six UAS platforms, spanning small quadcopters at 10m altitude to medium-sized fixed-wing aircraft at 360m. Mixed-effects logistic regression identified image resolution (head and body pixel counts), head height, sensor characteristics, and algorithm selection as primary determinants of verification success, whereas demographic attributes and mission gait were not significant predictors. Activity type and collection site influenced performance in close-range data but had negligible impact on UAS imagery. These results clarify modality-specific strengths and limitations and highlight opportunities to enhance cross-domain biometric verification.

Peluso, Alina [ORNL] (ORCID:0000000328950406)

Explaining word embeddings with perfect fidelity: a case study in predicting research impact

The best-performing approaches for scholarly document quality prediction are based on embedding models. In addition to their performance when used in classifiers, embedding models can also provide predictions even for words that were not contained in the labelled training data for the classification model, which is important in the context of the ever-evolving research terminology. Although model-agnostic explanation methods, such as Local interpretable model-agnostic explanations, can be applied to explain machine learning classifiers trained on embedding models, these produce results with questionable correspondence to the model. We introduce a new feature importance method, Self-Model Entities Rated (SMER), for logistic regression-based classification models trained on word embeddings. We show that SMER has theoretically perfect fidelity with the explained model, as the average of logits of SMER scores for individual words (SMER explanation) exactly corresponds to the logit of the prediction of the explained model. Quantitative and qualitative evaluation is performed through five diverse experiments conducted on 50,000 research articles (papers) from the CORD-19 corpus. In conclusion, through an AOPC curve analysis, we experimentally demonstrate that SMER produces better explanations than LIME, SHAP and global tree surrogates.

Coarse-grained models

Historical and Future Global Irrigation Energy Consumption by Fuel and Region

Irrigation energy use is a significant component of agricultural production costs, contributing directly to the energy and emissions intensity of crop production and ultimately to food prices. Understanding the existing structure of irrigation energy consumption help achieve food-energy-water security and environmental goals. We present a comprehensive global data set detailing country-level irrigation energy consumption, emphasizing the comparative use of electric, diesel, and emerging solar pumps. To our knowledge, no such data set exists. We draw from a literature review to develop a logistic transformed regression model to estimate the shares of fuel sources for irrigation across countries over historical years to construct a global data set of country-level irrigation energy consumption by multiple fuel sources. Additionally, we compare our estimates of irrigation energy use with agricultural energy use as reported by the International Energy Agency and other external sources. We then use this data to project future irrigation energy use with the Global Change Analysis Model, which is a multisector dynamics model, to showcase the usage of this data set. Projections under the reference scenario show a global shift in fuel types for irrigation pumping, while patterns vary across regions, with India and Pakistan leading in solar-powered irrigation growth and countries like the USA and China continuing to rely primarily on grid electricity. This data set provides a resource to understand the role of irrigation fuel choices within the broader energy sector, as well as the connected agricultural, land use, and water sectors under alternative future scenarios, enabling informed decision making toward efficient agricultural practices.

Global Change Analysis Model (GCAM)

Nonlinearity of the post-spinel transition and its expression in slabs and plumes worldwide

Phase transitions in the mantle control its internal dynamics and structure. The post-spinel transition marks the upper–lower mantle boundary, where ringwoodite dissociates into bridgmanite plus ferropericlase, and its Clapeyron slope regulates mantle flow across it. This interaction has previously been assumed to have no lateral spatial variations, based on the assumption of a linear post-spinel boundary in pressure and temperature. Here we present laser-heated diamond anvil cell experiments with synchrotron X-ray diffraction to better constrain this boundary, especially at higher temperatures. Combining our data with results from the literature, and using a global analysis based on machine learning, we find a pronounced nonlinearity in the post-spinel boundary, with its slope ranging from –4 MPa/K at 2100 K, to –2 MPa/K at 1950 K, and to 0 MPa/K at 1600 K. Changes in temperature over time and space can therefore cause the post-spinel transition to have variable effects on mantle convection and the movement of subducting slabs and upwelling plumes.

58 GEOSCIENCES