Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq↗

Chatter detection in simulated machining data: a simple refined approach to vibration data

Vibration monitoring is a critical aspect of assessing the health and performance of machinery and industrial processes. This study explores the application of machine learning techniques, specifically the Random Forest (RF) classification model, to predict and classify chatter—a detrimental self-excited vibration phenomenon—during machining operations. While sophisticated methods have been employed to address chatter, this research investigates the efficacy of a novel approach to an RF model. The study leverages simulated vibration data, bypassing resource-intensive real-world data collection, to develop a versatile chatter detection model applicable across diverse machining configurations. The feature extraction process combines time-series features and Fast Fourier Transform (FFT) data features, streamlining the model while addressing challenges posed by feature selection. By focusing on the RF model’s simplicity and efficiency, this research advances chatter detection techniques, offering a practical tool with improved generalizability, computational efficiency, and ease of interpretation. The study demonstrates that innovation can reside in simplicity, opening avenues for wider applicability and accelerated progress in the machining industry.

42 ENGINEERING↗

A new method for predicting hurricane rapid intensification based on co-occurring environmental parameters

Abstract Tropical cyclones (TCs) that undergo Rapid Intensification (RI) can pose serious socioeconomic threats and can potentially result in major damaging impacts along coastal areas. Considering the complexity of various physical mechanisms that play a role in RI and its relatively low probability of occurrence, predicting RI remains a major operational challenge. In this study, we propose a simple deterministic binary classification model based on the co-occurrence of environmental parameters (MCE) to predict an RI event. More specifically, the model determines the possibility of RI based on a simple count of the number of environmental predictors deemed favorable and unfavorable. We compare our model results to logistic regression (LR) and decision tree (DT) models, well-trained using the same set of environmental predictors. Results reveal that at an RI threshold of 30 kt, the MCE exhibits a critical success index score of 0.233 which is 14% higher than DT and LR model performances. When tested at multiple RI thresholds, the MCE displays relatively higher skill scores across multiple metrics. By simultaneously evaluating the favorability of predictors, the MCE is able to comparatively reduce the number of false alarms predicted when certain predictors are unfavorable toward RI. Interpreting these model results to gain a physical understanding of how co-occurring environmental parameters can affect RI, we highlight future directions for using models based on the MCE approach to understand and predict TC RI as well as other meteorological extremes.

54 ENVIRONMENTAL SCIENCES↗

Smart detection of indoor occupant thermal state via infrared thermography, computer vision, and machine learning

The ability to measure occupants’ thermal state in real time will enable major advances in the control of air conditioning systems. This study proposes predicting occupant thermal state by a combination of infrared thermography, computer vision, and machine learning. The approach (1) uses cheek, nose, and hand temperatures because they are least subject to blockage by hair, glasses, and clothing; (2) measures the distribution of skin temperatures within geometrically defined sub-areas of the face and hand; and (3) uses temperature differences within and between these areas to eliminate the effects of calibration drift that are unavoidable in thermal infrared (TIR) cameras. Two series of tests were conducted, respectively in an outdoor carport and an indoor environmental chamber, collecting a total of 48,422 sets of cheek, nose, and hand skin temperatures using a TIR camera and computer-vision technology, coupled with 715 subjective responses of thermal sensations. To predict occupant thermal state, Random Forest classification models were built using either absolute skin temperatures (the maximum and median temperatures of cheek and hand segments, and the temperature of the central spot on the nose), or intra- and inter-segment temperature differences of cheeks, hands, and nose. These measurements were found to accurately predict occupant thermal state. Using the maximum and median temperatures for cheek and nose, or for cheek and hand, predicts thermal state with an accuracy of 92–96%. In conclusion, using only the intra- and inter-segment temperature differences from cheek and nose is 83% accurate; adding the hand temperature differences increases the accuracy to 96%.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A machine learning approach for identifying variables associated with risk of developing neutralizing antidrug antibodies to factor VIII

A key unmet need in the management of hemophilia A (HA) is the lack of clinically validated markers that are associated with the development of neutralizing antibodies to Factor VIII (FVIII) (commonly referred to as inhibitors). This study aimed to identify relevant biomarkers for FVIII inhibition using Machine Learning (ML) and Explainable AI (XAI) using the My Life Our Future (MLOF) research repository. The dataset includes biologically relevant variables such as age, race, sex, ethnicity, and the variants in the F8 gene. In addition, we previously carried out Human Leukocyte Antigen Class II (HLA-II) typing on samples obtained from the MLOF repository. Using this information, we derived other patient-specific biologically and genetically important variables. These included identifying the number of foreign FVIII derived peptides, based on the alignment of the endogenous FVIII and infused drug sequences, and the foreign-peptide HLA-II molecule binding affinity calculated using NetMHCIIpan. The data were processed and trained with multiple ML classification models to identify the top performing models. The top performing model was then chosen to apply XAI via SHAP, (SHapley Additive exPlanations) to identify the variables critical for the prediction of FVIII inhibitor development in a hemophilia A patient. Using XAI we provide a robust and ranked identification of variables that could be predictive for developing inhibitors to FVIII drugs in hemophilia A patients. These variables could be validated as biomarkers and used in making clinical decisions and during drug development. The top five variables for predicting inhibitor development based on SHAP values are: (i) the baseline activity of the FVIII protein, (ii) mean affinity of all foreign peptides for HLA DRB 3, 4, & 5 alleles, (iii) mean affinity of all foreign peptides for HLA DRB1 alleles), (iv) the minimum affinity among all foreign peptides for HLA DRB1 alleles, and (v) F8 mutation type.

60 APPLIED LIFE SCIENCES↗

Actinides in complex reactive media: A combined ab initio molecular dynamics and machine learning analytics study of transuranic ions in molten salts

The predominant ionic chemistry and the similarity in ionic radius of actinides make it very difficult to structurally distinguish them in liquids. In this work, while ab initio molecular dynamics shows that the f-states clearly affects the electronic properties, their impact on structural properties is not obvious. For the series of trivalent actinides U 3+ , Pu 3+ , Cm 3+ , Cf 3+ , Fm 3+ in molten NaCl and FLiBe, actinide ligand bonds have a higher degree of covalency in NaCl (than in FLiBe), and a higher degree of ionicity in FLiBe. Furthermore, a machine learned classification model can distinguish atomic environments of chemically similar actinides, as long as atoms beyond the first solvation shells are considered. Our work shows that only two types of descriptors are necessary to account for all the fluctuations in heavy metal/molten salt mixtures: The first descriptor represents the electronic state of the heavy metal, while the second encompasses the local coordination environment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting measures of soil health using the microbiome and supervised machine learning

Soil health encompasses a range of biological, chemical, and physical soil properties that sustain the commercial and ecological value of agroecosystems. Monitoring soil health requires a comprehensive set of diagnostics that can be cost-prohibitive for routine analyses. The soil microbiome provides a rich source of information about soil properties, which can be assayed in a high-throughput, cost-effective way. We evaluated the accuracy of random forest (RF) and support vector machine (SVM) regression and classification models in predicting 12 measures of soil health, tillage status, and soil texture from 16S rRNA gene amplicon data with an operationally relevant sample set. We validated the efficacy of the best performing models against independent datasets and also tested best practices for processing microbiome data for use in machine learning. Soil health metrics could be predicted from microbiome data with the best models achieving a Kappa value of ~0.65, for categorical assessments, and a R2 value of ~0.8, for numerical scores. Biological health ratings were better predicted than chemical or physical ratings. Validation with independent datasets revealed that models had general predictive value for soil properties, including yield. The ecological profiles of several taxa important for model accuracy matched the observed relationships with soil health, including Pyrinomonadaceae, Nitrososphaeraceae, and Candidatus Udeaobacter. Models trained at the highest taxonomic resolution proved most accurate, with losses in accuracy resulting from rarefying, sparsity filtering, and aggregating at higher taxonomic ranks. Furthermore, our study provides the groundwork for developing scalable technology to use microbiome-based diagnostics for the assessment of soil health.

16S rRNA gene↗

Morphotype-resolved characterization of microalgal communities in a nutrient recovery process with ARTiMiS flow imaging microscopy

Microalgae-driven nutrient recovery represents a promising technology for phosphorus removal from wastewater while simultaneously generating biomass that can be valorized to offset treatment costs. As full-scale processes come online, system parameters including biomass composition must be carefully monitored to optimize performance and prevent culture crashes. In this study, flow imaging microscopy (FIM) was leveraged to characterize microalgal community composition in near real-time at a full-scale municipal wastewater treatment plant (WWTP) in Wisconsin, USA, and population and morphotype dynamics were examined to identify relationships between water chemistry, biomass composition, and system performance. Two FIM technologies, FlowCam and ARTiMiS, were evaluated as monitoring tools. ARTiMiS provided a more accurate estimate of total system biomass, and estimates derived from particle area as a proxy for biovolume yielded better approximations than particle counts. Deep learning classification models trained on annotated image libraries demonstrated equivalent performance between FlowCam and ARTiMiS, and convolutional neural network (CNN) classifiers proved significantly more accurate when compared to feature table-based dense neural network (DNN) models. Across a two-year study period, Scenedesmus spp. appeared most important for phosphorus removal, and were negatively impacted by elevated temperatures and increase in nitrite/nitrate concentrations. Chlorella and Monoraphidium also played an important role in phosphorus removal. For both Scenedesmus and Chlorella, smaller morphological types were more often associated with better system performance, whereas larger morphotypes likely associated with stress response(s) correlated with poor phosphorus recovery rates. Furthermore, these results demonstrate the potential of FIM as a critical technology for high-resolution characterization of industrial microalgal processes.

59 BASIC BIOLOGICAL SCIENCES↗

A Machine Learning Approach for the Prediction of Formability and Thermodynamic Stability of Single and Double Perovskite Oxides

Perovskite oxides continue to attract huge interest due to their fascinating and wide-ranging properties for diverse applications. The tunability of these properties may be further enhanced by increasing their compositional complexity via double perovskite-ordered configurations containing multiple cations. In this work, we focus on an exhaustive chemical space of single and double oxide perovskites and optimally explore this space to identify novel compositions that are likely to form stable compounds. Critically, we examine the relationship between formability, the practical ability to synthesize a compound, and stability, the thermodynamic preference to form the structure. Our formability and stability training data sets were enumerated from the available experimental literature and in-house density functional theory computations and contained 1505 and 3469 examples, respectively, representing state-of-the-art in the current open literature in perovskite and double perovskite compounds. Subsequently, cross-validated and highly accurate machine learning classification models are built using these training data sets and employed to screen for novel stable oxide perovskites. The study identifies (1) atomic features relevant to prediction of formability and stability in perovskite and double perovskite compounds, (2) the importance of including energy contributions due to local structural relaxations going beyond the high symmetry perovskite phase, and (3) 437,828 double perovskite compounds that are likely to be stable and 891,188 compounds that are likely to be formable. From the intersection of this large chemical space of formable and stable oxide perovskites, 414 compositions are identified as the most promising candidates for future experimental synthesis of novel oxide perovskites. The developed models may be generalized and have implications beyond perovskite discovery if applied to other families of compounds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluating Material Design Principles for Calcium-Ion Mobility in Intercalation Cathodes

Multivalent-ion batteries offer an alternative to Li-based technologies, with the potential for greater sustainability, improved safety, and higher energy density, primarily due to their rechargeable system featuring a passivating metal anode. Although a system based on the Ca 2+ /Ca couple is particularly attractive given the low electrochemical plating potential of Ca 2+ , the remaining challenge for a viable rechargeable Ca battery is to identify Ca cathodes with fast ion transport. In this work, a high-throughput computational pipeline is adapted to (1) discover novel Ca cathodes in a largely unexplored space of empty intercalation hosts and (2) develop material design rules for Ca-ion mobility. One candidate from the screening, W 2 O 3 (PO 4 ) 2 , is confirmed to have a low Nudged Elastic Band (NEB) barrier of 168 meV within a one-dimensional (1D) ion percolation topology. This candidate is subsequently synthesized and electrochemically tested, achieving reversible Ca cycling with a capacity of 25 mA h/g. To further accelerate the screening for promising Ca intercalation electrodes, machine learning (ML) Random Forest (RF) and Extreme Gradient Boosting (XGB) classification models are created with local environment descriptors based on a large, structurally and chemically diverse dataset of minimum energy pathways, spanning over 5,000 density functional theory (DFT) site energy calculations. Accuracies of 92% are achieved, material design metrics are quantified, ML force-fields are leveraged in an accelerated iteration of the screening, and a total of 27 novel Ca cathode materials are highlighted for further investigation.

25 ENERGY STORAGE↗

Using Data Science Tools to Reveal and Understand Subtle Relationships of Inhibitor Structure in Frontal Ring-Opening Metathesis Polymerization

The rate of frontal ring-opening metathesis polymerization (FROMP) using the Grubbs generation II catalyst is impacted by both the concentration and choice of monomers and inhibitors, usually organophosphorus derivatives. Herein we report a data-science-driven workflow to evaluate how these factors impact both the rate of FROMP and how long the formulation of the mixture is stable (pot life). Using this workflow, we built a classification model using a single-node decision tree to determine how a simple phosphine structural descriptor (V bur-near ) can bin long versus short pot life. Additionally, we applied a nonlinear kernel ridge regression model to predict how the inhibitor and selection/concentration of comonomers impact the FROMP rate. Furthermore, the analysis provides selection criteria for material network structures that span from highly cross-linked thermosets to non-cross-linked thermoplastics as well as degradable and nondegradable materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Leveraging 13C-Labeling to Assign Molecular Formulas to Unknown Yeast Metabolites

Mass spectrometry analyses have identified tens of thousands of unknown small molecule-associated peaks in different biological specimens. Notably, even the simplest and best studied organisms like Escherichia coli and Saccharomyces cerevisiae yield thousands of unknown peaks. A key question is how many of these reflect actual novel endogenous metabolites. To explore this, Mahieu and Patti used complete 13 C -labeling in E. coli to credential peaks as biological. This reduced the number of unknowns by more than 90%. Here, we carry out similar uniform 13 C-labeling in the Baker’s yeast S. cerevisiae and two less-studied bioenergy-relevant yeasts Rhodotorula toruloides (lipid producer) and Issatchenkia orientalis (organic acid producer). Identification of unknown metabolite peaks and their molecular formulas is facilitated through software tailored for 13 C labeling data and resulting knowledge of carbon atom count. A classification model evaluates the plausibility of each candidate formula, with peaks lacking plausible candidate formulas unlikely to reflect metabolite molecular ions. This approach prioritizes about one hundred candidate abundant unknown metabolites with logical molecular formulas. Most of these are species-specific rather than conserved across yeasts, and more are found in the nonmodel yeasts than S. cerevisiae. Thus, 13 C-labeling data on unknown metabolites highlights the potential for discovering new metabolites and pathways in nonmodel yeasts.

Carbon↗

CHQ- SocioEmo: Identifying Social and Emotional Support Needs in Consumer-Health Questions

General public, often called consumers, are increasingly seeking health information online. To be satisfactory, answers to health-related questions often have to go beyond informational needs. Automated approaches to consumer health question answering should be able to recognize the need for social and emotional support. Recently, large scale datasets have addressed the issue of medical question answering and highlighted the challenges associated with question classification from the standpoint of informational needs. However, there is a lack of annotated datasets for the non-informational needs. We introduce a new dataset for non-informational support needs, called CHQ-SocioEmo. The Dataset of Consumer Health Questions was collected from a community question answering forum and annotated with basic emotions and social support needs. This is the first publicly available resource for understanding non-informational support needs in consumer health-related questions online. We benchmark the corpus against multiple state-of-the-art classification models to demonstrate the dataset’s effectiveness.

60 APPLIED LIFE SCIENCES↗

Property space mapping of Pseudomonas aeruginosa permeability to small molecules

Two membrane cell envelopes act as selective permeability barriers in Gram-negative bacteria, protecting cells against antibiotics and other small molecules. Significant efforts are being directed toward understanding how small molecules permeate these barriers. In this study, we developed an approach to analyze the permeation of compounds into Gram-negative bacteria and applied it to Pseudomonas aeruginosa, an important human pathogen notorious for resistance to multiple antibiotics. The approach uses mass spectrometric measurements of accumulation of a library of structurally diverse compounds in four isogenic strains of P. aeruginosa with varied permeability barriers. We further developed a machine learning algorithm that generates a deterministic classification model with minimal synonymity between the descriptors. This model predicted good permeators into P. aeruginosa with an accuracy of 89% and precision above 58%. The good permeators are broadly distributed in the property space and can be mapped to six distinct regions representing diverse chemical scaffolds. We posit that this approach can be used for more detailed mapping of the property space and for rational design of compounds with high Gram-negative permeability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting transcriptional responses to cold stress across plant species

Although genome-sequence assemblies are available for a growing number of plant species, gene-expression responses to stimuli have been cataloged for only a subset of these species. Many genes show altered transcription patterns in response to abiotic stresses. However, orthologous genes in related species often exhibit different responses to a given stress. Accordingly, data on the regulation of gene expression in one species are not reliable predictors of orthologous gene responses in a related species. Here, we trained a supervised classification model to identify genes that transcriptionally respond to cold stress. A model trained with only features calculated directly from genome assemblies exhibited only modest decreases in performance relative to models trained by using genomic, chromatin, and evolution/diversity features. Models trained with data from one species successfully predicted which genes would respond to cold stress in other related species. Cross-species predictions remained accurate when training was performed in cold-sensitive species and predictions were performed in cold-tolerant species and vice versa. Models trained with data on gene expression in multiple species provided at least equivalent performance to models trained and tested in a single species and outperformed single-species models in cross-species prediction. These results suggest that classifiers trained on stress data from well-studied species may suffice for predicting gene-expression patterns in related, less-studied species with sequenced genomes.

54 ENVIRONMENTAL SCIENCES↗

Organ-delimited gene regulatory networks provide high accuracy in candidate transcription factor selection across diverse processes

Organ-specific gene expression datasets that include hundreds to thousands of experiments allow the reconstruction of organ-level gene regulatory networks (GRNs). However, creating such datasets is greatly hampered by the requirements of extensive and tedious manual curation. Here, we trained a supervised classification model that can accurately classify the organ-of-origin for a plant transcriptome. This K-Nearest Neighbor-based multiclass classifier was used to create organ-specific gene expression datasets for the leaf, root, shoot, flower, and seed in Arabidopsis thaliana . A GRN inference approach was used to determine the: i. influential transcription factors (TFs) in each organ and, ii. most influential TFs for specific biological processes in that organ. These genome-wide, organ-delimited GRNs (OD-GRNs), recalled many known regulators of organ development and processes operating in those organs. Importantly, many previously unknown TF regulators were uncovered as potential regulators of these processes. As a proof-of-concept, we focused on experimentally validating the predicted TF regulators of lipid biosynthesis in seeds, an important food and biofuel trait. Of the top 20 predicted TFs, eight are known regulators of seed oil content, e.g., WRI1, LEC1, FUS3. Importantly, we validated our prediction of MybS2, TGA4, SPL12, AGL18, and DiV2 as regulators of seed lipid biosynthesis. We elucidated the molecular mechanism of MybS2 and show that it induces purple acid phosphatase family genes and lipid synthesis genes to enhance seed lipid content. This general approach has the potential to be extended to any species with sufficiently large gene expression datasets to find unique regulators of any trait-of-interest.

09 BIOMASS FUELS↗

Radioisotope Identification with List-Mode Gamma-Ray Data

This work explores the potential of utilizing temporal data from gamma-ray detectors, known as list-mode data, to enhance radioisotope identification. Traditional identification methods, which rely on full gamma-ray spectrum analysis, often require long dwell times and struggle with spectra containing similarly spaced spectral peaks. We hypothesize that by leveraging the probabilistic nature of nuclear decay and the time-encoded information from decay sequences and interactions with surrounding materials, we can improve classification accuracy over static spectral analysis. This research examines the temporal content of list-mode data through exploratory data analysis via correlation discovery and qualitative distribution analysis. Additionally, we propose a probabilistic classification model that can utilize spectral data, temporal data, or both to determine if the incorporation of temporal information improves radioisotope identification. Our findings suggest that the temporal information present in list-mode gamma-ray data has merit and should be further investigated to develop more robust and optimal methods for utilizing this temporal information in applications requiring radioisotope identification.

List-mode data↗

Snekmer: a scalable pipeline for protein sequence fingerprinting based on amino acid recoding

Abstract Motivation The vast expansion of sequence data generated from single organisms and microbiomes has precipitated the need for faster and more sensitive methods to assess evolutionary and functional relationships between proteins. Representing proteins as sets of short peptide sequences (kmers) has been used for rapid, accurate classification of proteins into functional categories; however, this approach employs an exact-match methodology and thus may be limited in terms of sensitivity and coverage. We have previously used similarity groupings, based on the chemical properties of amino acids, to form reduced character sets and recode proteins. This amino acid recoding (AAR) approach simplifies the construction of protein representations in the form of kmer vectors, which can link sequences with distant sequence similarity and provide accurate classification of problematic protein families. Results Here, we describe Snekmer, a software tool for recoding proteins into AAR kmer vectors and performing either (i) construction of supervised classification models trained on input protein families or (ii) clustering for de novo determination of protein families. We provide examples of the operation of the tool against a set of nitrogen cycling families originally collected using both standard hidden Markov models and a larger set of proteins from Uniprot and demonstrate that our method accurately differentiates these sequences in both operation modes. Availability and implementation Snekmer is written in Python using Snakemake. Code and data used in this article, along with tutorial notebooks, are available at http://github.com/PNNL-CompBio/Snekmer under an open-source BSD-3 license. Supplementary information Supplementary data are available at Bioinformatics Advances online.

59 BASIC BIOLOGICAL SCIENCES↗