Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AuC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Effects of Substance Use and Antisocial Personality on Neuroimaging-Based Machine Learning Prediction of Schizophrenia

Abstract Background and hypothesis Neuroimaging-based machine learning (ML) algorithms have the potential to aid the clinical diagnosis of schizophrenia. However, literature on the effect of prevalent comorbidities such as substance use disorder (SUD) and antisocial personality (ASPD) on these models’ performance has remained unexplored. We investigated whether the presence of SUD or ASPD affects the performance of neuroimaging-based ML models trained to discern patients with schizophrenia (SCH) from controls. Study design We trained an ML model on structural MRI data from public datasets to distinguish between SCH and controls (SCH = 347, controls = 341). We then investigated the model’s performance in two independent samples of individuals undergoing forensic psychiatric examination: sample 1 was used for sensitivity analysis to discern ASPD (N = 52) from SCH (N = 66), and sample 2 was used for specificity analysis to discern ASPD (N = 26) from controls (N = 25). Both samples included individuals with SUD. Study results In sample 1, 94.4% of SCH with comorbid ASPD and SUD were classified as SCH, followed by patients with SCH + SUD (78.8% classified as SCH) and patients with SCH (60.0% classified as SCH). The model failed to discern SCH without comorbidities from ASPD + SUD (AUC = 0.562, 95%CI = 0.400–0.723). In sample 2, the model’s specificity to predict controls was 84.0%. In both samples, about half of the ASPD + SUD were misclassified as SCH. Data-driven functional characterization revealed associations between the classification as SCH and cognition-related brain regions. Conclusion Altogether, ASPD and SUD appear to have effects on ML prediction performance, which potentially results from converging cognition-related brain abnormalities between SCH, ASPD, and SUD.

99 GENERAL AND MISCELLANEOUS↗

The mechanism driving a solid–solid phase transition in a biomacromolecular crystal

A solid-solid phase transition (SSPT) occurs between distinguishable crystalline forms. Because of its importance in application and theory in material science and condensed matter physics, SSPT has been studied most extensively in metallic alloys, inorganic salt or small organic molecular crystals, but much less so in biomacromolecular crystals. In general, the mechanism of SSPT at atomic and molecular levels is not well understood. Here, we describe the ordered molecular rearrangements in biomacromolecular crystals of the adenine riboswitch (riboA) aptamer using real-time serial crystallography and solution atomic force microscopy (AFM). The large, ligand-induced conformational changes drive the initial phase transition from the apo unit cell (AUC) to the trans unit cell 1 (TUC1). During this transition, coaxial stacking of P1 duplexes becomes the dominant packing interface, whereas P2-P2 interactions are almost completely disrupted, resulting in “floating” layers of molecules. The coupling points in TUC1 and their local conformational flexibility allow the molecules to reorganize to achieve the more densely packed and energetically favorable bound unit cell (BUC). Our study thus reveals the interplay between the conformational changes and the crystal phases—the underlying mechanism that drives the phase transition. Using polarized video microscopy (PVM) to monitor the SSPT in small crystals at high ligand concentration, we have identified the time window during which the major conformational changes take place, and simulated the in crystallo kinetics. Together, these results provide the spatiotemporal information necessary for informing time-resolved crystallography (TRX) experiments. Moreover, this study illustrates a practical approach to characterize SSPT in transparent crystals.

36 MATERIALS SCIENCE↗

Errant Beam Detection Using the AMD Versal ACAP and Vitis AI

The prevalence of ML and AI-powered solutions along with the slowing of Moore's Law has given rise to novel hardware platforms aimed at accelerating ML and AI. While programming these hardware platforms can be difficult, particularly for non-hardware experts, hardware vendors provide high-level tooling in an effort to address this difficulty. The Versal ACAP is an SoC designed by AMD that combines CPU cores, FPGA fabric, and a tiled, vector architecture called an AI engine all on the same socket. In an effort to more easily program this heterogeneous system, AMD has provided the Vitis AI development stack. In this work, we leverage Vitis AI to program a Versal ACAP to perform errant beam detection in the Spallation Neutron Source at Oak Ridge National Laboratory. Our initial work shows that after quantization and compilation of the model for the Versal ACAP, the classification accuracy, as measured by the AUC metric, is over 95% accurate while achieving this accuracy in 46 microseconds on average.

Cabrera, Anthony↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

A comparison of histopathology imaging comprehension algorithms based on multiple instance learning

Whole slide imaging (WSI), also called digital virtual microscopy, is a new imaging modality. It allows for the application of AI and machine learning methods to cancer pathology to help establish a means for the automatic diagnosis of cancer cases. However, designing machine-learning models for WSI is computationally challenging due to its required ultra-high resolution. The current state-of-the-art models use multiple instance learning (MIL). MIL is a weakly-supervised learning method in which the model uses an array of inferences from many smaller instances to make a final classification about the entire set. In the context of WSI, researchers divide the ultra-high-resolution image into many patches. The model then classifies the slide based on an array of inferences from the patches. Among several ways of making the final classification, attention-based mechanisms have resulted in superb accuracy scores. The Transformer, one attention-based algorithm, has reported substantial improvements for WSI comprehension tasks. In this project, we studied and compared several WSI comprehension algorithms. We used the following three datasets: CAMELYON16+17, TCGALung, and TCGA-Kidney. We found that attention-based MIL algorithms performed better than standard MIL algorithms for classifying WSI images, achieving a higher mean accuracy and AUC. However, none of the attention-based algorithms performed significantly better than the others, reporting accuracy scores that varied widely. Presumably, it is due to the limited availability of training samples in the data corpus. Since it is not easy to increase the samples from human subjects, some machine learning techniques like transfer learning could help mitigate this issue.

Saunders, Adam↗

SourceFinder: a Machine-Learning-Based Tool for Identification of Chromosomal, Plasmid, and Bacteriophage Sequences from Assemblies

High-throughput genome sequencing technologies enable the investigation of complex genetic interactions, including the horizontal gene transfer of plasmids and bacteriophages. However, identifying these elements from assembled reads remains challenging due to genome sequence plasticity and the difficulty in assembling complete sequences. In this study, we developed a classifier, using random forest, to identify whether sequences originated from bacterial chromosomes, plasmids, or bacteriophages. The classifier was trained on a diverse collection of 23,211 chromosomal, plasmid, and bacteriophage sequences from hundreds of bacterial species. In order to adapt the classifier to incomplete sequences, each complete sequence was subsampled into 5,000 nucleotide fragments and further subdivided into k-mers. This three-class classifier succeeded in identifying chromosomes, plasmids, and bacteriophages using k-mer distributions of complete and partial genome sequences, including simulated metagenomic scaffolds with minimum performance of 0.939 area under the receiver operating characteristic curve (AUC). This classifier, implemented as SourceFinder, has been made available as an online web service to help the community with predicting the chromosomal, plasmid, and bacteriophage sources of assembled bacterial sequence data (https://cge.food.dtu.dk/services/SourceFinder/).

59 BASIC BIOLOGICAL SCIENCES↗

Characterizing the Spread of COVID-19 from Human Mobility Patterns and SocioDemographic Indicators

Mobility is an indicator of human movement through space and time. With the increasing availability of geolocated data (from GPS, accelerometers, etc.), it is now possible to examine individual as well as group human mobility patterns. Human mobility is influenced by both intrinsic (i.e. personal motivations) and extrinsic (i.e., events like natural hazards or a pandemic like the COVID-19) factors. However, the intricate relationships between human mobility patterns and sociodemographic characteristics in the context of a pandemic are yet to be fully explored. Our goal is to overcome this gap by using human mobility data at the census block group level from mobile phones and combining those with social vulnerability indicators to examine the overall spread of COVID-19 at local spatial scales. We used 585,878 weekly visits to 37,871 points of interests (POIs) from Safegraph to quantify mobility indices and social distancing metrics in 2,820 census block groups in the city of Los Angeles (LA) - before and during lockdown as well as during the phase1 and phase 2 reopening. Finally, using supervised machine learning algorithms, we classified the census block groups in LA into High, Medium and Low categories that represented the vulnerability of these block groups based on the cumulative number of occurrences of COVID-19 cases till July 24, 2020. Our results indicate that the tree-based classifiers performed well in comparison to the Support Vector Machines and Multinomial Logit models. Gradient Boosting had the highest classification accuracy of 97.4% COVID-19 with an AUC score of 0.987. The block groups with high COVID-19 cases also had a high concentration of socially vulnerable populations, high human mobility index and a low social distancing index.

Roy, Avipsa↗

Scaling Resolution of Gigapixel Whole Slide Images Using Spatial Decomposition on Convolutional Neural Networks

Gigapixel images are prevalent in scientific domains ranging from remote sensing, and satellite imagery to microscopy, etc. However, training a deep learning model at the natural resolution of those images has been a challenge in terms of both, overcoming the resource limit (e.g. HBM memory constraints), as well as scaling up to a large number of GPUs. In this paper, we trained Residual neural Networks (ResNet) on 22,528 x 22,528-pixel size images using a distributed spatial decomposition method on 2,304 GPUs on the Summit Supercomputer. We applied our method on a Whole Slide Imaging (WSI) dataset from The Cancer Genome Atlas (TCGA) database. WSI images can be in the size of 100,000 x 100,000 pixels or even larger, and in this work we studied the effect of image resolution on a classification task, while achieving state-of-the-art AUC scores. Moreover, our approach doesn't need pixel-level labels, since we're avoiding patching from the WSI images completely, while adding the capability of training arbitrary large-size images. This is achieved through a distributed spatial decomposition method, by leveraging the non-block fat-tree interconnect network of the Summit architecture, which enabled GPU-to-GPU direct communication. Finally, detailed performance analysis results are shown, as well as a comparison with a data-parallel approach when possible.

Tsaris, Aristeidis (aris)↗

The Prediction Model of Risk Factors for COVID-19 Developing into Severe Illness Based on 1046 Patients with COVID-19

This study analyzed the risk factors for patients with COVID-19 developing severe illnesses and explored the value of applying the logistic model combined with ROC curve analysis to predict the risk of severe illnesses at COVID-19 patients’ admissions. The clinical data of 1046 COVID-19 patients admitted to a designated hospital in a certain city from July to September 2020 were retrospectively analyzed, the clinical characteristics of the patients were collected, and a multivariate unconditional logistic regression analysis was used to determine the risk factors for severe illnesses in COVID-19 patients during hospitalization. Based on the analysis results, a prediction model for severe conditions and the ROC curve were constructed, and the predictive value of the model was assessed. Logistic regression analysis showed that age (OR = 3.257, 95% CI 10.466–18.584), complications with chronic obstructive pulmonary disease (OR = 7.337, 95% CI 0.227–87.021), cough (OR = 5517, 95% CI 0.258–65.024), and venous thrombosis (OR = 7322, 95% CI 0.278–95.020) were risk factors for COVID-19 patients developing severe conditions during hospitalization. When complications were not taken into consideration, COVID-19 patients’ ages, number of diseases, and underlying diseases were risk factors influencing the development of severe illnesses. The ROC curve analysis results showed that the AUC that predicted the severity of COVID-19 patients at admission was 0.943, the optimal threshold was −3.24, and the specificity was 0.824, while the sensitivity was 0.827. The changes in the condition of severe COVID-19 patients are related to many factors such as age, clinical symptoms, and underlying diseases. This study has a certain value in predicting COVID-19 patients that develop from mild to severe conditions, and this prediction model is a useful tool in the quick prediction of the changes in patients’ conditions and providing early intervention for those with risk factors.

Lian, Zhichuang↗

Codon2Vec v1.0

Background: Codon2Vec is an embedding neural network that predicts 'high' or 'low' gene expression directly from the protein-coding sequences. Embedding neural networks are commonly used for natural language processing (NLP) applications. Analogous to how an English sentence is a string of words, a gene can be thought of as a string of codons. Similar to how NLP neural networks model English sentences as a non-random sequence of words, we considered a coding sequence as a non-random non-overlapping array of codons (k-mers of length = 3). Value Proposition: - Codon2Vec achieved a high median AUC-ROC score of 83.8% when trained and applied to transcriptomic data from 300 fungal species - Unlike Codo2Vec, conventional methods predicting for expression based on codon usage rely on a priori knowledge of optimal codons or a set of reference genes. - Unlike Codon2vec, these methods do not account for the effect of codon order on gene expression. - Codon2Vec neural network bypasses the need for artisanal feature selection step that is necessary for traditional machine learning models.

Wint, Rhondene↗

Time-Based CAN IDS Paper Results Code

Modern vehicles are complex cyber-physical systems made of hundreds of electronic control units (ECUs) that communicate over controller area networks (CANs). This inherited complexity has expanded the CAN attack surface which is vulnerable to message injection attacks. These injections change the overall timing characteristics of messages on the bus, and thus, to detect these malicious messages, time-based intrusion detection systems (IDSs) have been proposed. However, time-based IDSs are usually trained and tested on low-fidelity datasets with unrealistic, labeled attacks. This makes difficult the task of evaluating, comparing, and validating IDSs. Here we detail and benchmark four time-based IDSs against the newly published ROAD dataset, the first open CAN IDS dataset with real (non-simulated) stealthy attacks with physically verified effects. We found that methods that perform hypothesis testing by explicitly estimating message timing distributions have lower performance than methods that seek anomalies in a distribution related statistic. In particular, these “distribution-agnostic” based methods outperform “distribution-based” methods by at least 55% in area under the precision-recall curve (AUC-PR). Our results expand the body of knowledge of CAN time-based IDSs by providing details of these methods and reporting their results when tested on datasets with real advanced attacks. Finally, we develop an after-market plug-in detector using lightweight hardware, which can be used to deploy the best performing IDS method on nearly any vehicle.

Moriano, Pablo [Oak Ridge National Lab. (ORNL), Oa↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗

Identification of carbohydrate gene clusters obtained from in vitro fermentations as predictive biomarkers of prebiotic responses

Prebiotic fibers are non-digestible substrates that modulate the gut microbiome by promoting expansion of microbes having the genetic and physiological potential to utilize those molecules. Although several prebiotic substrates have been consistently shown to provide health benefits in human clinical trials, responder and non-responder phenotypes are often reported. These observations had led to interest in identifying, a priori, prebiotic responders and non-responders as a basis for personalized nutrition. In this study, we conducted in vitro fecal enrichments and applied shotgun metagenomics and machine learning tools to identify microbial gene signatures from adult subjects that could be used to predict prebiotic responders and non-responders. Using short chain fatty acids as a targeted response, we identified genetic features, consisting of carbohydrate active enzymes, transcription factors and sugar transporters, from metagenomic sequencing of in vitro fermentations for three prebiotic substrates: xylooligosacharides, fructooligosacharides, and inulin. A machine learning approach was then used to select substrate-specific gene signatures as predictive features. These features were found to be predictive for XOS responders with respect to SCFA production in an in vivo trial. Our results confirm the bifidogenic effect of commonly used prebiotic substrates along with inter-individual microbial responses towards these substrates. We successfully trained classifiers for the prediction of prebiotic responders towards XOS and inulin with robust accuracy (≥ AUC 0.9) and demonstrated its utility in a human feeding trial. Overall, the findings from this study highlight the practical implementation of pre-intervention targeted profiling of individual microbiomes to stratify responders and non-responders.

59 BASIC BIOLOGICAL SCIENCES↗

Time-Based CAN Intrusion Detection Benchmark

Modern vehicles are complex cyber-physical systems made of hundreds of electronic control units (ECUs) that communicate over controller area networks (CANs). This inherited complexity has expanded the CAN attack surface by the injection of malicious messages that vary their time-based characteristics. To detect these malicious messages, time-based intrusion detection systems (IDS) have been proposed. However, time-based IDS are usually trained and tested on low-fidelity datasets with unrealistic labeled attacks. This makes difficult the task of evaluating, comparing, and validating IDS. Here we detail and benchmark four time-based IDS in a dataset with real and advanced attacks. We found that methods with strong assumptions regarding the distribution of inter-arrival times have lower performance than distribution agnostic based methods. In particular, distribution agnostic based methods outperform distribution based methods at least on $55\%$ in area under the precision-recall (AUC-PR) curve. Our results expand the body of knowledge of CAN time-based IDS by providing details of these methods and reporting their results when tested on datasets with real and advanced attacks. We describe limitations, open challenges, and how lessons learnt from this research can inform the design of deployable time-based IDS in modern vehicles.

Blevins, Deborah↗

Exploring unsupervised top tagging using Bayesian inference

Recognizing hadronically decaying top-quark jets in a sample of jets, or even its total fraction in the sample, is an important step in many LHC searches for Standard Model and Beyond Standard Model physics as well. Although there exists outstanding top-tagger algorithms, their construction and their expected performance rely on Montecarlo simulations, which may induce potential biases. For these reasons we develop two simple unsupervised top-tagger algorithms based on performing Bayesian inference on a mixture model. In one of them we use as the observed variable a new geometrically-based observable \tilde{A}_{3} A ̃ 3 , and in the other we consider the more traditional \tau_{3}/\tau_{2} τ 3 / τ 2 N N -subjettiness ratio, which yields a better performance. As expected, we find that the unsupervised tagger performance is below existing supervised taggers, reaching expected Area Under Curve AUC \sim 0.80-0.81 ∼ 0.80 − 0.81 and accuracies of about 69% - − 75% in a full range of sample purity. However, these performances are more robust to possible biases in the Montecarlo that their supervised counterparts. Our findings are a step towards exploring and considering simpler and unbiased taggers.

Alvarez, Ezequiel↗

A perfused multi-well bioreactor platform to assess tumor organoid response to a chemotherapeutic gradient

There is an urgent need to develop new therapies for colorectal cancer that has metastasized to the liver and, more fundamentally, to develop improved preclinical platforms of colorectal cancer liver metastases (CRCLM) to screen therapies for efficacy. To this end, we developed a multi-well perfusable bioreactor capable of monitoring CRCLM patient-derived organoid response to a chemotherapeutic gradient. CRCLM patient-derived organoids were cultured in the multi-well bioreactor for 7 days and the subsequently established gradient in 5-fluorouracil (5-FU) concentration resulted in a lower IC 50 in the region near the perfusion channel versus the region far from the channel. We compared behaviour of organoids in this platform to two commonly used PDO culture models: organoids in media and organoids in a static (no perfusion) hydrogel. The bioreactor IC 50 values were significantly higher than IC 50 values for organoids cultured in media whereas only the IC 50 for organoids far from the channel were significantly different than organoids cultured in the static hydrogel condition. Using finite element simulations, we showed that the total dose delivered, calculated using area under the curve (AUC) was similar between platforms, however normalized viability was lower for the organoid in media condition than in the static gel and bioreactor. Our results highlight the utility of our multi-well bioreactor for studying organoid response to chemical gradients and demonstrate that comparing drug response across these different platforms is nontrivial.

59 BASIC BIOLOGICAL SCIENCES↗

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES↗

Establishment of a reverse transcription real-time quantitative PCR method for Getah virus detection and its application for epidemiological investigation in Shandong, China

Getah virus (GETV) is a mosquito-borne, single-stranded, positive-sense RNA virus belonging to the genus Alphavirus of the family Togaviridae . Natural infections of GETV have been identified in a variety of vertebrate species, with pathogenicity mainly in swine, horses, bovines, and foxes. The increasing spectrum of infection and the characteristic causing abortions in pregnant animals pose a serious threat to public health and the livestock economy. Therefore, there is an urgent need to establish a method that can be used for epidemiological investigation in multiple animals. In this study, a real-time reverse transcription fluorescent quantitative PCR (RT-qPCR) method combined with plaque assay was established for GETV with specific primers designed for the highly conserved region of GETV Nsp1 gene. The results showed that after optimizing the condition of RT-qPCR reaction, the minimum detection limit of the assay established in this study was 7.73 PFU/mL, and there was a good linear relationship between viral load and Cq value with a correlation coefficient ( R 2 ) of 0.998. Moreover, the method has good specificity, sensitivity, and repeatability. The established RT-qPCR is 100-fold more sensitive than the conventional RT-PCR. The best cutoff value for the method was determined to be 37.59 by receiver operating characteristic (ROC) curve analysis. The area under the curve (AUC) was 0.956. Meanwhile, we collected 2,847 serum specimens from swine, horses, bovines, sheep, and 17,080 mosquito specimens in Shandong Province in 2022. The positive detection rates by RT-qPCR were 1%, 1%, 0.2%, 0%, and 3%, respectively. In conclusion, the method was used for epidemiological investigation, which has extensive application prospects.

Cao, Xinyu↗