Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “ROC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A comparison of machine learning methods to classify radioactive elements using prompt-gamma-ray neutron activation data

The detection of illicit radiological materials is critical to establishing a robust second line of defence in nuclear security. Neutron-capture prompt-gamma activation analysis (PGAA) can be used to detect multiple radioactive materials across the entire Periodic Table. However, long detection times and a high rate of false positives pose a significant hindrance in the deployment of PGAA-based systems to identify the presence of illicit substances in nuclear forensics. In the present work, six different machine-learning algorithms were developed to classify radioactive elements based on the PGAA energy spectra. The model performance was evaluated using standard classification metrics and trend curves with an emphasis on comparing the effectiveness of algorithms that are best suited for classifying imbalanced datasets. We analyse the classification performance based on Precision, Recall, F1-score, Specificity, Confusion matrix, ROC-AUC curves, and Geometric Mean Score (GMS) measures. The tree-based algorithms (Decision Trees, Random Forest and AdaBoost) have consistently outperformed Support Vector Machine and K-Nearest Neighbours. Based on the results presented, AdaBoost is the preferred classifier to analyse data containing PGAA spectral information due to the high recall and minimal false negatives reported in the minority class.

97 MATHEMATICS AND COMPUTING↗

Deep neural network uncertainty quantification for LArTPC reconstruction

We evaluate uncertainty quantification (UQ) methods for deep learning applied to liquid argon time projection chamber (LArTPC) physics analysis tasks. As deep learning applications enter widespread usage among physics data analysis, neural networks with reliable estimates of prediction uncertainty and robust performance against overconfidence and out-of-distribution (OOD) samples are critical for their full deployment in analyzing experimental data. While numerous UQ methods have been tested on simple datasets, performance evaluations for more complex tasks and datasets are scarce. Here we assess the application of selected deep learning UQ methods on the task of particle classification using the PiLArNet monte carlo 3D LArTPC point cloud dataset. We observe that UQ methods not only allow for better rejection of prediction mistakes and OOD detection, but also generally achieve higher overall accuracy across different task settings. We assess the precision of uncertainty quantification using different evaluation metrics, such as distributional separation of prediction entropy across correctly and incorrectly identified samples, receiver operating characteristic curves (ROCs), and expected calibration error from observed empirical accuracy. We conclude that ensembling methods can obtain well calibrated classification probabilities and generally perform better than other existing methods in deep learning UQ literature.

47 OTHER INSTRUMENTATION↗

Deeplasmid: deep learning accurately separates plasmids from bacterial chromosomes

Plasmids are mobile genetic elements that play a key role in microbial ecology and evolution by mediating horizontal transfer of important genes, such as antimicrobial resistance genes. Many microbial genomes have been sequenced by short read sequencers and have resulted in a mix of contigs that derive from plasmids or chromosomes. New tools that accurately identify plasmids are needed to elucidate new plasmid-borne genes of high biological importance. We have developed Deeplasmid, a deep learning tool for distinguishing plasmids from bacterial chromosomes based on the DNA sequence and its encoded biological data. It requires as input only assembled sequences generated by any sequencing platform and assembly algorithm and its runtime scales linearly with the number of assembled sequences. Deeplasmid achieves an AUC–ROC of over 89%, and it was more accurate than five other plasmid classification methods. Finally, as a proof of concept, we used Deeplasmid to predict new plasmids in the fish pathogen Yersinia ruckeri ATCC 29473 that has no annotated plasmids. Deeplasmid predicted with high reliability that a long assembled contig is part of a plasmid. Using long read sequencing we indeed validated the existence of a 102 kb long plasmid, demonstrating Deeplasmid's ability to detect novel plasmids.

59 BASIC BIOLOGICAL SCIENCES↗

Toward an AI-Powered Software Pipeline for Real-Time Tracking and Analysis of Wildfire and Smoke

Real-time tracking of wildfires and smoke is crucial for effective response, minimizing damage, protecting lives, and efficiently managing resources during fire emergencies. We develop a web-based AI-powered pipeline that detects wildfires in aerial video and estimates deployment-relevant behavior metrics, including cumulative burned area, burned-area growth rate, fire spread direction, and smoke dispersion. The system combines a YOLO-based detector with YCbCr-based fire segmentation, HSV-based smoke segmentation, Farneback optical flow, and centroid-based spatiotemporal tracking. Using ground sampling distance (GSD), pixel-level fire masks are converted to physical burned-area measurements by correlating fire pixel counts with camera altitude and tilt angle. We benchmark YOLO variants and non-YOLO baselines (GoogLeNet, CNN, DBN, Autoencoder, U-Net, and AlexNet) on the IEEE FLAME dataset and a newly created aerial frame dataset, Wildfire-DB. Cross-dataset evaluation uses a strict threshold-transfer protocol: decision thresholds are selected on FLAME validation and transferred unchanged to Wildfire-DB to quantify generalization under domain shift. YOLOv6 achieves the strongest cross-dataset frame-level fire detection on Wildfire-DB (ROC-AUC 0.8200, PR-AUC 0.8044, and transferred-threshold F1 0.7596). For tracking-oriented deployment requiring oriented localization, YOLO11-OBB provides the most reliable cross-dataset behavior among OBB-capable models while remaining computationally feasible. To analyze the feasibility of UAV deployment, we further measure inference efficiency using synchronized GPU and CPU power logs on a fixed workload of 1569 frames. YOLO-family models process the video in 5.73–12.47 seconds with net energy of 1247.28–1775.39 J, substantially lower latency and energy than heavier classification and reconstruction baselines. Overall, model optimality depends on operational objectives: YOLOv6 is best for cross-dataset detection robustness, whereas YOL...

Color segmentation↗

Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems

Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1

Marfo, William [University of Texas at El Paso,Dep↗

Different CT slice thickness and contrast‐enhancement phase in radiomics models on the differential performance of lung adenocarcinoma

Abstract Background To investigate the effects of computed tomography (CT) reconstruction slice thickness and contrast‐enhancement phase on the differential diagnosis performance of radiomic signature in lung adenocarcinoma. Methods A total of 187 patients who had been pathologically confirmed with lung adenocarcinoma and nonadenocarcinoma were divided into a training cohort ( n = 149) and validation cohort ( n = 38). All the patients underwent contrast‐enhanced CT and the images were reconstructed with different slice thickness. The radiomic features were extracted from different slice thickness and scan phase. The logistic regression (LR) algorithm was used to build a machine learning model for each group. The area under the curve (AUC) obtained from the receiver operating characteristic (ROC) curve and DeLong test was used to evaluate its discriminating performance. Results Finally, 34 image features and five semantic features were selected to establish a radiomics model. Based on the three contrast‐enhanced CT phases and four reconstruction slice thickness, 12 groups of radiomics models showed good discrimination ability with the AUCs range from 0.9287 to 0.9631, sensitivity range from 0.8349 to 0.9083, specificity range from 0.825 to 0.925 in the training group. Similar results were observed in the validation group. However, there was no statistical significance between the different CT scan phase groups and different slice thickness ( p > 0.05). Conclusions The radiomic analysis of contrast‐enhanced CT can be used for the differential diagnosis of lung adenocarcinoma. Moreover, different slice thickness and contrast‐enhanced scan phase did not affect the discriminating ability in the radiomics models.

Wang, Yang↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗

Baseline [ 18 F]GTP1 tau PET imaging is associated with subsequent cognitive decline in Alzheimer’s disease

Background: The role and implementation of tau PET imaging for predicting subsequent cognitive decline in Alzheimer’s disease (AD) remains uncertain. This study was designed to evaluate the relationship between baseline [ 18 F]GTP1 tau PET and subsequent longitudinal change across multiple cognitive measures over 18 months. Methods: Our analyses incorporated data from 67 participants, including cognitively normal controls (n = 10) and β-amyloid (Aβ)-positive individuals ([ 18 F] florbetapir Aβ PET) with prodromal (n = 26), mild (n = 16), or moderate (n = 15) AD. Baseline measurements included cortical volume (MRI), tau burden ([ 18 F]GTP1 tau PET), and cognitive assessments [Mini-Mental State Examination (MMSE), Clinical Dementia Rating (CDR), 13-item version of the Alzheimer’s Disease Assessment Scale-Cognitive Subscale (ADAS-Cog13), and Repeatable Battery for the Assessment of Neuropsychological Status (RBANS)]. Cognitive assessments were repeated at 6-month intervals over an 18-month period. Associations between baseline [ 18 F]GTP1 tau PET indices and longitudinal cognitive performance were assessed via univariate (Spearman correlations) and multivariate (linear mixed effects models) approaches. The utility of potential prognostic tau PET cut points was assessed with ROC curves. Results: Univariate analyses indicated that greater baseline [ 18 F]GTP1 tau PET signal was associated with faster rates of subsequent decline on the MMSE, CDR, and ADAS-Cog13 across regions of interest (ROIs). In multivariate analyses adjusted for baseline age, cognitive performance, cortical volume, and Aβ PET SUVR, the prognostic performance of [ 18 F]GTP1 SUVR was most robust in the whole cortical gray ROI. When AD participants were dichotomized into low versus high tau subgroups based on baseline [ 18 F]GTP1 PET standardized uptake value ratios (SUVR) in the temporal (cutoff = 1.325) or whole cortical gray (cutoff = 1.245) ROIs, high tau subgroups demonstrated significantly more decline on the MMSE, CDR, and ADAS-Cog13. Conclusions: Our results suggest that [ 18 F]GTP1 tau PET represents a prognostic biomarker in AD and are consistent with data from other tau PET tracers. Tau PET imaging may have utility for identifying AD patients at risk for more rapid cognitive decline and for stratification and/or enrichment of participant selection in AD clinical trials.

60 APPLIED LIFE SCIENCES↗

Association of Serum Bile Acid and Unsaturated Fatty Acid Profiles with the Risk of Diabetic Retinopathy in Type 2 Diabetic Patients

Aim: We aimed to identify the ability of serum bile acids (BAs) and unsaturated fatty acids (UFAs) profiles to predict the development of diabetic retinopathy (DR) in type 2 diabetes mellitus (T2DM) patients. Methods: We first used univariate and multivariate analysis to compare 15 serum BA and 11 UFA levels in healthy control (HC) group (n = 82), T2DM patients with DR (n = 58) and T2DM patients without DR (n = 60). Forty T2DM patients were considered for validation. Then, the receiver operating characteristic curve (ROC) and decision curve analysis were used to assess the diagnostic value and clinical benefit of serum biomarkers alone, clinical variables alone or in combination, and the area under the curve (AUC), integrated discrimination improvement (IDI), and net reclassification improvement (NRI) were used to further assess whether the addition of biomarkers significantly improved the predictive ability of the model. Results: Orthogonal partial least squares-discriminant analysis (OPLS-DA) of serum BAs and UFAs separated the three cohorts including HC, T2DM patients with or without DR. The difference in serum BA and UFA profiles of T2DM patients with or without DR was mainly manifested in the three metabolites of taurolithocholic acid (TLCA), tauroursodeoxycholic acid (TUDCA) and arachidonic acid (AA). Together, they had an AUC of 0.785 (0.918 for validation cohort) for predicting DR in T2DM patients. After adjusting for numerous confounding factors, TLCA, TUDCA, and AA were independent predictors that differentiated T2DM with or without DR. The results of AUC, IDI, and NRI demonstrated that adding these three biomarkers to a model with clinical variables statistically increased their predictive value and were replicated in our independent validation cohort. Conclusion: These findings highlight the association of three metabolites, TLCA, TUDCA and AA, with DR and may indicate their potential value in the pathogenesis of DR.

60 APPLIED LIFE SCIENCES↗

High-Throughput Custom Monitoring for the Mu2e TDAQ System

In this project we are studying the application of programmable network hardware to provide a custom monitoring capability for the Mu2e Trigger and Data Acquisition System (TDAQ) system. The goal of the Mu2e experiment is to search for a charged-lepton flavor violating processes where a negative muon converts into an electron in the field of an aluminum nucleus. This experiment is intended to improve by four orders of magnitude the search sensitivity reached so far. We have a working prototype of a system that provides high-throughput, custom monitoring for the Mu2e TDAQ system. The custom Mu2e network packet header format is parsed as it crosses the network switch. Parsing extracts bits that convey information about error states at read-out controllers (ROCs). This information is periodically relayed to the switch controller, which in turn alerts experiment operators.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Test Beam Results of Planar Pixel Sensor for the CMS Phase 2 Inner Tracker Upgrade

Results of the test beam measurements that characterise the performance of CMS Readout Chip (CROC) sensors to be used in the High Luminosity era of the Large Hadron Collider (HL-LHC) are presented. The HL-LHC peak instantaneous luminosity of $7.5 \times 10^{34} \ \text{cm}^{-2} s^{-1}$ corresponds to an average of around 200 inelastic proton-proton collisions per beam-crossing every 25 ns. In order to efficiently reconstruct and track particles in this extreme and challenging conditions, the present CMS tracking detector will be completely replaced. The new tracking detector consists of an Inner Tracker closest to the beamline and an Outer Tracker surrounding it. These are populated with modules that comprise of readout chips and silicon sensors. The test beam measurements of these modules are vital to understand the performance of the related technologies. Using a primary 120 GeV proton beam from the Main Injector at Fermilab, data was collected at the Fermilab Test Beam Facility (FTBF) using the silicon tracker telescope that provides a precision position measurement of the track impact point with less than 5 $\mu$m uncertainty. The proton beam was incident on a 1x2 planar CROC module developed by Hamamatsu. The sensor has 100x25 $\mu m^2$ standard pixels and also a smaller number of 225x25 $\mu m^2$ longer pixels at the boundary between the two ROCs. We present characterization of these modules that includes pixel efficiency, resolution, cluster size and charge distributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Commissioning of the Mu2e Data AcQuisition System and the Vertical Slice Test of the Straw Tracker

The report discusses the commissioning of the DAQ system of the Mu2e experiment. We studied how the tracker works and especially the tracker readout. We used a generator to send pulses and we tried to understand the output and non-output of the DTC, in order to validate the event and run format and all the failure modes. In section one, we describe Charged Lepton Flavour Violation, the main purpose of Mu2e, in section two we describe Mu2e experiment. Section three presents the tracker description and readout, meanwhile section four the DAQ system and event building. In section five, we start explaining the analysis we have done on events and all the failure modes we discovered. In section six, we have done an analysis of the logger and boardreader rate, we validated the generator frequency in section seven and studied the number of hits in function of the channel number in section eight. To conclude we investigated the channel to channel time differences in section nine and in section ten we explain our Monte Carlo simulation, reproducing ROCs readout.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES↗

Establishment of a reverse transcription real-time quantitative PCR method for Getah virus detection and its application for epidemiological investigation in Shandong, China

Getah virus (GETV) is a mosquito-borne, single-stranded, positive-sense RNA virus belonging to the genus Alphavirus of the family Togaviridae . Natural infections of GETV have been identified in a variety of vertebrate species, with pathogenicity mainly in swine, horses, bovines, and foxes. The increasing spectrum of infection and the characteristic causing abortions in pregnant animals pose a serious threat to public health and the livestock economy. Therefore, there is an urgent need to establish a method that can be used for epidemiological investigation in multiple animals. In this study, a real-time reverse transcription fluorescent quantitative PCR (RT-qPCR) method combined with plaque assay was established for GETV with specific primers designed for the highly conserved region of GETV Nsp1 gene. The results showed that after optimizing the condition of RT-qPCR reaction, the minimum detection limit of the assay established in this study was 7.73 PFU/mL, and there was a good linear relationship between viral load and Cq value with a correlation coefficient ( R 2 ) of 0.998. Moreover, the method has good specificity, sensitivity, and repeatability. The established RT-qPCR is 100-fold more sensitive than the conventional RT-PCR. The best cutoff value for the method was determined to be 37.59 by receiver operating characteristic (ROC) curve analysis. The area under the curve (AUC) was 0.956. Meanwhile, we collected 2,847 serum specimens from swine, horses, bovines, sheep, and 17,080 mosquito specimens in Shandong Province in 2022. The positive detection rates by RT-qPCR were 1%, 1%, 0.2%, 0%, and 3%, respectively. In conclusion, the method was used for epidemiological investigation, which has extensive application prospects.

Cao, Xinyu↗

Automatic Search of Cataclysmic Variables Based on LightGBM in LAMOST-DR7

The search for special and rare celestial objects has always played an important role in astronomy. Cataclysmic Variables (CVs) are special and rare binary systems with accretion disks. Most CVs are in the quiescent period, and their spectra have the emission lines of Balmer series, HeI, and HeII. A few CVs in the outburst period have the absorption lines of Balmer series. Owing to the scarcity of numbers, expanding the spectral data of CVs is of positive significance for studying the formation of accretion disks and the evolution of binary star system models. At present, the research for astronomical spectra has entered the era of Big Data. The Large Sky Area Multi-Object Fiber Spectroscopy Telescope (LAMOST) has produced more than tens of millions of spectral data. the latest released LAMOST-DR7 includes 10.6 million low-resolution spectral data in 4926 sky regions, providing ideal data support for searching CV candidates. To process and analyze the massive amounts of spectral data, this study employed the Light Gradient Boosting Machine (LightGBM) algorithm, which is based on the ensemble tree model to automatically conduct the search in LAMOST-DR7. Finally, 225 CV candidates were found and four new CV candidates were verified by SIMBAD and published catalogs. This study also built the Gradient Boosting Decision Tree (GBDT), Adaptive Boosting (AdaBoost), and eXtreme Gradient Boosting (XGBoost) models and used Accuracy, Precision, Recall, the F1-score, and the ROC curve to compare the four models comprehensively. Experimental results showed that LightGBM is more efficient. The search for CVs based on LightGBM not only enriches the existing CV spectral library, but also provides a reference for the data mining of other rare celestial objects in massive spectral data.

79 ASTRONOMY AND ASTROPHYSICS↗

Edge ML for CAN bus intrusion detection in AVs

Autonomous Vehicles (AVs) are revolutionizing transportation, but their reliance on interconnected cyber-physical systems exposes them to unprecedented cybersecurity risks. This study addresses the critical challenge of detecting real-time cyber intrusions in self-driving vehicles by leveraging a dataset from the Udacity self-driving car project. We simulate four high-impact attack vectors, Denial of Service (DoS), spoofing, replay, and fuzzy attacks, by injecting noise into spatial features (e.g., bounding box coordinates) to replicate adversarial scenarios. We develop and evaluate two lightweight neural network architectures (NN-1 and NN-2) alongside a logistic regression baseline (LG-1) for intrusion detection. The models achieve exceptional performance, with NN-2 attaining an AUC score of 93.15% and 93.15% accuracy, demonstrating their suitability for edge deployment in AV environments. Through explainable AI techniques, we uncover unique forensic fingerprints of each attack type, such as spatial corruption in fuzzy attacks and temporal anomalies in replay attacks, offering actionable insights for feature engineering and proactive defense. Visual analytics, including confusion matrices, ROC curves, and feature importance plots, validate the models' robustness and interpretability. This research sets a new benchmark for AV cybersecurity, delivering a scalable, field-ready toolkit for Original Equipment Manufacturers (OEMs) and policymakers. By aligning intrusion fingerprints with SAE J3061 automotive security standards, we provide a pathway for integrating machine learning into safety-critical AV systems. Our findings underscore the urgent need for security-by-design AI, ensuring that AVs not only drive autonomously but also defend autonomously. This work bridges the gap between theoretical cybersecurity and life-preserving engineering, offering a leap toward safer, more secure autonomous transportation.

97 MATHEMATICS AND COMPUTING↗

Psychophysical Models for Signal Detection with Time Varying Uncertainty

Psychophysical models for the behavior of the human operator in detection tasks which include change in detectability, correlation between observations and deferred decisions are developed. Classical Signal Detection Theory (SDT) is discussed and its emphasis on the sensory processes is contrasted to decision strategies. The analysis of decision strategies utilizes detection tasks with time varying signal strength. The classical theory is modified to include such tasks and several optimal decision strategies are explored. Two methods of classifying strategies are suggested. The first method is similar to the analysis of ROC curves, while the second is based on the relation between the criterion level (CL) and the detectability. Experiments to verify the analysis of tasks with changes of signal strength are designed. The results show that subjects are aware of changes in detectability and tend to use strategies that involve changes in the CL's.

Gai, E.↗