Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Average Accuracy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Automating galaxy morphology classification using k -nearest neighbours and non-parametric statistics

ABSTRACT Morphology is a fundamental property of any galaxy population. It is a major indicator of the physical processes that drive galaxy evolution and in turn the evolution of the entire Universe. Historically, galaxy images were visually classified by trained experts. However, in the era of big data, more efficient techniques are required. In this work, we present a k-nearest neighbours based approach that utilizes non-parametric morphological quantities to classify galaxy morphology in Sloan Digital Sky Survey images. Most previous studies used only a handful of morphological parameters to identify galaxy types. In contrast, we explore 1023 morphological spaces (defined by up to 10 non-parametric statistics) to find the best combination of morphological parameters. Additionally, while most previous studies broadly classified galaxies into early types and late types or ellipticals, spirals, and irregular galaxies, we classify galaxies into 11 morphological types with an average accuracy of ${\sim} 80\!-\!90 \, {{\rm per\, cent}}$ per T-type. Our method is simple, easy to implement, and is robust to varying sizes and compositions of the training and test samples. Preliminary results on the performance of our technique on deeper images from the Hyper Suprime-Cam Subaru Strategic Survey reveal that an extension of our method to modern surveys with better imaging capabilities might be possible.

Mukundan, Kavya↗

Effects-Based Monitoring of Geomagnetically-Induced Current Using a Convolutional Neural Network

Geomagnetically-induced current (GIC) due to space weather can flow in the power grid causing undesirable effects such as transformer overheating, misoperation of protection devices, and potential blackouts. It is therefore important to monitor GIC in the power grid to improve online situational awareness and decision-making of system operators during a geomagnetic disturbance. To avoid the costly installation of GIC monitors at transformers’ neutrals, it is desirable to find correlations between GIC and already-monitored parameters. Hence, this work proposed the use of a convolutional neural network (CNN) to compute GIC amplitudes from learned patterns in the time-series data of transformer even harmonic currents. Using an electromagnetic transient program, GIC injection simulations were performed for a modeled Dominion Energy Virginia (DEV) substation with two 504 MVA, 500/230 kV transformers. Data collected from these offline simulations were used to train the CNN to provide online GIC monitoring. Testing the CNN performance involved using real GIC measurements from published literature and from a physical GIC monitor in the DEV area. Finally, the results showed that the proposed method was able to provide GIC readings with a root mean squared error of 1.56 A/phase (equivalent to an average accuracy of 94%) for these real GIC waveforms.

42 ENGINEERING↗

DP-TwoLevel: two-stage gradient subspace learning for differentially private federated learning

Federated learning (FL) enables collaborative model training across distributed data sources without sharing raw data, but faces fundamental challenges in communication efficiency and privacy. Differentially private (DP) training mitigates information leakage but introduces noise that degrades model performance, especially in high-dimensional settings. We propose DP-TwoLevel, a hierarchical gradient projection method that improves utility under fixed DP constraints by exploiting low-dimensional structure in model updates. Our approach learns a two-level PCA-based representation of gradients and applies DP noise in a reduced-dimensional subspace, thereby lowering the effective noise magnitude while preserving dominant signal components. We evaluate the method across three datasets (MNIST, Fashion-MNIST, CIFAR-10) and three privacy regimes (ϵ∈0.5, 1.0, 2.0). Across nine experimental settings, DP-TwoLevel consistently outperforms DP-FedAvg, achieving an average accuracy improvement of 9.44%, with larger gains observed in lower ϵ(higher-noise) regimes (up to +22.31%). We further analyze scalability across models ranging from 100K to 1.49M parameters and identify a variance-based success criterion: performance remains strong when the projection preserves more than 75% of gradient variance, degrades in a marginal regime (65–75%), and fails below this threshold. Our results demonstrate that structure-aware dimensionality reduction can significantly improve the privacy–utility tradeoff in FL without modifying formal privacy guarantees. We also provide empirical evidence of scaling limitations for global projections and motivate per-layer extensions for larger models.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data

Zero-shot and prompt-based models have excelled at visual reasoning tasks by leveraging large-scale natural image corpora, but they often fail on sparse and domain-specific scientific image data. We introduce Zenesis, a no-code interactive computer vision platform designed to reduce data readiness bottlenecks in scientific imaging workflows. Zenesis integrates lightweight multimodal adaptation for zero-shot inference on raw scientific data, human-in-the-loop refinement, and heuristic-based temporal enhancement. We validate our approach on Focused Ion Beam Scanning Electron Microscopy (FIB-SEM) datasets of catalyst-loaded membranes. Zenesis outperforms baselines, achieving an average accuracy of 0.947, Intersection over Union (IoU) of 0.858, and Dice score of 0.923 on amorphous catalyst samples; and 0.987 accuracy, 0.857 IoU, and 0.923 Dice on crystalline samples. These results represent a significant performance gain over conventional methods such as Otsu thresholding and standalone models like the Segment Anything Model (SAM). Zenesis enables effective image segmentation in domains where annotated datasets are limited, offering a scalable solution for scientific discovery.

Mukherjee, Shubhabrata↗

HLA-Clus: HLA class I clustering based on 3D structure

In a previous paper, we classified populated HLA class I alleles into supertypes and subtypes based on the similarity of 3D landscape of peptide binding grooves, using newly defined structure distance metric and hierarchical clustering approach. Compared to other approaches, our method achieves higher correlation with peptide binding specificity, intra-cluster similarity (cohesion), and robustness. Here we introduce HLA-Clus, a Python package for clustering HLA Class I alleles using the method we developed recently and describe additional features including a new nearest neighbor clustering method that facilitates clustering based on user-defined criteria. The HLA-Clus pipeline includes three stages: First, HLA Class I structural models are coarse grained and transformed into clouds of labeled points. Second, similarities between alleles are determined using a newly defined structure distance metric that accounts for spatial and physicochemical similarities. Finally, alleles are clustered via hierarchical or nearest-neighbor approaches. We also interfaced HLA-Clus with the peptide:HLA affinity predictor MHCnuggets. By using the nearest neighbor clustering method to select optimal allele-specific deep learning models in MHCnuggets, the average accuracy of peptide binding prediction of rare alleles was improved. The HLA-Clus package offers a solution for characterizing the peptide binding specificities of a large number of HLA alleles. This method can be applied in HLA functional studies, such as the development of peptide affinity predictors, disease association studies, and HLA matching for grafting. HLA-Clus is freely available at our GitHub repository (https://github.com/yshen25/HLA-Clus).

59 BASIC BIOLOGICAL SCIENCES↗

Classification of bacterial plasmid and chromosome derived sequences using machine learning

Plasmids are important genetic elements that facilitate horizonal gene transfer between bacteria and contribute to the spread of virulence and antimicrobial resistance. Most bacterial genome sequences in the public archives exist in draft form with many contigs, making it difficult to determine if a contig is of chromosomal or plasmid origin. Using a training set of contigs comprising 10,584 chromosomes and 10,654 plasmids from the PATRIC database, we evaluated several machine learning models including random forest, logistic regression, XGBoost, and a neural network for their ability to classify chromosomal and plasmid sequences using nucleotide k-mers as features. Based on the methods tested, a neural network model that used nucleotide 6-mers as features that was trained on randomly selected chromosomal and plasmid subsequences 5kb in length achieved the best performance, outperforming existing out-of-the-box methods, with an average accuracy of 89.38% ± 2.16% over a 10-fold cross validation. The model accuracy can be improved to 92.08% by using a voting strategy when classifying holdout sequences. In both plasmids and chromosomes, subsequences encoding functions involved in horizontal gene transfer—including hypothetical proteins, transporters, phage, mobile elements, and CRISPR elements—were most likely to be misclassified by the model. This study provides a straightforward approach for identifying plasmid-encoding sequences in short read assemblies without the need for sequence alignment-based tools.

59 BASIC BIOLOGICAL SCIENCES↗

Grid Event Signature Library Analytics Report: Signature Matching Tool Development Efforts

This report describes the purpose and features of the Signature Matching Tool (SMT), employed in the Department of Energy (DOE) Grid Event Signature Library (GESL). The SMT supports a user of GESL to identify snippets of electric signatures, usually from sensor devices measuring electric characteristics such as phase voltages and currents, frequency, etc., suspected to represent certain events in the power grid but are not known to the user. The SMT uses a classification method to identify an event of the unknown signature, using the repository of known and labeled signatures in the GESL. The classifier applies a local binary classifier per node (LCN) approach to the unique event tag taxonomy used in the GESL, where training phases are separated based on the Primary labels in the taxonomy, sensor type, and voltage level. Results show that this method helps with computing time during training, in comparison to a flat, multinomial classifier, and produces acceptable average accuracy of 83% across all Primary labels. The report concludes with planned future work including integration to the web interface and API.

97 MATHEMATICS AND COMPUTING↗

Discriminative analysis of schizophrenia patients using graph convolutional networks: A combined multimodal MRI and connectomics analysis

Introduction Recent studies in human brain connectomics with multimodal magnetic resonance imaging (MRI) data have widely reported abnormalities in brain structure, function and connectivity associated with schizophrenia (SZ). However, most previous discriminative studies of SZ patients were based on MRI features of brain regions, ignoring the complex relationships within brain networks. Methods We applied a graph convolutional network (GCN) to discriminating SZ patients using the features of brain region and connectivity derived from a combined multimodal MRI and connectomics analysis. Structural magnetic resonance imaging (sMRI) and resting-state functional magnetic resonance imaging (rs-fMRI) data were acquired from 140 SZ patients and 205 normal controls. Eighteen types of brain graphs were constructed for each subject using 3 types of node features, 3 types of edge features, and 2 brain atlases. We investigated the performance of 18 brain graphs and used the TopK pooling layers to highlight salient brain regions (nodes in the graph). Results The GCN model, which used functional connectivity as edge features and multimodal features (sMRI + fMRI) of brain regions as node features, obtained the highest average accuracy of 95.8%, and outperformed other existing classification studies in SZ patients. In the explainability analysis, we reported that the top 10 salient brain regions, predominantly distributed in the prefrontal and occipital cortices, were mainly involved in the systems of emotion and visual processing. Discussion Our findings demonstrated that GCN with a combined multimodal MRI and connectomics analysis can effectively improve the classification of SZ at an individual level, indicating a promising direction for the diagnosis of SZ patients. The code is available at https://github.com/CXY-scut/GCN-SZ.git .

Chen, Xiaoyi↗

Deep-Learning-Based Segmentation of Keyhole in In-Situ X-ray Imaging of Laser Powder Bed Fusion

In laser powder bed fusion processes, keyholes are the gaseous cavities formed where laser interacts with metal, and their morphologies play an important role in defect formation and the final product quality. The in-situ X-ray imaging technique can monitor the keyhole dynamics from the side and capture keyhole shapes in the X-ray image stream. Keyhole shapes in X-ray images are then often labeled by humans for analysis, which increasingly involves attempting to correlate keyhole shapes with defects using machine learning. However, such labeling is tedious, time-consuming, error-prone, and cannot be scaled to large data sets. To use keyhole shapes more readily as the input to machine learning methods, an automatic tool to identify keyhole regions is desirable. In this paper, a deep-learning-based computer vision tool that can automatically segment keyhole shapes out of X-ray images is presented. The pipeline contains a filtering method and an implementation of the BASNet deep learning model to semantically segment the keyhole morphologies out of X-ray images. The presented tool shows promising average accuracy of 91.24% for keyhole area, and 92.81% for boundary shape, for a range of test dataset conditions in Al6061 (and one AliSi10Mg) alloys, with 300 training images/labels and 100 testing images for each trial. Prospective users may apply the presently trained tool or a retrained version following the approach used here to automatically label keyhole shapes in large image sets.

36 MATERIALS SCIENCE↗

Using Downwelling Far- and Thermal-Infrared Hyperspectral Radiance for Cloud Phase Classification in the Antarctic

The cloud phase is one of the most important parameters of clouds. In this paper, we propose a method for cloud phase classification that synergistically utilizes the far- and thermal-infrared bands based on the Atmospheric Emitted Radiance Interferometer (AERI) at the Atmospheric Radiation Measurement West Antarctic Radiation Experiment (AWARE) observatory in 2016. The possible features in the far- and thermal-infrared bands are analyzed based on the differences in the simulated cloud brightness temperature (BT) spectra with different cloud phases. Using the support vector machine (SVM) algorithm, four features are determined to identify the cloud phase, which include the BT at 900 cm -1 , the slope of the fitted function of BT in the 900–1000 cm -1 interval, the BT difference (BTD) between 512 cm -1 and 726 cm -1 , and the BTD between 550 cm -1 and 726 cm -1 . Here, the performance of the proposed method is evaluated with Shupe’s and Turner’s method. The monthly average accuracy of the proposed method, the method without the two far-infrared features, and Turner’s method are about 76%, 36%, and 49%, respectively, which infer the good performance of the proposed method and also indicate that the far-infrared band features can effectively enhance cloud phase classification. It is notable that, compared to Shupe’s method, the accuracy for the proposed method is only 61% during the Antarctic summer, which results from the definitions of cloud phase and radiative effect. In addition, the accuracy is only 44% for Turner’s method in seasons with a low frequency of mixed clouds due to the significant effect of water vapor.

54 ENVIRONMENTAL SCIENCES↗

Decoding the EEG patterns induced by sequential finger movement for brain-computer interfaces

Objective In recent years, motor imagery-based brain–computer interfaces (MI-BCIs) have developed rapidly due to their great potential in neurological rehabilitation. However, the controllable instruction set limits its application in daily life. To extend the instruction set, we proposed a novel movement-intention encoding paradigm based on sequential finger movement. Approach Ten subjects participated in the offline experiment. During the experiment, they were required to press a key sequentially [i.e., Left→Left (LL), Right→Right (RR), Left→Right (LR), and Right→Left (RL)] using the left or right index finger at about 1 s intervals under an auditory prompt of 1 Hz. The movement-related cortical potential (MRCP) and event-related desynchronization (ERD) features were used to investigate the electroencephalography (EEG) variation induced by the sequential finger movement tasks. Twelve subjects participated in an online experiment to verify the feasibility of the proposed paradigm. Main results As a result, both the MRCP and ERD features showed the specific temporal–spatial EEG patterns of different sequential finger movement tasks. For the offline experiment, the average classification accuracy of the four tasks was 71.69%, with the highest accuracy of 79.26%. For the online experiment, the average accuracies were 83.33% and 82.71% for LL-versus-RR and LR-versus-RL, respectively. Significance This paper demonstrated the feasibility of the proposed sequential finger movement paradigm through offline and online experiments. This study would be helpful for optimizing the encoding method of motor-related EEG information and providing a promising approach to extending the instruction set of the movement intention-based BCIs.

Liu, Chang↗

Orbit-averaging and deposition accuracy for runaway electron beams in hybrid kinetic-MHD simulations of the runaway plateau

We develop a new procedure that combines the kinetic orbit runaway electron code (KORC) and the NIMROD extended-magnetohydrodynamic code to simulate runaway electrons (REs) in the post-disruption plateau. KORC integrates guiding-center orbits, with a barycentric-based binary search strategy providing initial guesses for the Newton–Raphson logical-to-physical coordinate inversion, ensuring reliable particle-to-mesh mapping in NIMROD, whose fields remain static for the present study. Samples are drawn in accord with experimental parallel current profiles of RE beams during the plateau phase. Deposition in NIMROD is verified through comparison with a Python-based finite-element code that ensures periodicity in the poloidal direction and continuity at the magnetic axis. Accurate representation of near-axis fields requires finer mesh resolution to prevent under- and overshoots in current density from orbit inaccuracies. Yet, at a fixed particle count, increasing mesh resolution amplifies statistical noise in the deposited fields. An orbit-averaging method accumulates partial current deposits over multiple kinetic steps and reduces the statistical noise with little added computational cost. By coupling kinetic routines from KORC directly into the NIMROD codebase, these developments lay essential groundwork for future self-consistent KORC–NIMROD coupling.

Algorithms and data structure↗

Detection and imaging of chemicals and hidden explosives using terahertz time-domain spectroscopy and deep learning

Detecting concealed chemicals and explosives remains a critical challenge in global security. Terahertz time-domain spectroscopy (THz-TDS) offers a promising non-invasive and stand-off detection technique owing to its ability to penetrate optically opaque materials without causing ionization damage. While many chemicals exhibit distinct spectral features in the terahertz range, conventional terahertz-based detection methods often struggle in real-world environments, where variations in sample geometry, thickness, and packaging can lead to inconsistent spectral responses. In this study, we present a chemical imaging system that integrates THz-TDS with deep learning to enable accurate pixel-level identification and classification of different explosives. Operating in reflection mode and enhanced with plasmonic nanoantenna arrays, our THz-TDS system achieves a peak dynamic range of 96 dB and a detection bandwidth of 4.5 THz, supporting practical, stand-off operation. By analyzing individual time-domain pulses with deep neural networks, the system exhibits strong resilience to environmental variations and sample inconsistencies. Blind testing across eight chemicals—including pharmaceutical excipients and explosive compounds—resulted in an average classification accuracy of 99.42% at the pixel level. Notably, the system maintained an average accuracy of 88.83% when detecting explosives concealed under opaque paper coverings, demonstrating its robust generalization capability. These results highlight the potential of combining advanced terahertz spectroscopy with neural networks for highly sensitive and specific chemical and explosive detection in diverse and operationally relevant scenarios.

Imaging and sensing↗

Machine Learning Models for Mapping Groundwater Pollution Risk: Advancing Water Security and Sustainable Development Goals in Georgia, USA

The widespread use of pesticides, such as atrazine and malathion, in agricultural systems raises significant concerns regarding the contamination of groundwater, which serves as a critical resource for drinking water. This study applies machine learning techniques to predict the concentrations of atrazine and malathion in groundwater across Georgia, USA, using 2019 data. A Random Forest classifier was employed to integrate various environmental and demographic factors, including pesticide application rates, precipitation, lithology, and population density, to predict pesticide contamination in groundwater. The models demonstrated high training accuracies of 100% and moderate average testing accuracy of 55% for atrazine and 60% for malathion across five iterations. The low test accuracy of the model, ranging from 50% to 75%, is likely due to overfitting, which can be attributed to the small dataset size and the complex nature of pesticide-contamination patterns, making it challenging for the model to generalize to unseen data. Feature importance analysis revealed that average pesticide usage emerged as the most influential factor for atrazine, while aquifer lithology and precipitation played crucial roles in both models. These results provide valuable insights into the dynamics of pesticide contamination, highlighting areas at greater risk of contamination. The findings underscore the importance of integrating environmental, geological, and agricultural variables for more effective groundwater management and sustainable agricultural practices, contributing to the protection of water resources and public health.

54 ENVIRONMENTAL SCIENCES↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

Dual-Image Color Normalization to Enable High-Performance Concentrating Solar Optical Metrology

Concentrating Solar Power (CSP) requires precision mirrors, and these in turn require metrology systems to measure their optical slope. In this project we studied a color-based approach to the correspondence problem, which is the association of points on an optical target with their corresponding points seen in a reflection. This is a core problem in deflectometry-based metrology, and a color solution would enable important new capabilities. We modeled color as a vector in the [R,G,B] space measured by a digital camera, and explored a dual-image approach to compensate for inevitable changes in illumination color. Through a series of experiments including color target design and dual-image setups both indoors and outdoors, we collected reference/measurement image pairs for a variety of configurations and light conditions. We then analyzed the resulting image pairs by selecting example [R,G,B] pixels in the reference image, and seeking matching [R,G,B] pixels in the measurement image. Modulating a tolerance threshold enabled us to assess both match reliability and match ambiguity, and for some configurations, orthorectification enabled us to assess match accuracy. Using direct-direct imaging, we demonstrated color correspondence achieving average match accuracy values of 0.004 h, where h is the height of the color pattern. We found that wide-area two-dimensional and linear one-dimensional color targets outperformed hybrid linear/lateral gradient targets in the cases studied. Introducing a mirror degraded performance under our current techniques, and we did not have time to evaluate whether matches could be reliably achieved despite varying light conditions. Nonetheless, our results thus far are promising.

14 SOLAR ENERGY↗

Rapid, antibiotic incubation-free determination of tuberculosis drug resistance using machine learning and Raman spectroscopy

Tuberculosis (TB) is the world’s deadliest infectious disease, with over 1.5 million deaths and 10 million new cases reported anually. The causative organism Mycobacterium tuberculosis (Mtb) can take nearly 40 d to culture, a required step to determine the pathogen’s antibiotic susceptibility. Both rapid identification and rapid antibiotic susceptibility testing of Mtb are essential for effective patient treatment and combating antimicrobial resistance. Here, we demonstrate a rapid, culture-free, and antibiotic incubation-free drug susceptibility test for TB using Raman spectroscopy and machine learning. We collect few-to-single-cell Raman spectra from over 25,000 cells of the Mtb complex strain Bacillus Calmette-Guérin (BCG) resistant to one of the four mainstay anti-TB drugs, isoniazid, rifampicin, moxifloxacin, and amikacin, as well as a pan-susceptible wildtype strain. By training a neural network on this data, we classify the antibiotic resistance profile of each strain, both on dried samples and on patient sputum samples. On dried samples, we achieve >98% resistant versus susceptible classification accuracy across all five BCG strains. In patient sputum samples, we achieve ~79% average classification accuracy. We develop a feature recognition algorithm in order to verify that our machine learning model is using biologically relevant spectral features to assess the resistance profiles of our mycobacterial strains. Finally, we demonstrate how this approach can be deployed in resource-limited settings by developing a low-cost, portable Raman microscope that costs <$5,000. We show how this instrument and our machine learning model enable combined microscopy and spectroscopy for accurate few-to-single-cell drug susceptibility testing of BCG.

60 APPLIED LIFE SCIENCES↗

Pedestal origin and extrapolation of high-density small edge-localised-modes peak parallel energy fluence in ITER and SPARC

Experimental analysis and simulations with the BOUT++ code show that small edge-localised modes (ELMs) in reactor-relevant high-density regimes originate in a region close to the separatrix and only marginally perturb the pedestal structure. The measured divertor peak parallel energy fluence (ε ∥,peak ) for a database of small ELM scenarios in DIII-D and ASDEX Upgrade can be reproduced, within 40 % accuracy on average, if an ad hoc modification of the Eich peak parallel ELM energy fluence model is applied to account for the small ELM pedestal birth location. This allows for first-order extrapolation of small-ELM divertor ε ∥,peak to ITER and SPARC, resulting in values that satisfy the nominal melting threshold of tungsten monoblocks of 12 MJ m −2 . The findings reported in this study, both via modelling and direct measurements, constitute a step forward in assessing small ELMs in high edge-collisionality scenarios as a viable plasma regime for the operation of next-generation fusion machines.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗