Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Identification and Photometric Classification of Extragalactic Transients in the Vera C. Rubin Observatory’s Data Preview 1

The Vera C. Rubin Observatory will soon survey the southern sky, delivering a depth and sky coverage that is unprecedented in time-domain astronomy. As part of commissioning, Data Preview 1 (DP1) has been released. It comprises a Legacy Survey of Space and Time (LSST) Commissioning Camera observing campaign between 2024 November and December with multiband imaging of seven fields, covering roughly 0.4 deg 2 each, providing a first glimpse into the data products that will become available once the LSST begins. In this work, we search three fields for extragalactic transients. We identify eight new likely supernovae (SNe), and three known ones from a sample of 369,644 difference image analysis objects. Photometric classification using Superphot+ assigns subclasses with >95% confidence to only one SN Ia and one SN II in this sample. Our findings are in agreement with SN detection rate predictions of 15 ± 4 SNe from simulations using simsurvey. The SN detection rate in the data is possibly affected by the lack of suitable templates. Nevertheless, this work demonstrates the quality of the data products delivered in DP1 and indicates that the Rubin Observatory’s LSST is well placed to fulfill its discovery potential in time-domain astronomy.

Freeburn, James [University of North Carolina, Cha↗

Interpretable boosted-decision-tree analysis for the Majorana Demonstrator

The Majorana Demonstrator is a leading experiment searching for neutrinoless double-beta decay with high purity germanium detectors (HPGe). Machine learning provides a new way to maximize the amount of information provided by these detectors, but the data-driven nature makes it less interpretable compared to traditional analysis. An interpretability study reveals the machine's decision-making logic, allowing us to learn from the machine to feedback to the traditional analysis. In this work, we have presented the first machine learning analysis of the data from the Majorana Demonstrator; this is also the first interpretable machine learning analysis of any germanium detector experiment. Two gradient boosted decision tree models are trained to learn from the data, and a game-theory-based model interpretability study is conducted to understand the origin of the classification power. By learning from data, this analysis recognizes the correlations among reconstruction parameters to further enhance the background rejection performance. By learning from the machine, this analysis reveals the importance of new background categories to reciprocally benefit the standard Majorana analysis. This model is highly compatible with next-generation germanium detector experiments like LEGEND since it can be simultaneously trained on a large number of detectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Plasma image classification using cosine similarity constrained convolutional neural network

Plasma jets are widely investigated both in the laboratory and in nature. Astrophysical objects such as black holes, active galactic nuclei and young stellar objects commonly emit plasma jets in various forms. With the availability of data from plasma jet experiments resembling astrophysical plasma jets, classification of such data would potentially aid in not only investigating the underlying physics of the experiments but also the study of astrophysical jets. In this work we use deep learning to process all of the laboratory plasma images from the Caltech Spheromak Experiment spanning two decades. We found that cosine similarity can aid in feature selection, classify images through comparison of feature vector direction and be used as a loss function for the training of AlexNet for plasma image classification. We also develop a simple vector direction comparison algorithm for binary and multi-class classification. Using our algorithm we demonstrate 93 % accurate binary classification to distinguish unstable columns from stable columns and 92 % accurate five-way classification of a small, labelled data set which includes three classes corresponding to varying levels of kink instability.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Search for single production of a vector-like T quark decaying to a top quark and a neutral scalar boson in the lepton+jets final state in proton-proton collisions at $\sqrt{s}=13$ TeV

A search for single production of a vector-like T quark with charge 2e/3, decaying to a top quark and a neutral scalar boson is presented. The boson can be a standard model Higgs boson (H) or a new scalar boson (ϕ). In the first case, a branching fraction of 25% is assumed for the decay T → tH, while in the second case the T quark is assumed to decay exclusively to tϕ. The top quark is identified via its lepton+jets decay, and the neutral boson via its decay into a bottom quark-antiquark pair. Final states with Lorentz-boosted topologies are considered and machine-learning techniques are exploited for optimal event classification. The analysis uses data collected by the CMS experiment in proton-proton collisions at a center-of-mass energy of 13 TeV, corresponding to an integrated luminosity of 138 fb −1 recorded at the CERN LHC in 2016–2018. Upper limits at 95% confidence level are set on the product of cross section and branching fraction for a T quark in a narrow-width approximation. They vary between 14.7 and 0.1 fb, for T quark masses in the range 1.3–3.0 TeV and ϕ boson masses in the range 25–250 GeV. These are the first exclusion limits set on the production of a single T quark decaying into a top quark and a new neutral scalar boson. For the decay channel into a top quark and a standard model Higgs boson, the results provide the best limits on production cross sections to date, for T quark masses above 2 TeV.

hadron-hadron scattering↗

Cropland abandonment between 1986 and 2018 across the United States: spatiotemporal patterns and current land uses

Knowing where and when croplands have been abandoned or otherwise removed from cultivation is fundamental to evaluating future uses of these areas, e.g. as sites for ecological restoration, recultivation, bioenergy production, or other uses. However, large uncertainties remain about the location and time of cropland abandonment and how this process and the availability of associated lands vary spatially and temporally across the United States. Here, we present a nationwide, 30 m resolution map of croplands abandoned throughout the period of 1986–2018 for the conterminous United States (CONUS). We mapped the location and time of abandonment from annual cropland layers we created in Google Earth Engine from 30 m resolution Landsat imagery using an automated classification method and training data from the U.S. Department of Agriculture Cropland Data Layer. Our abandonment map has overall accuracies of 0.91 and 0.65 for the location and time of abandonment, respectively. From 1986 to 2018, 12.3 (±2.87) million hectares (Mha) of croplands were abandoned across CONUS, with areas of greatest change over the Ogallala Aquifer, the southern Mississippi Alluvial Plain, the Atlantic Coast, North Dakota, northern Montana, and eastern Washington state. The average annual nationwide abandoned area across our study period was 0.51 Mha per year. Annual abandonment peaked between 1997 and 1999 at a rate of 0.63 Mha year –1 , followed by a continuous decrease to 0.41 Mha year –1 in 2009–2011. Among the abandoned croplands, 53% (6.5 Mha) changed to grassland and pasture, 18.6% (2.28 Mha) to shrubland and forest, 8.4% (1.03 Mha) to wetlands, and 4.6% (0.56 Mha) to non-vegetated lands. Of the areas that we mapped as abandoned, 19.6% (2.41 Mha) were enrolled in the Conservation Reserve Program as of 2020. Our new map highlights the long-term dynamic nature of agricultural land use and its relation to various competitive pressures and land use policies in the United States.

54 ENVIRONMENTAL SCIENCES↗

Light-Duty Vehicle Trip Classification Using One-Class Novelty Detection and Exhaustive Feature Extraction

Travel mode classification within travel survey data sets, especially light-duty vehicle (LDV) trips, is foundational, though nontrivial, to emerging mobility systems, travel behavior analysis, and fuel consumption estimation. Current travel mode detection approaches require well-sampled and balanced data sets with ground truth travel mode labels. The detection approaches are rarely applied and validated on large-scale, real-world data sets, which may not satisfy the dataset requirements. This work proposes an LDV trip detection model as a supplement to current travel mode detection methods, for the case when the training set is highly (and/or completely) unbalanced, to the extent that classical machine-learning approaches become difficult or impossible to deploy. The proposed model uses a novelty detection technique - one-class support vector machines (OCSVMs) - and a novel exhaustive feature extraction (EFE) technique on continuous time series data (i.e., Global Positioning System [GPS] speed profiles) for single-mode trip trajectories. Training and validation of the model are conducted on a large-scale, real-world data set. The proposed method accurately identifies LDV trips from a broad set of multimodal trips by leveraging a wealth of preexisting in-vehicle GPS travel data. Additional sensitivity analysis sheds light on the optimal training size and feature selection, which will benefit applications limited by highly imbalanced data. The paper also discusses performance comparison with regular machine-learning approaches, the model's robustness, and the potential to extend the proposed model to multi-modal trip prediction.

33 ADVANCED PROPULSION SYSTEMS↗

Elliptically-Contoured Tensor-variate Distributions with Application to Image Learning

Statistical analysis of tensor-valued data has largely used the tensor-variate normal (TVN) distribution that may be inadequate for data arising from distributions with heavier or lighter tails. We study a general family of elliptically contoured (EC) TV distributions and derive its characterizations, moments, marginal, and conditional distributions. We describe procedures for maximum likelihood estimation from data that are (1) uncorrelated draws from an EC distribution, (2) from a scale mixture of the TVN distribution, and (3) from an underlying but unknown EC distribution, for which we extend Tyler’s robust estimator. A detailed simulation study highlights the benefits of choosing an EC distribution over the TVN for heavier-tailed data. We develop TV classification rules using discriminant analysis and EC errors and show that they better predict cats and dogs from images in the Animal Faces-HQ dataset than the TVN-based rules. A novel tensor-on-tensor regression and TV analysis of variance (TANOVA) framework under EC errors is also demonstrated to better characterize gender, age, and ethnic origin than the usual TVN-based TANOVA in the celebrated labeled faces of the wild dataset.

97 MATHEMATICS AND COMPUTING↗

Concentric Spherical GNN for 3D Representation Learning

Learning 3D representations that generalize well to arbitrarily oriented inputs is a challenge of practical importance in applications varying from computer vision to physics and chemistry. We propose a novel multi-resolution convolutional architecture for learning over concentric spherical feature maps, of which the single sphere representation is a special case. Our hierarchical architecture is based on alternatively learning to incorporate both intra-sphere and inter-sphere information. We show the applicability of our method for two different types of 3D inputs, mesh objects, which can be regularly sampled, and point clouds, which are irregularly distributed. We also propose an efficient mapping of point clouds to concentric spherical images, thereby bridging spherical convolutions on grids with general point clouds. We demonstrate the effectiveness of our approach in improving state-of-the-art performance on 3D classification tasks with rotated data.

97 MATHEMATICS AND COMPUTING↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A new aerial approach for quantifying and attributing methane emissions: implementation and validation

Methane (CH 4 ) is a powerful greenhouse gas that is produced by a diverse set of natural and anthropogenic emission sources. Biogenic methane sources generally involve anaerobic decay processes such as those occurring in wetlands, melting permafrost, or the digestion of organic matter in the guts of ruminant animals. Thermogenic CH 4 sources originate from the breakdown of organic material at high temperatures and pressure within the Earth's crust, a process which also produces more complex trace hydrocarbons such as ethane (C 2 H 6 ). Here, we present the development and deployment of an uncrewed aerial system (UAS) that employs a fast (1 Hz) and sensitive (1–0.5 ppb s -1 ) CH 4 and C 2 H 6 sensor and ultrasonic anemometer. The UAS platform is a vertical-takeoff, hexarotor drone (DJI Matrice 600 Pro, M600P) capable of vertical profiling to 120 m altitude and plume sampling across scales up to 1 km. Simultaneous measurements of CH 4 and C 2 H 6 concentrations, vector winds, and positional data allow for source classification (biogenic versus thermogenic), differentiation, and emission rates without the need for modeling or a priori assumptions about winds, vertical mixing, or other environmental conditions. The system has been used for direct quantification of methane point sources, such as orphan wells, and distributed emitters, such as landfills and wastewater treatment facilities. With detectable source rates as low as 0.04 and up to ~1500 kg h -1 , this UAS offers a direct and repeatable method of horizontal and vertical profiling of emission plumes at scales that are complementary to regional aerial surveys and localized ground-based monitoring.

54 ENVIRONMENTAL SCIENCES↗

The DESI Transients Survey: Legacy Classifications and Methodology

We present the first systematic spectroscopic observations of extragalactic transients from the Dark Energy Spectroscopic Instrument (DESI), as part of the DESI Transients Survey program. With 5,000 fibers and an ${\sim} 8$ deg$^2$ field of view, we exploit DESI as a machine for the discovery and classification of transients. We present transient classifications from archival DESI data in Data Releases 1 and 2, relying on a combination of a secondary target program and serendipitous observations. We also present observations from the first 6 months of the DESI spare fiber program dedicated to transients. The program is run in coordination with a dedicated DECam time-domain survey, serving as a pathfinder for what we will be able to achieve in conjunction with the Rubin Observatory Legacy Survey of Space and Time (LSST). We classify over 250 transients, of which the majority were previously unclassified. The sample comprises thermonuclear and core-collapse supernovae and tidal disruption events (TDEs), including a TDE observed before its discovery in imaging. We demonstrate DESI's ability to classify a population of faint transients down to $r\sim 22.5$ mag during main survey operations, with negligible impacts on DESI's main observations.

Hall, Xander J. [Carnegie Mellon U.] (ORCID:000000↗

Updates to Relevance Vector Machine: Multiclass Classification, Variable Selection, and Proof-of-Concept Application to Safeguards Fresh Fuel Verification using List-Mode Neutron Collar Data

To expand the capabilities of safeguards authorities to verify the integrity of fresh fuel assemblies, Oak Ridge National Laboratory has retrofit the existing electronics of the JCC-71 uranium neutron coincidence collar, which contains 18 3 He neutron detectors and an external 241 AmLi(α, n) neutron interrogation source arranged to surround a fresh nuclear fuel assembly. The new electronics system allows analysts to record list-mode neutron multiplicity data in addition to the singles and doubles rates that are currently measured. Based on previous proof-of-concept research, analysis of these new data will identify off-normal fuel configurations in an assembly and characterize or localize the specific partial fuel defects. The purpose of this report it to document the analysis algorithm development and then to demonstrate its capability for the safeguards verification of fresh fuel assemblies using list mode neutron collar data. To analyze the complex list-mode data collected with the upgraded uranium neutron collar, multivariate classification algorithms are being developed using a novel classification method, the relevance vector machine. This approach may be applied to multiclass problems to estimate the probability that test data belongs to one of many possible classes of data. In addition, our method identifies the most useful variables/channels for making predictions, which illuminates the basis for the model’s predictions, and this interpretability is largely unique among data analytics methods. Variable selection occurs during model training and parameter tuning and does not need any external hyperparameter tuning routines. Finally, we apply the modified relevance vector machine to a simulated dataset of list-mode neutron collar data generated with the radiation transport code MCNP. The method can correctly identify off-normal fuel configurations, categorize the data according to four fuel defect scenarios, and rank the channels in the data according to prediction utility. For nuclear safeguards applications, it is concluded that this method has the potential to increase the sensitivity and reliability to detect missing fuel rods from a standard 17 x 17 Pressurized Water Reactor (PWR) fresh fuel assembly. Within this analysis, “off-normal” (i.e., missing fuel rods) were correctly classified in 17 simulated test scenarios with one quarter (25%) of the fresh fuel rods missing using a training data set of 58 simulated measurements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry

Advancing automated classification of atmospheric aerosols from Single-Particle Mass Spectrometry (SPMS) data remains challenging due to overlapping ion signatures, compositional diversity, and limited labeled data. This study evaluates supervised and semi-supervised learning frameworks to enhance aerosol identification by jointly leveraging labeled and unlabeled spectra. Four models were compared: a supervised Support Vector Machine (SVM), a self-training SVM, a stacked autoencoder classifier, and a stacked autoencoder trained using a temporal-ensembling Mean Teacher approach. All models achieved high and stable accuracies (90.0 %–91.1 %), surpassing previous results on the same dataset (87 %) and matching the performance of state-of-the-art deep learning methods. Despite small global metric differences (≤ 1 %), semi-supervised variants yielded up to 5 %–10 % improvements for compositionally rare particle types – such as soot (0.77 % of spectra, F1-score: 0.93–0.97) and hazelnut pollen (0.98 % of spectra, F1-score: 0.97–1.00) – equating to roughly ∼ 187 additional correctly classified spectra. These gains are scientifically significant, as such rare particles exert disproportionate influence on radiative absorption and ice nucleation processes; their improved detection reduces modeled uncertainties in aerosol absorption optical depth and mixed-phase cloud ice nucleation rates. The models' residual misclassifications (≈ 9 %) largely arise from true spectral overlap among chemically adjacent species (e.g., Na- vs. K-feldspar, coated vs. uncoated feldspars), reflecting physical compositional continuity rather than algorithmic error. Collectively, these findings demonstrate that leveraging unlabeled data to learn robust spectral representations and refine classification enhances both fidelity and interpretability, bridging data-driven analysis with aerosol–climate process understanding.

54 ENVIRONMENTAL SCIENCES↗

Quasi-Distributed Fiber Sensor-Based Approach for Pipeline Health Monitoring: Generating and Analyzing Physics-Based Simulation Datasets for Classification

This study presents a framework for detecting mechanical damage in pipelines, focusing on generating simulated data and sampling to emulate distributed acoustic sensing (DAS) system responses. The workflow transforms simulated ultrasonic guided wave (UGW) responses into DAS or quasi-DAS system responses to create a physically robust dataset for pipeline event classification, including welds, clips, and corrosion defects. This investigation examines the effects of sensing systems and noise on classification performance, emphasizing the importance of selecting the appropriate sensing system for a specific application. The framework shows the robustness of different sensor number deployments to experimentally relevant noise levels, demonstrating its applicability in real-world scenarios where noise is present. Overall, this study contributes to the development of a more reliable and effective method for detecting mechanical damage to pipelines by emphasizing the generation and utilization of simulated DAS system responses for pipeline classification efforts. The results on the effects of sensing systems and noise on classification performance further enhance the robustness and reliability of the framework.

36 MATERIALS SCIENCE↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

Identification of integrated proteomics and transcriptomics signature of alcohol-associated liver disease using machine learning

Distinguishing between alcohol-associated hepatitis (AH) and alcohol-associated cirrhosis (AC) remains a diagnostic challenge. In this study, we used machine learning with transcriptomics and proteomics data from liver tissue and peripheral mononuclear blood cells (PBMCs) to classify patients with alcohol-associated liver disease. The conditions in the study were AH, AC, and healthy controls. We processed 98 PBMC RNAseq samples, 55 PBMC proteomic samples, 48 liver RNAseq samples, and 53 liver proteomic samples. First, we built separate classification and feature selection pipelines for transcriptomics and proteomics data. The liver tissue models were validated in independent liver tissue datasets. Next, we built integrated gene and protein expression models that allowed us to identify combined gene-protein biomarker panels. For liver tissue, we attained 90% nested-cross validation accuracy in our dataset and 82% accuracy in the independent validation dataset using transcriptomic data. We attained 100% nested-cross validation accuracy in our dataset and 61% accuracy in the independent validation dataset using proteomic data. For PBMCs, we attained 83% and 89% accuracy with transcriptomic and proteomic data, respectively. The integration of the two data types resulted in improved classification accuracy for PBMCs, but not liver tissue. We also identified the following gene-protein matches within the gene-protein biomarker panels: CLEC4M-CLC4M, GSTA1-GSTA2 for liver tissue and SELENBP1-SBP1 for PBMCs. In this study, machine learning models had high classification accuracy for both transcriptomics and proteomics data, across liver tissue and PBMCs. The integration of transcriptomics and proteomics into a multi-omics model yielded improvement in classification accuracy for the PBMC data. The set of integrated gene-protein biomarkers for PBMCs show promise toward developing a liquid biopsy for alcohol-associated liver disease.

60 APPLIED LIFE SCIENCES↗

Automated solar collector installation design including version management

Embodiments may include systems and methods to create and edit a representation of a worksite, to create various data objects, to classify such objects as various types of predefined “features” with attendant properties and layout constraints. As part of or in addition to classification, an embodiment may include systems and methods to create, associate, and edit intrinsic and extrinsic properties to these objects. A design engine may apply of design rules to the features described above to generate one or more solar collectors installation design alternatives, including generation of on-screen and/or paper representations of the physical layout or arrangement of the one or more design alternatives. Some embodiments may provide viewing, creating, and manipulating of multiple versions of a solar collector layout design for a particular installation worksite. The use of versions may allow analysis of alternative layouts, alternative feature classifications, and cost and performance data corresponding to alternative design choices. Version summary information providing a representative comparison between versions across a number of dimensions may be provided.

14 SOLAR ENERGY↗