Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning↗

WBS 2.1.5.401 - Model Validation and Site Characterization for Early Deployment MHK Sites and Establishment of Wave Classification Scheme

The "Resource Characterization" project delivers the data and tools needed to engineer robust marine renewable energy devices and projects. The project measures resource details at commercially promising sites, runs high resolution models of promising sites and regions, and develops classification schemes that streamline device engineering, project development, and increase investor confidence.

ENGINEERING,TIDAL AND WAVE POWER↗

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS↗

Calorimetric classification of track-like signatures in liquid argon TPCs using MicroBooNE data

The MicroBooNE liquid argon time projection chamber located at Fermilab is a neutrino experiment dedicated to the study of short-baseline oscillations, the measurements of neutrino cross sections in liquid argon, and to the research and development of this novel detector technology. Accurate and precise measurements of calorimetry are essential to the event reconstruction and are achieved by leveraging the TPC to measure deposited energy per unit length along the particle trajectory, with mm resolution. We describe the non-uniform calorimetric reconstruction performance in the detector, showing dependence on the angle of the particle trajectory. Such non-uniform reconstruction directly affects the performance of the particle identification algorithms which infer particle type from calorimetric measurements. This work presents a new particle identification method which accounts for and effectively addresses such non-uniformity. The newly developed method shows improved performance compared to previous algorithms, illustrated by a 93.7% proton selection efficiency and a 10% muon mis-identification rate, with a fairly loose selection of tracks performed on beam data. The performance is further demonstrated by identifying exclusive final states in ν μ CC interactions. While developed using MicroBooNE data and simulation, this method is easily applicable to future LArTPC experiments, such as SBND, ICARUS, and DUNE.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Avian Activity Classification Using Recurrent Networks to Fuse Videos with Metadata on Imbalanced Datasets

Activity classification plays a crucial role in various real-life scenarios involving both humans and animals. There is an increasing need for precise activity classification focused on avian-solar interactions, as the usage of solar energy facilities, such as photovoltaic array power stations, has been observed to impact bird species richness, behavior, and activity. However, there has been no work to develop an automated system to monitor and classify these avian-solar interactions. All current methods rely on human observers, which is time and human resources costly and subject to errors related to searcher efficiency. With the recent success of Deep Learning models in activity classification problems, this paper develops a recurrent neural network-based model to automatically classify six avian activities around solar energy facilities. Our proposed model integrates critical feature engineering metadata with video frame data, enabling improved learning and more accurate activity classification. Furthermore, we address the challenge of data imbalance during training and demonstrate the efficacy of our model in detecting and classifying different activities within video tracks. Additionally, we analyze the saliency/backpropagation map of the trained proposed model and validate its decision-making rationale.

Avian activity classification; bidirectional LSTM;↗

Exploring New Ways to Classify Industries for Energy Analysis and Modeling

As the US moves closer to embracing a net zero greenhouse gas emissions position, combustion processes outside the power sector are becoming urgent concerns. Industry is an important end user of energy and relies on fossil fuels used directly for process heating and as feedstocks for a diverse range of applications. Fuel and energy use by industry is heterogeneous, meaning that even a single product group can vary broadly in its production routes and associated energy usage. In the US, the North American Industry Classification System (NAICS) serves as the basis for data collection and reporting. In turn, data based on NAICS is the foundation of most US energy modeling. Thus, the effectiveness of NAICS at representing energy use is a limiting condition for plans to improve energy efficiency and alternatives to fossil fuels in industry. Facility-level data to build more detail into heterogeneous sectors is scarce. This work explores alternative classification schemes for industry based on energy use characteristics, and provides a validation of an approach to make facility-level energy use estimates based on publicly available data from the greenhouse gas reporting program. First, several approaches to industrial taxonomies and their usefulness for industrial energy modeling are summarized. Data from Industrial Assessment Centers is analyzed using unsupervised machine learning techniques to detect clusters. Cladistics, an approach from biology, is adapted to energy and process characteristics of industries. A cladogram is presented for evolutionary directions in the iron and steel sector. Cladograms are a promising tool for constructing scenarios and summarizing directions of sectoral innovation. Finally, validation is performed for facility-level energy estimates from the US EPA Greenhouse Gas Reporting Program. This validation assists in making this data source available for use in energy modeling. Together, this work explores alternative approaches for categorizing industries in a way that aids understanding energy use, and presenting pathways for the future.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Data Driven Fault Detection of Premixer Centerbody Degradation in a Swirl Combustor

This paper introduces a data-driven framework for combustor-focused, performance-based condition monitoring of gas turbines. Commercial condition monitoring systems typically generate huge amounts of data that make efficient onboard monitoring challenging. This paper focuses on quantifying combustor component degradation, using premixer centerbody degradation in a swirl stabilized combustor as a case study. The input for these analyses is acoustic pressure measurements acquired at various locations on the combustor. The diagnosis methodology is based on a classification framework and consists of 3 steps: 1) Data curation, 2) Feature Engineering, and 3) Diagnosis. Data curation ensures good quality of the data that is passed through the algorithm. Feature engineering deals with the extraction of the most informative features, from the most informative sensors, that can accurately capture the introduced fault. To perform diagnosis, the classification model is trained using experimentally acquired data and is then tested on a separate data set. The framework was able to achieve high classification accuracy (>99%) for training size as low as 30% of the total recorded observations. The low number of features required to achieve this accuracy suggests high potential for integration into existing onboard condition monitoring systems.

data driven methods, fault detection, swirl flames↗

DeepMerge: Classifying high-redshift merging galaxies with deep neural networks

In this work, we investigate and demonstrate the use of convolutional neural networks (CNNs) for the task of distinguishing between merging and non-merging galaxies in simulated images, and for the first time at high redshifts (i.e. $z=2$). We extract images of merging and non-merging galaxies from the Illustris-1 cosmological simulation and apply observational and experimental noise that mimics that from the Hubble Space Telescope; the data without noise form a "pristine" data set and that with noise form a "noisy" data set. The test set classification accuracy of the CNN is $79\%$ for pristine and $76\%$ for noisy. The CNN outperforms a Random Forest classifier, which was shown to be superior to conventional one- or two-dimensional statistical methods (Concentration, Asymmetry, the Gini, $M_{20}$ statistics etc.), which are commonly used when classifying merging galaxies. We also investigate the selection effects of the classifier with respect to merger state and star formation rate, finding no bias. Finally, we extract Grad-CAMs (Gradient-weighted Class Activation Mapping) from the results to further assess and interrogate the fidelity of the classification model.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluation of data collection bias of third molar stages of mineralisation for age estimation in the living

Abstract Age assessment of the living is a fundamental procedure in the process of human identification, in order to guarantee fair treatment of individuals, which has ethical, civil, legal, and medical repercussions. The careful selection of the appropriate methods requires evaluation of several parameters: accuracy, precision of the method, as well as its reproducibility. The approach proposed by Mincer et al. adapted from Demirjian et al. exploring third molar mineralisation, is one of the most frequently considered for age estimation of the living. Thus, this work aims to assess potential bias in the data collection when applying the classification stages for dental mineralisation adapted by Mincer et al. A total of 102 orthopantomographs, of clinical origin, belonging to individuals aged between 12 and 25 years ($ \bar{\textit x} $ = 20.12 years, SD = 3.49 years; 65 females, 37 males, all of Portuguese nationality) were included and a retrospective analysis performed by five observers with different levels of experience (high, average, and basic). The performance and agreement between five observers were evaluated using Weighted Cohen’s Kappa and the Intraclass Correlation Coefficient. To access the influence of impaction on third molar classification, variables were tested using ordinal logistic regression Generalised Linear Model. It was observed that there were variations in the number of teeth identified among the observers, but the agreement levels ranged from moderate to substantial (0.4–0.8). Upon closer examination of the results, it was observed that although there were discernible differences between highly experienced observers and those with less experience, the gap was not as significant as initially hypothesised, and a greater disparity between the classifications of the upper (0.24–0.49) and lower third molars (>0.55) was observed. When bone superimposition is present, the classification process is not significantly influenced; however, variation in teeth angulation affects the assessment. The results suggest that with an efficient preparation, the level of experience as a factor can be overcome. Mincer and colleague's classification system can be replicated with ease and consistency, even though the classification of upper and lower third molars presents distinct challenges.

de Oliveira Santos, Inês (ORCID:0000000267324347)↗

Supervised learning with word embeddings derived from PubMed captures latent knowledge about protein kinases and cancer

Abstract Inhibiting protein kinases (PKs) that cause cancers has been an important topic in cancer therapy for years. So far, almost 8% of >530 PKs have been targeted by FDA-approved medications, and around 150 protein kinase inhibitors (PKIs) have been tested in clinical trials. We present an approach based on natural language processing and machine learning to investigate the relations between PKs and cancers, predicting PKs whose inhibition would be efficacious to treat a certain cancer. Our approach represents PKs and cancers as semantically meaningful 100-dimensional vectors based on word and concept neighborhoods in PubMed abstracts. We use information about phase I-IV trials in ClinicalTrials.gov to construct a training set for random forest classification. Our results with historical data show that associations between PKs and specific cancers can be predicted years in advance with good accuracy. Our tool can be used to predict the relevance of inhibiting PKs for specific cancers and to support the design of well-focused clinical trials to discover novel PKIs for cancer therapy.

59 BASIC BIOLOGICAL SCIENCES↗

Convolutional Variational Autoencoder-based Unsupervised Learning for Power Systems Faults

Classification of power system event data is a growing need, particularly where non-protective relaying-based sensors are used to monitor grid performance. Given the high burden of obtaining event data with appropriate labeling, an unsupervised approach is highly valuable. This approach enables using event data without labeling, which is far easier to obtain. This paper presents an unsupervised learning method to classify and label transients observed in the distribution grid. A Convolutional Variational Autoencoder (CVAE) was developed for this purpose. We demonstrate the efficacy of our approach using the transient data generated from the simulations. The simulation data is used to train the CVAE that identifies different faults as different clusters in the latent space. The clusters are then used as the foundation model to categorize the real-world data.

Alam, Maksudul↗

Genomad v1.0

Genomad aims to identify mobile genetic elements (namely, viruses and plasmids) from DNA sequence data. It uses a combination of marker gene identification and machine learning models to find likely virus/plasmid candidates in environmental DNA sequencing data. It has an improved classification performance over similar tools.

Camargo, Antonio↗

Three-Dimensional Convective–Stratiform Echo-Type Classification and Convectivity Retrieval from Radar Reflectivity

The Echo Classification from COnvectivity (ECCO) algorithm identifies convective and stratiform types of radar echo in three dimensions. It is based on the calculation of reflectivity texture—a combination of the intensity and the heterogeneity of the radar echoes on each horizontal plane in a 3D Cartesian volume. Reflectivity texture is translated into convectivity, which is designed to be a quantitative measure of the convective nature of each 3D radar grid point. It ranges from 0 (100% stratiform) to 1 (100% convective). By thresholding convectivity, a more traditional qualitative categorization is obtained, which classifies radar echoes as convective, mixed, or stratiform. In contrast to previous algorithms, these echo-type classifications are provided on the full 3D grid of the reflectivity field. The vertically resolved classifications, in combination with temperature data, allow for subclassifications into shallow, mid-, deep, and elevated convective features, and low, mid-, and high stratiform regions—again in three dimensions. The algorithm was validated using datasets collected over the U.S. Great Plains during the PECAN field campaign. An analysis of lightning counts shows ~90% of lightning occurring in regions classified as convective by ECCO. A statistical comparison of ECCO echo types with the well-established GPM radar precipitation-type categories show 84% (88%) of GPM stratiform (convective) echo being classified as stratiform (convective) or mixed by ECCO. ECCO was applied to radar grids for the continental United States, the United Arab Emirates, Australia, and Europe, illustrating its robustness and adaptability to different radar grid characteristics and climatic regions.

Radars/Radar observations↗

Water Observations of Flow/No-Flow for the East-Taylor Watershed, Colorado (June-July 2025 and 2026)

This dataset provides multi-year, ground-truth visual observations of surface water flow/no-flow conditions within the East-Taylor Watershed, Colorado, collected during June and July of 2025 and 2026. In June and July 2025, on-the-ground visual observations of flow/no-flow were collected as part of the Watershed Function Scientific Focus Area (SFA) and Rocky Mountain Biological Laboratory (RMBL) Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign (further details are provided within the CHESS Project Description). We obtained 377 water observations of flow/no-flow within the East-Taylor Watershed, Colorado. These ground-truth observations were collected to validate classification maps from remote sensing data and model results within the East-Taylor Watershed. In 2025, flow/no-flow measurements were collected using a field-based app for the CHESS Campaign (Zerion iForm). Within the field app, a water observation form was created to collect coordinates and metadata about the observation. Information collected for the water observation points included information about visually-assessed streamflow presence/absence (standard question obtained from Colorado State University’s StreamTracker project), flow estimate, stream or ponded area width, canopy cover, manganese films, iron seeps, and beaver activity. For 2025 water observations, this dataset contains: (1) a data file with the water observations and coordinates (2025_Water_Observations.csv); (2) a Keyhole Markup Language Zipped (KMZ) with the water observation locations and metadata (2025_Water_Observations_Locations.kmz); (3) photos (.jpg and .jpeg) of the water observation points, organized by location, contained within 2025_Water_Observations_FieldPhotographs.zip file; and (4) water observation protocols and figures (2025_Water_Observation_Protocols.pdf). In June and July 2026, on-the-ground visual observations of flow/no-flow were collected as part of the Watershed Function SFA project. We obtained 365 water observations of flow/no-flow within the East-Taylor Watershed, Colorado. The 2026 observations focused on collecting repeat measurements at the 2025 flow/no-flow observation locations conducted as part of the CHESS campaign. These ground-truth observations were collected to understand differences in flow/no-flow in 2026, given the unprecedented 2026 drought in Colorado. In 2026, flow/no-flow measurements were collected using ArcGIS (Geographic Information System) Survey123. Within the field app, a water observation form was created to collect coordinates and metadata about the observation. Information collected for the water observation points included repeat information from the 2025 water observation effort, including visually-assessed streamflow presence/absence (standard question obtained from Colorado State University’s StreamTracker project), flow estimate, stream or ponded area width, canopy cover, manganese films, iron seeps, beaver activity, and a new metadata component of estimated stream depth (for select locations). For 2026 water observations, this dataset contains: (1) a data file with the water observations and coordinates (2026_Water_Observations.csv); (2) a Keyhole Markup Language Zipped (KMZ) with the water observation locations and metadata (2026_Water_Observations_Locations.kmz); (3) photos (.jpg) of the water observation points, organized by location, contained within 2026_Water_Observations_FieldPhotographs.zip file; and (4) water observation protocols and figures (2026_Water_Observation_Protocols.pdf). For 2025 and 2026 water observations, this dataset contains: (1) a location metadata file (locations.csv); (6) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and (7) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. 2026-09-02: This dataset was updated to include 2026 water observation measurements. The 2025 observation files were also updated to ensure a consistent file naming convention across water observation years.

2018 NEON and 2025 CHESS Campaigns↗

Seismic Spatial Gradients and Machine Learning-Based Classifiers for Explosion Monitoring (LDRD 218327)

This final report summarizes the work completed under the Laboratory Directed Research and Development (LDRD) project “Seismic Spatial Gradients as a Machine Learning-Based Classifier for Explosion Monitoring.” The overarching goal of the project was to explore the efficacy of using machine learning-based classification algorithms where the input data are the spatial gradient of the seismic wavefield collected at a single point on the Earth’s surface. The methods that I describe here are in direct contrast to conventional methods of seismic discrimination which typically rely on a spatially extended network of instruments and physics-based wavefield attributes such as, for example, the ratio between $\textit{P}$ and $\textit{S}$ waves. Rather, we use the spatial gradient of the seismic wavefield observed at a single point on the Earth’s surface and data processing approaches inspired by the machine learning community. We tested two algorithms, a neural network and a modified version of principal component analysis termed Spectrally Filtered Principal Component Analysis (SFPCA). To test these algorithms, we first conducted a series of numerical tests using synthetic data and then conducted a small-scale controlled field experiment. The tests using synthetic data showed that both algorithms had high success rates on gradiometric data, even when simulated noise was added to the signal. Furthermore, we found that using seismic spatial gradients increased the performance of our discrimination algorithms when compared to using just the traditional translational motion seismic data. The tests with field data also showed a high degree of discriminative success.

58 GEOSCIENCES↗

Evaluation of Multi-Fidelity Soil Moisture Products Across the Continental United States

We have aggregated the most recent soil moisture datasets from a diverse range of sources, encompassing the Continental United States (CONUS). These sources encompass gridded data from remote sensing products, reanalysis products, machine learning-based projects, and land surface modeling products. Additionally, we have obtained and processed in-situ soil moisture observations from the International Soil Moisture Network. The collected datasets exhibit variations in both temporal and spatial resolutions. Among the 20 datasets, six are available at a spatial resolution of 0.25 degrees, while three are at a coarser spatial resolution of 25 km. To minimize spatial interpolation, we conducted data uncertainty evaluations at the 0.25-degree spatial resolution. For our data evaluations, we maintained a monthly temporal resolution, which effectively captures soil moisture seasonality and interannual variability. Our data processing strategy preserves the raw data and interpolated data at their original temporal resolutions. Datasets with higher temporal resolutions, including daily, three-hourly, and hourly datasets, are set aside for subsequent analyses. These analyses will delve into topics such as soil moisture changes and recovery during extreme weather events. Furthermore, we have processed auxiliary data to enhance our evaluation, leveraging tools such as Google Earth Engine. This includes incorporating topography data, land use land cover data, Köppen-Geiger climate classification, and more to provide a comprehensive assessment from multiple sources.

Li, Lingcheng↗

Feature identification or classification using task-specific metadata

Innovations in the identification or classification of features in a data set are described, such as a data set representing measurements taken by a scientific instrument. For example, a task-specific processing component, such as a video encoder, is used to generate task-specific metadata. When the data set includes video frames, metadata can include information regarding motion of image elements between frames, or other differences between frames. A feature of the data set, such as an event, can be identified or classified based on the metadata. For example, an event can be identified when metadata for one or more elements of the data set exceed one or more threshold values. When the feature is identified or classified, an output, such as a display or notification, can be generated. Although the metadata may be useable to generate a task-specific output, such as compressed video data, the identifying or classifying is not used solely in production of, or the creation of an association with, the task-specific output.

Teuton, Jeremy R.↗

TopTemp: Parsing Precipitate Structure from Temper Topology

Technological advances are in part enabled by the development of novel manufacturing processes that give rise to new materials or material property improvements. Development and evaluation of new manufacturing methodologies is labor-, time-, and resource-intensive expensive due to complex, poorly defined relationships between advanced manufacturing process parameters and the resulting microstructures. In this work, we present a topological representation of temper (heat-treatment) dependent material micro-structure, as captured by scanning electron microscopy, called TopTemp. We show that this topological representation is able to support temper classification of microstructures in a data limited setting, generalizes well to previously unseen samples, is robust to image perturbations, and captures domain interpretable features. The presented work outperforms conventional deep learning baselines and is a first step towards improving understanding of process parameters and resulting material properties.

Kassab, Lara↗