Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Automated solar collector installation design including version management

Embodiments may include systems and methods to create and edit a representation of a worksite, to create various data objects, to classify such objects as various types of predefined “features” with attendant properties and layout constraints. As part of or in addition to classification, an embodiment may include systems and methods to create, associate, and edit intrinsic and extrinsic properties to these objects. A design engine may apply of design rules to the features described above to generate one or more solar collectors installation design alternatives, including generation of on-screen and/or paper representations of the physical layout or arrangement of the one or more design alternatives. Some embodiments may provide viewing, creating, and manipulating of multiple versions of a solar collector layout design for a particular installation worksite. The use of versions may allow analysis of alternative layouts, alternative feature classifications, and cost and performance data corresponding to alternative design choices. Version summary information providing a representative comparison between versions across a number of dimensions may be provided.

14 SOLAR ENERGY↗

Evaluation of artificial neural network performance for classification of potato plants infected with potato virus Y using spectral data on multiple varieties and genotypes

Potato virus Y (Potyviridae, PVY) is a plant virus that poses a significant threat to potato producers on a global basis. The pathogen has disrupted seed potato supplies and negatively impacted yield and quality of commercial potato crops. The potato industry currently manages PVY infection levels via insecticide applications, regional seed certification programs that rely on field scouting to visually assess individual plants for infection status, and destructive and costly tissue sampling coupled with laboratory assays. Despite these efforts, PVY continues to confound potato industry stakeholders resulting in economic harm. Remote sensing and machine learning provide for the development of new tools to more accurately detect and spatially quantify PVY-infected plants versus the current state of the art. However, there is a need to understand how the occurrence of many different potato varieties impact the dynamics of developing models to detect potato plants impacted with PVY and their potential effectiveness. This study evaluates classification modelling outcomes using spectral datasets collected in different temporal and spatial environments (greenhouse and a production field) on multiple potato varieties consisting of labelled instances of plants infected with PVY and those not infected with the virus. A modelling framework was developed to support iterative modelling runs using artificial neural network (ANN) architectures configured as binary classifiers to develop sample populations to support statistical analysis on model performance using specific spectral subsets. When using spectral data to detect PVY-infected plants, ANN models achieved the highest mean accuracy of 0.894 on a single variety. Conversely, the same ANN model architecture only achieved a mean accuracy of 0.575 on a spectral data set representing 29 potato breeding lines. Additionally, statistical analysis indicates spectral regions including the red edge, near infrared and shortwave infrared contain more important spectral features for the ANN classifier introduced in this research.

60 APPLIED LIFE SCIENCES↗

Using active learning to improve quasar identification for the DESI spectra processing pipeline

The Dark Energy Spectroscopic Instrument (DESI) survey uses an automatic spectral classification pipeline to classify spectra. QuasarNET is a convolutional neural network used as part of this pipeline originally trained using data from the Baryon Oscillation Spectroscopic Survey (BOSS). In this paper we implement an active learning algorithm to optimally select spectra to use for training a new version of the QuasarNET weights file using only DESI data, with the goal of improving classification accuracy. This active learning algorithm includes a novel outlier rejection step using a Self-Organizing Map to ensure we label spectra representative of the larger quasar sample observed in DESI. We perform two iterations of the active learning pipeline, assembling a final dataset of 5600 labeled spectra, a small subset of the approximately 1.3 million quasar targets in DESI's Data Release 1. When splitting the spectra into training and validation subsets we achieve similar performance to the previously trained weights file in completeness and purity calculated on the validation dataset but do so with less than one tenth of the amount of training data. The new weights also more consistently classify objects in the same way when used on unlabeled data compared to the old weights file. In the process of improving QuasarNET's classification accuracy we discovered a systemic error in QuasarNET's redshift estimation and used our findings to improve our understanding of QuasarNET's redshifts.

Machine learning↗

Logistic classification for tool life modeling in machining

This paper describes the application of logistic classification for tool life modeling and prediction in an industrial setting using shop floor data. Tool life is treated as a classification problem since tool wear can only be measured at the time of tool replacement in a production environment. Laboratory tool wear experiments are used to simulate shop floor wear data by two states: not worn (class 0); and worn (class 1). To incorporate non-linearity in logistic classification, a log-transformation of input features is performed. The logistic classification approach, results, and interpretability of the logistic model are presented.

Karandikar, Jaydeep↗

Question-answering system extracts information on injection drug use from clinical notes

Background. Injection drug use (IDU) can increase mortality and morbidity. Therefore, identifying IDU early and initiating harm reduction interventions can benefit individuals at risk. However, extracting IDU behaviors from patients’ electronic health records (EHR) is difficult because there is no other structured data available, such as International Classification of Disease (ICD) codes, and IDU is most often documented in unstructured free-text clinical notes. Although natural language processing can efficiently extract this information from unstructured data, there are no validated tools. Methods. Here, to address this gap in clinical information, we design a question-answering (QA) framework to extract information on IDU from clinical notes for use in clinical operations. Our framework involves two main steps: (1) generating a gold-standard QA dataset and (2) developing and testing the QA model. We use 2323 clinical notes of 1145 patients curated from the US Department of Veterans Affairs (VA) Corporate Data Warehouse to construct the gold-standard dataset for developing and evaluating the QA model. We also demonstrate the QA model’s ability to extract IDU-related information from temporally out-of-distribution data. Results. Here, we show that for a strict match between gold-standard and predicted answers, the QA model achieves a 51.65% F1 score. For a relaxed match between the gold-standard and predicted answers, the QA model obtains a 78.03% F1 score, along with 85.38% Precision and 79.02% Recall scores. Moreover, the QA model demonstrates consistent performance when subjected to temporally out-of-distribution data. Conclusions. Our study introduces a QA framework designed to extract IDU information from clinical notes, aiming to enhance the accurate and efficient detection of people who inject drugs, extract relevant information, and ultimately facilitate informed patient care.

60 APPLIED LIFE SCIENCES↗

CMOS-Based Single-Cycle in-Memory XOR/XNOR

Big data applications are on the rise, and so is the number of data centers. The ever-increasing massive data pool needs to be periodically backed up in a secure environment. Moreover, a massive amount of securely backed-up data is required for training binary convolutional neural networks for image classification. XOR and XNOR operations are essential for large-scale data copy verification, encryption, and classification algorithms. The disproportionate speed of existing compute and memory units makes the von Neumann architecture inefficient to perform these Boolean operations. Compute-in-memory (CiM) has proved to be an optimum approach for such bulk computations. The existing CiM-based XOR/XNOR techniques either require multiple cycles for computing or add to the complexity of the fabrication process. Here, we propose a CMOS-based hardware topology for single-cycle in-memory XOR/XNOR operations. Our design provides at least 2× improvement in the latency compared with other existing CMOS-compatible solutions. We verify the proposed system through circuit/system-level simulations and evaluate its robustness using a 5000-point Monte Carlo variation analysis. This all-CMOS design paves the way for practical implementation of CiM XOR/XNOR at scaled technology nodes.

97 MATHEMATICS AND COMPUTING↗

AI-NERD: Elucidation of relaxation dynamics beyond equilibrium through AI-informed X-ray photon correlation spectroscopy

Abstract Understanding and interpreting dynamics of functional materials in situ is a grand challenge in physics and materials science due to the difficulty of experimentally probing materials at varied length and time scales. X-ray photon correlation spectroscopy (XPCS) is uniquely well-suited for characterizing materials dynamics over wide-ranging time scales. However, spatial and temporal heterogeneity in material behavior can make interpretation of experimental XPCS data difficult. In this work, we have developed an unsupervised deep learning (DL) framework for automated classification of relaxation dynamics from experimental data without requiring any prior physical knowledge of the system. We demonstrate how this method can be used to accelerate exploration of large datasets to identify samples of interest, and we apply this approach to directly correlate microscopic dynamics with macroscopic properties of a model system. Importantly, this DL framework is material and process agnostic, marking a concrete step towards autonomous materials discovery.

36 MATERIALS SCIENCE↗

Citizen science for IceCube: Name that Neutrino

Name that Neutrino is a citizen science project where volunteers aid in classification of events for the IceCube Neutrino Observatory, an immense particle detector at the geographic South Pole. From March 2023 to September 2023, volunteers did classifications of videos produced from simulated data of both neutrino signal and background interactions. Name that Neutrino obtained more than 128,000 classifications by over 1800 registered volunteers that were compared to results obtained by a deep neural network machine-learning algorithm. Possible improvements for both Name that Neutrino and the deep neural network are discussed.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Spectroscopic Data Processing Pipeline for the Dark Energy Spectroscopic Instrument

Abstract We describe the spectroscopic data processing pipeline of the Dark Energy Spectroscopic Instrument (DESI), which is conducting a redshift survey of about 40 million galaxies and quasars using a purpose-built instrument on the 4 m Mayall Telescope at Kitt Peak National Observatory. The main goal of DESI is to measure with unprecedented precision the expansion history of the universe with the baryon acoustic oscillation technique and the growth rate of structure with redshift space distortions. Ten spectrographs with three cameras each disperse the light from 5000 fibers onto 30 CCDs, covering the near-UV to near-infrared (3600–9800 Å) with a spectral resolution ranging from 2000 to 5000. The DESI data pipeline generates wavelength- and flux-calibrated spectra of all the targets, along with spectroscopic classifications and redshift measurements. Fully processed data from each night are typically available to the DESI collaboration the following morning. We give details about the pipeline’s algorithms, and provide performance results on the stability of the optics, the quality of the sky background subtraction, and the precision and accuracy of the instrumental calibration. This pipeline has been used to process the DESI Survey Validation data set, and has exceeded the project’s requirements for redshift performance, with high efficiency and a purity greater than 99% for all target classes.

79 ASTRONOMY AND ASTROPHYSICS↗

Vegetation classification map and covariates associated with NEON AOP survey, East River, CO 2018

This package includes geospatial data layers developed to investigate how environmental gradients—specifically topography and near-surface soil properties—drive the spatial arrangement of dominant plant communities in mountainous watersheds. The geospatial products, which support the analysis of these ecological relationships, are derived from airborne hyperspectral and LiDAR datasets acquired by the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP), in conjunction with an extensive ground field campaign conducted in summer 2018. This work is part of the DOE Watershed Function Science Focus Area (SFA) and features geospatial datasets developed based on observations and ground data collected at East River, Colorado, in collaboration with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey in June 2018. Classification Map: - Classification Map (PNG, GeoTIFF): Derived from hyperspectral and LiDAR airborne data using a machine learning approach. - Class Code Mapper (CSV): Associates pixel values with corresponding vegetation/non-vegetation classes. - Classification Reference Data (CSV): Reference data used in the machine learning procedure. LiDAR-Derived Products: - Topographical Metrics (GeoTIFFs): Elevation, slope, curvature, TWI, TPI, solar insolation, and canopy height model (CHM), smoothed with a 5x5 pixel window. Vegetation Indices: - GeoTIFFs of NDVI, NDNI, NDWI: Vegetation indices derived from hyperspectral data. Urban Masks: - Urban Mask (GeoTIFF): Applied to the mapping to convert bare soil classes to urban classes. Software Compatibility: GeoTIFFs: Can be visualized with GIS software or libraries that support GeoTIFF images. CSV Files: Can be opened with any software that handles comma-separated values. The FLMD file provides details and links to the source datasets used to derive the products. The manuscript (in the Method session) provides details on how each product was derived. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Update on 2026-03-25: Since the original dataset publication date of 02/28/2020, this package has a new classification map derived by an improved methodology. This update also includes additional ground data that improved the representation of some of the communities. See the methods for further details on what has changed between versions.

2018 NEON and 2025 CHESS Campaigns↗

Confidentiality-preserving machine learning algorithms for soft-failure detection in optical communication networks

Automated fault management is at the forefront of next-generation optical communication networks. The increase in complexity of modern networks has triggered the need for programmable and software-driven architectures to support the operation of agile and self-managed systems. In these scenarios, the European Telecommunications Standards Institute zero-touch network and service management approach is imperative. The need for machine learning algorithms to process the large volume of telemetry data brings safety concerns as distributed cloud-computing solutions become the preferred approach for deploying reliable communication network automation. This paper’s contribution is twofold. First, we propose a simple yet effective method to guarantee the confidentiality of the telemetry data based on feature scrambling. The method allows the operation of third-party computational services without direct access to the full content of the collected data. Additionally, the effectiveness of four unsupervised machine learning algorithms for soft-failure detection is evaluated when applied to the scrambled telemetry data. The methods are based on factor analysis, principal component analysis, nonlinear principal component analysis, and singular value decomposition. Most dimensionality reduction algorithms have the common property that they can maintain similar levels of fault classification performance while hiding the data structure from unauthorized access. Evaluations of the proposed algorithms demonstrate this capability.

97 MATHEMATICS AND COMPUTING↗

Federated Machine Learning-Based Anomaly Detection System for Synchrophasor Network Using Heterogeneous Data Sets: Preprint

Synchrophasor technology is widely deployed in the energy management system to monitor the grid health at micro level and perform necessary corrective actions in real time; however, integrated phasor devices and data aggregators are exposed to several cybersecurity threats. This paper proposes a federated ML(FML)-based ADS to detect several data integrity attacks in the synchrophasor network. The proposed approach integrates the horizontal FML technique and consists of substation-based local models and a control center-based global model. The proposed methodology includes training local models using heterogeneous data sets that include network and grid information and updating the global model through multiple iterations by sharing model gradients. Finally, the trained global model is applied to identify cyberattacks, normal operation, and physical events. To validate the proof of concept, we used synthetic data sets generated by Mississippi State University and Oak Ridge National Laboratory for training and testing the classification models using the National Renewable Energy Laboratory's high performance computing resources. Our experimental results, computed through several performance measures, reveal that the proposed approach shows consistent performance during the binary, three-class, and multiclass classifications while ensuring privacy of synchrophasor data.

anomaly detection system↗

ATAT: Astronomical Transformer for time series and Tabular data

Context. The advent of next-generation survey instruments, such as theVera C. RubinObservatory and its Legacy Survey of Space and Time (LSST), is opening a window for new research in time-domain astronomy. The Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC) was created to test the capacity of brokers to deal with a simulated LSST stream. Aims. Our aim is to develop a next-generation model for the classification of variable astronomical objects. We describe ATAT, the Astronomical Transformer for time series And Tabular data, a classification model conceived by the ALeRCE alert broker to classify light curves from next-generation alert streams. ATAT was tested in production during the first round of the ELAsTiCC campaigns. Methods. ATAT consists of two transformer models that encode light curves and features using novel time modulation and quantile feature tokenizer mechanisms, respectively. ATAT was trained on different combinations of light curves, metadata, and features calculated over the light curves. We compare ATAT against the current ALeRCE classifier, a balanced hierarchical random forest (BHRF) trained on human-engineered features derived from light curves and metadata. Results. When trained on light curves and metadata, ATAT achieves a macro F1 score of 82.9 ± 0.4 in 20 classes, outperforming the BHRF model trained on 429 features, which achieves a macro F1 score of 79.4 ± 0.1. Conclusions. The use of transformer multimodal architectures, combining light curves and tabular data, opens new possibilities for classifying alerts from a new generation of large etendue telescopes, such as theVera C. RubinObservatory, in real-world brokering scenarios.

Astronomy & Astrophysics↗

Time projection chamber for GADGET II

The established Gaseous Detector with Germanium Tagging (GADGET) detection system is used to measure weak, low-energy 𝛽-delayed proton decays. It consists of the Gaseous Proton Detector equipped with a MICROMEGAS (MM) readout to detect protons and other charged particles calorimetrically, surrounded by the Segmented Germanium Array (SeGA) for high-resolution detection of prompt 𝛾 rays. To upgrade GADGET's Proton Detector to operate as a compact time projection chamber (TPC) for the detection, three-dimensional imaging and identification of low-energy 𝛽-delayed single- and multiparticle emissions mainly of interest to astrophysical studies. A new high granularity MM board with 1024 pads has been designed, fabricated, installed, and tested. A high-density data acquisition system based on generic electronics for TPCs (GET) has been installed and optimized to record and process the gas avalanche signals collected on the readout pads. The TPC's performance has been tested using a 220 Rn 𝛼-particle source and cosmic-ray muons. In addition, decay events in the TPC have been simulated by adapting the attpcroot data analysis framework. Furthermore, a novel application of two-dimensional convolutional neural networks for GADGET II event classification is introduced. The optimization of data throughput is also addressed. The GADGET II TPC is capable of detecting and identifying 𝛼 particles as well as measuring their track direction, range, and energy. The extracted energy resolution of the GADGET II TPC using P10 gas is about 5.4% at 6.288 MeV ( 220 Rn 𝛼 events), computed using charge integration. Based on a systematic simulation study, we estimated the detection efficiency of the GADGET II TPC for protons and 𝛼 particles, respectively. It has also been demonstrated that the GADGET II TPC is capable of tracking minimum-ionizing particles (i.e., cosmic-ray muons). From these measurements, the electron drift velocity was measured under typical operating conditions. In addition to being one of the first generation of micropattern gaseous detectors (MPGDs) to utilize a resistive anode applied to low-energy nuclear physics, the GADGET II TPC will also be the first TPC surrounded by a high-efficiency array of high-purity germanium 𝛾-ray detectors. As a result, the TPC of GADGET II has been designed, fabricated, and tested and is ready for operation at the Facility for Rare Isotope Beams for radioactive-beam-line experiments.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

On the transferability of residence time distributions in two 10-km long river sections with similar hydromorphic units

Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.

54 ENVIRONMENTAL SCIENCES↗

Life-cycle energy use and greenhouse gas emissions of palm fatty acid distillate derived renewable diesel

This study aims to quantify life-cycle fossil energy use and greenhouse gas (GHG) emissions for palm fatty acid distillate (PFAD) derived renewable diesel (RD) taking into consideration different feedstock classifications that are applicable to PFAD (residue, byproduct, or coproduct), and incorporating updated data for key processes. Under the three classifications, the PFAD to RD pathway was modeled using the Greenhouse gases, Regulated Emissions, and Energy Use in Technologies (GREET®) model. PFAD-derived RD could reduce fossil energy consumption by 77%–88%, relative to petroleum diesel. GHG emissions are very sensitive to PFAD classification and coproduct handling methods. Considering the production of palm oil and PFAD and economic value, we maintain that PFAD should be treated as a byproduct in palm oil refineries. With this treatment, PFAD-derived RD could achieve 84% GHG emissions reductions, compared to the emissions of petroleum diesel. We also employed a substitution method to address the substitution of PFAD by other materials in the marketplace. Compared to coproduct allocation results, we found substituting PFAD by tallow, soy oil, barley, and canola oil results in lower GHG emissions. Due to high induced land-use change emissions associated with palm farming, if PFAD is treated as a coproduct with refined palm oil, PFAD-derived RD may not deliver GHG reductions. A sensitivity analysis identified key parameters such as palm fruit yield, oil extraction efficiency in oil mills, and energy use intensity for RD production affects LCA results significantly; future efforts to improve these parameters could result in further GHG reductions.

09 BIOMASS FUELS↗

Multimodal X-ray nano-spectromicroscopy analysis of chemically heterogeneous systems

Abstract Understanding the nanoscale chemical speciation of heterogeneous systems in their native environment is critical for several disciplines such as life and environmental sciences, biogeochemistry, and materials science. Synchrotron-based X-ray spectromicroscopy tools are widely used to understand the chemistry and morphology of complex material systems owing to their high penetration depth and sensitivity. The multidimensional (4D+) structure of spectromicroscopy data poses visualization and data-reduction challenges. This paper reports the strategies for the visualization and analysis of spectromicroscopy data. We created a new graphical user interface and data analysis platform named XMIDAS (X-ray multimodal image data analysis software) to visualize spectromicroscopy data from both image and spectrum representations. The interactive data analysis toolkit combined conventional analysis methods with well-established machine learning classification algorithms (e.g. nonnegative matrix factorization) for data reduction. The data visualization and analysis methodologies were then defined and optimized using a model particle aggregate with known chemical composition. Nanoprobe-based X-ray fluorescence (nano-XRF) and X-ray absorption near edge structure (nano-XANES) spectromicroscopy techniques were used to probe elemental and chemical state information of the aggregate sample. We illustrated the complete chemical speciation methodology of the model particle by using XMIDAS. Next, we demonstrated the application of this approach in detecting and characterizing nanoparticles associated with alveolar macrophages. Our multimodal approach combining nano-XRF, nano-XANES, and differential phase-contrast imaging efficiently visualizes the chemistry of localized nanostructure with the morphology. We believe that the optimized data-reduction strategies and tool development will facilitate the analysis of complex biological and environmental samples using X-ray spectromicroscopy techniques.

36 MATERIALS SCIENCE↗

Predicting Drug Effects from High-dimensional Asymmetric Drug Data Sets using Graph Neural Networks: A Comprehensive Analysis of Multi-target Drug Effect Prediction

Graph neural networks (GNNs) have emerged as one of the most effective Machine learning (ML) techniques for drug effect prediction from drug molecular graphs. Despite having immense potential, GNN models lack performance when using data sets that contain high dimensional asymmetrically co-occurrent drug effects as targets with complex correlations between them. Training individual learning models for each drug effect and incorporating every prediction result for a wide spectrum of drug effects is beyond practicality. Such an implication provides a testbed to address this challenge as multi-target prediction problems, aiming to predict all drug effects at a time. We develop standard and hybrid graph neural networks (GNNs)to perform two separate tasks that are multi-regression for continuous values and multi-label classification for categorical values contained in our data sets. Since this step makes the target data even more sparse and introduces asymmetric label co-occurrence, the learning of multi-label classification models becomes difficult and heavily impacts the GNN's performance. To address these challenges, we propose a new data oversampling technique to improve multi-label classification performances on all the given imbalanced molecular graph data sets. Using the technique, we improve the data imbalance ratio of the drug effects better than before while protecting the data set's integrity. Finally, we evaluate multi-label classification performance using the best-performant hybrid GNN model on all the oversampled data sets obtained from the proposed oversampling technique. These results outperform those of other ML models including GNN models when they are trained on the original data sets or oversampled data sets using MLSMOTE (a well-known oversampling technique) in all evaluation metrics precision, recall, and F1 score by a significant margin.

Bose, Avishek [ORNL]↗