Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection

Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection Description This dataset contains input and output data for the manuscript Mongird, K. et al. (under review) titled "Energy Infrastructure Futures: A Multiscale Evaluation of Projected Power Plant Siting Across the Western Interconnection". Input data corresponds to gridded spatial siting attributes that are necessary to conduct a random forest machine learning analysis of siting feature importance. Output data includes SHAP feature analysis outputs, and classification report values. For data on power plant siting results referred to in the manuscript, please refer to the CERF: IM3 Projected Western US Power Plant Locations data download page. The downloadable data includes values for eight different future scenarios for the Western US. The scenarios include combinations of two Shared Socioeconomic Pathways (SSP3 and SSP5) with four high-resolution climate projections specific to the United States (see, https://tgw-data.msdlive.org/). These climate projections include "hotter" and "cooler" variants for two Representative Concentration Pathways (RCP4.5 and RCP8.5). The resulting eight simulations are: rcp45cooler_ssp3 rcp45cooler_ssp5 rcp45hotter_ssp3 rcp45hotter_ssp5 rcp85cooler_ssp3 rcp85cooler_ssp5 rcp85hotter_ssp3 rcp85hotter_ssp5 Technical Information The dataset includes two sets of data files: (1) CERF gridded siting parameters and (2) Feature analysis outputs and classification reports. All downloadable data is in csv file format. Files with x/y coordinate information use the Albers Equal Area Conic projection (ESRI:102003). 1. CERF Gridded Siting Parameters This directory provides a balanced sample of gridded CERF siting parameters data for eight different scenarios for the Western US through 2055, seven different technologies, and eight timesteps. This data serves as input to the feature analysis. It contains the following parameters. region_name - name of region (i.e., state) sited - binary value representing whether the grid cell received a siting of that technology type (1=True) rcp - binary value representing scenario resource concentration pathway (0 = RCP4.5, 1 = RCP8.5) ssp - binary value representing scenario shared socioeconomic pathway (0 = SSP3, 1 = SSP5) climate - binary value representing cooler (0) or hotter (1) GCM forcing tech_name - generation technology name sited_year - year that values correspond to transmission_cost - cost of transmission interconnection pipeline_cost - cost of natural gas pipeline interconnection interconnection_cost - total interconnection cost (sum of transmission cost and gas pipeline cost) lmp - associated locational marginal value ($/MWh) associated with the grid cell, timestep, scenario, and technology xcoord - x-coordinate of location ycoord - y-coordinate of location 2a. Feature Analysis Output The dataset includes the feature analysis shap output for locational marginal price and interconnection cost. It contains the following parameters. technology - generator technology name scenario - name of scenario feature - name of feature, either locational_marginal_price or interconnection_cost value - the mean of absolute value of SHAP values for given feature 2b. Feature Analysis Classification Report This download includes the classification report associated with each random forest model. The dataset contains the following parameters. technology - generation technology name scenario - name of scenario test - one of precision (the proportion of predicted positives that are actually correct), recall (the proportion of actual positives that were correctly identified), f1-score (the harmonic mean of precision and recall) 0.0 - value of test for classification of 0 (grid cell not chosen for siting) 1.0 - value of test for classification of 1 (grid cell chosen for siting) accuracy - accuracy of model (i.e., fraction of all predictions that were right) macro avg - Simple average of test values for all classes weighted avg - Weighted average of test values for all classes, weighted based on Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program. License This data is made available under a CCBY4 License Disclaimer This material was prepared as an account of work sponsored by an agency of the United States Government. Neither the United States Government nor the United States Department of Energy, nor the Contractor, nor any or their employees, nor any jurisdiction or organization that has cooperated in the development of these materials, makes any warranty, express or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness or any information, apparatus, product, software, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States Government or any agency thereof, or Battelle Memorial Institute. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States Government or any agency thereof. PACIFIC NORTHWEST NATIONAL LABORATORYoperated byBATTELLEfor theUNITED STATES DEPARTMENT OF ENERGYunder Contract DE-AC05-76RL01830

Mongird, Kendall [Pacific Northwest National Labor↗

Identifying Climate Patterns Using Clustering Autoencoder Techniques

Abstract The complexity of growing spatiotemporal resolution of climate simulations produces a variety of climate patterns under different projection scenarios. This paper proposes a new data-driven climate classification workflow via an unsupervised deep learning technique that can dimensionally reduce the vast volume of spatiotemporal numerical climate projection data into a compact representation. We aim to identify distinct zones that capture multiple climate variables as well as their future changes under different climate change scenarios. Our approach leverages convolutional autoencoders combined with k -means clustering (standard autoencoder) and online clustering based on the Sinkhorn–Knopp algorithm (clustering autoencoder) across the conterminous United States (CONUS) to capture unique climate patterns in a data-driven fashion from the Geophysical Fluid Dynamics Laboratory Earth System Model with GOLD component (GFDL-ESM2G). The developed approach compresses 70 years of GFDL-ESM2G simulation at 0.125° spatial resolution across the CONUS under multiple warming scenarios to a lower-dimensional space by a factor of 660 000 and then tested on 150 years of GFDL-ESM2G simulation data. The results show that five climate clusters capture physically reasonable and spatially stable climatological patterns matched to known climate classes defined by human experts. Results also show that using a clustering autoencoder can reduce the computational time for clustering by up to 9.2 times when compared to using a standard autoencoder. Our five unique climate patterns resulting from the deep learning–based clustering of the lower-dimensional space thereby enable us to provide insights on hydrometeorology and its spatial heterogeneity across the conterminous United States immediately without downloading large climate datasets. Significance Statement This paper presents a data-driven climate classification approach using unsupervised deep learning to dimensionally reduce climate model outputs and to identify distinct climate regions for their future changes. Our approach compresses climate information for 70 years of Geophysical Fluid Dynamics Laboratory Earth System Model data across the conterminous United States (CONUS) at 0.125° spatial resolution. The results reveal that five climate clusters capture reasonable and stable climatological patterns matched to known climate patterns. The embedded clustering process in deep learning provides ×9.2 times faster execution than the k -means clustering technique. These results give us insight about climate spatial patterns and heterogeneity of hydrological patterns across the conterminous United States without downloading large climate datasets.

Kurihana, Takuya↗

Multi-Source Data Aggregation and Real-Time Anomaly Classification and Localization in Power Distribution Systems

This paper proposes a real-time anomaly location and classification framework for power distribution systems to simultaneously determine the type of anomaly (i.e., short-circuit fault, cyber attack, DER switching) and its location. The proposed framework employs the data aggregation module to collect the measurement data from multiple field devices operating at different sampling rates, such as protection relays and D-PMUs. The output of the data aggregation is then fed into a multi-task learning-based long-based short-term memory (MTL-LSTM) to classify the type of anomaly and the location in two separate tasks. The proposed MTL-LSTM approach can be utilized in real-time operation in order to distinguish between normal and several anomalous operations and locate the anomaly. The proposed framework is tested on a modified IEEE 33-bus test feeder benchmark that integrates solar generation and energy storage. Furthermore, the results show that the proposed framework can locate and classify anomalies for several operation conditions with more than 96% accuracy. Further experiments highlight the impact of aggregating multiple sources of data on the performance of the proposed model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Image feature extraction and galaxy classification: a novel and efficient approach with automated machine learning

ABSTRACT In this work, we explore the possibility of applying machine learning methods designed for 1D problems to the task of galaxy image classification. The algorithms used for image classification typically rely on multiple costly steps, such as the point spread function deconvolution and the training and application of complex Convolutional Neural Networks of thousands or even millions of parameters. In our approach, we extract features from the galaxy images by analysing the elliptical isophotes in their light distribution and collect the information in a sequence. The sequences obtained with this method present definite features allowing a direct distinction between galaxy types. Then, we train and classify the sequences with machine learning algorithms, designed through the platform Modulos AutoML. As a demonstration of this method, we use the second public release of the Dark Energy Survey (DES DR2). We show that we are able to successfully distinguish between early-type and late-type galaxies, for images with signal-to-noise ratio greater than 300. This yields an accuracy of $86{{\ \rm per\ cent}}$ for the early-type galaxies and $93{{\ \rm per\ cent}}$ for the late-type galaxies, which is on par with most contemporary automated image classification approaches. The data dimensionality reduction of our novel method implies a significant lowering in computational cost of classification. In the perspective of future data sets obtained with e.g. Euclid and the Vera Rubin Observatory, this work represents a path towards using a well-tested and widely used platform from industry in efficiently tackling galaxy classification problems at the peta-byte scale.

79 ASTRONOMY AND ASTROPHYSICS↗

One of These Things IS Like the Other: Pursuing a New Taxonomy of Industry for Improved Energy System Modeling

Industrial processes drive the exchange of materials, energy, and currency throughout the economy. These processes are powered by electricity and direct combustion, with variation in their operation even within the same industry. This heterogeneity makes it difficult for large models, including the National Energy Modeling System (US), to project their energy use while remaining tractable. Decarbonization and ensuing changes to the energy system require changes to industrial processes while offering opportunities for process innovation, but the extent and nature of changes are difficult to model with current classification schemes and corresponding data. The North American Industrial Classification (NAICS) is an economic taxonomy of industries, but its categories are less meaningful from an energy and material flow perspective. For example, a facility that makes steel from iron ore in a blast furnace/basic oxygen furnace is categorized under the same NAICS code as a facility that makes steel from scrap in an electric arc furnace despite the scale, use of recycled scrap versus iron ore, and energy use differences in the two facility types. Exploratory analysis is performed on a large dataset used for plant-level energy assessment in order to detect clusters that can aid in better modeling of industry for energy analysis in an evolving system with breakthrough technologies.

28 EE - Advanced Manufacturing Office (EE-5A)↗

Classification and Fusion of Two Disparate Data Streams and Nuclear Dissolutions Application

We consider two streams of data or measurements with disparate qualities and time resolutions that need to be classified. The first stream consists of higher quality data at a coarser time resolution, and the other consists of lower quality data at a finer time resolution. We present a fuser-switch method that fuses the set of classifiers of each stream separately and switches between them. We show that this method provides classification decisions at a finer time resolution with superior detection and false alarm probabilities compared to individual classifiers, under the statistical independence and time resolution ratio conditions. When classifiers are trained using machine learning methods, we show that this superior performance is guaranteed with a confidence probability specified by the classifiers' generalization equations. We use these results to provide analytical foundations for previous practical results that achieved significant performance improvements in classifying Pu/Np target dissolution events at a radiochemical processing facility.

Rao, Nageswara↗

Addressing the dynamic nature of reference data: a new nucleotide database for robust metagenomic classification

Accurate metagenomic classification relies on comprehensive, up-to-date, and validated reference databases. While the NCBI BLAST Nucleotide (nt) database, encompassing a vast collection of sequences from all domains of life, represents an invaluable resource, its massive size—currently exceeding 10 12 nucleotides—and exponential growth pose significant challenges for researchers seeking to maintain current nt-based indices for metagenomic classification. Recognizing that no current nt-based indices exist for the widely used Centrifuge classifier, and the last public version currently available was released in 2018, we addressed this critical gap by leveraging advanced high-performance computing resources. We present new Centrifuge-compatible nt databases, meticulously constructed using a novel pipeline incorporating different quality control measures, including reference decontamination and filtering. These measures demonstrably reduce spurious classifications, as shown through our reanalysis of published metagenomic data where Plasmodium annotations were dramatically reduced using our decontaminated database, highlighting how database quality can significantly impact research conclusions. Through temporal comparisons, we also reveal how our approach minimizes inconsistencies in taxonomic assignments stemming from asynchronous updates between public sequence and taxonomy databases. These discrepancies are particularly evident in taxa such as Listeria monocytogenes and Naegleria fowleri, where classification accuracy varied significantly across database versions. These new databases, made available as pre-built Centrifuge indexes, respond to the need for an open, robust, nt-based pipeline for taxonomic classification in metagenomics. Applications such as environmental metagenomics, forensics, and clinical metagenomics, which require comprehensive taxonomic coverage, will benefit from this resource. Our work highlights the importance of treating reference databases as dynamic entities, subject to ongoing quality control and validation akin to software development best practices. This approach is crucial for ensuring accuracy and reliability of metagenomic analysis, especially as databases continue to expand in size and complexity.

59 BASIC BIOLOGICAL SCIENCES↗

Automated classification of big X-ray diffraction data using deep learning models

Abstract In current in situ X-ray diffraction (XRD) techniques, data generation surpasses human analytical capabilities, potentially leading to the loss of insights. Automated techniques require human intervention, and lack the performance and adaptability required for material exploration. Given the critical need for high-throughput automated XRD pattern analysis, we present a generalized deep learning model to classify a diverse set of materials’ crystal systems and space groups. In our approach, we generate training data with a holistic representation of patterns that emerge from varying experimental conditions and crystal properties. We also employ an expedited learning technique to refine our model’s expertise to experimental conditions. In addition, we optimize model architecture to elicit classification based on Bragg’s Law and use evaluation data to interpret our model’s decision-making. We evaluate our models using experimental data, materials unseen in training, and altered cubic crystals, where we observe state-of-the-art performance and even greater advances in space group classification.

Chemistry↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

Multimodal Few-Shot Segmentation of Electron Micrographs

Scanning transmission electron microscopy (STEM) is one of the most used methods of analyzing the chemistry and composition of materials. By analyzing microstructures, these microscopes can help scientists better understand the molecular underpinnings of microelectronics, batteries, and more. However, STEM data can be difficult to interpret, so recent developments have been made in applications of machine learning to analyze these images. The PNNL-developed pyCHIP Classifier has achieved results in segmenting STEM these images via few-shot learning, a method which requires little data and human input, perfect for quickly analysis. In my internship I (Eli Meyers) investigated a multimodal improvement of this classifier by incorporating energy dispersive x-ray spectroscopy (EDS) data into the classification process for a more accurate segmentation. Furthermore, I encoded the spectral data by training a mass spectrometry encoder on the EDS data to extract a more meaningful representation of the data.

36 MATERIALS SCIENCE↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CNN-Based Phase Fault Classification in Real and Simulated Power Systems Data

This study proposes a convolutional neural network (CNN)–based two-step phase fault detection and identification method to classify anomalies in the power grid signal. Specifically, the first step checks the fault’s existence and determines the need for the second step. Subsequently, in the case of anomalies in the power grid signal, the second step identifies the type of fault, including line-to-line, single-line-to-ground, double-line-to-ground, and triple-line. Accordingly, the CNN architecture is both designed for the classification layers and trained with simulated data. To provide maximum prediction accuracy with minimum processing time, this study investigates the combinations of various feature extraction (FE) techniques, such as fast Fourier transform (FFT), amplitude and phase (AP), auto-correlation function, power spectral density, and wavelet transform (WT). Consequently, simulated and real-world results demonstrate that the proposed two-step method outperforms conventional one-step techniques, with the best performance obtained by using the combination of AP-AP, AP-WT, FFT-AP, and FFT-WT–based FE methods.

Alaca, Ozgur↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

Discovering Ca II absorption lines with a neural network

Quasar absorption line analysis is critical for studying gas and dust components and their physical and chemical properties as well as the evolution and formation of galaxies in the early universe. Calcium II (Ca II ) absorbers, which are one of the dustiest absorbers and are located at lower redshifts than most other absorbers, are especially valuable when studying physical processes and conditions in recent galaxies. However, the number of known quasar Ca II absorbers is relatively low due to the difficulty of detecting them with traditional methods. In this work, we developed an accurate and quick approach to search for Ca II absorption lines using deep learning. In our deep learning model, a convolutional neural network, tuned using simulated data, is used for the classification task. The simulated training data are generated by inserting artificial Ca II absorption lines into original quasar spectra from the Sloan Digital Sky Survey (SDSS), while an existing Ca II catalogue is adopted as the test set. The resulting model achieves an accuracy of 96 per cent on the real data in the test set. Our solution runs thousands of times faster than traditional methods, taking a fraction of a second to analyse thousands of quasars, while traditional methods may take days to weeks. The trained neural network is applied to quasar spectra from SDSS’s DR7 and DR12 and discovered 399 new quasar Ca II absorbers. In addition, we confirmed 409 known quasar Ca II absorbers identified previously by other research groups through traditional methods.

79 ASTRONOMY AND ASTROPHYSICS↗

Rotational Equivariance for Object Classification Using xView

With the recent addition of large, curated and labeled data sets to the remote sensing discipline, deep learning models have largely surpassed the performance of classical techniques. These deep models, typically Convolutional Neural Networks, are invariant to translation through the use of successive convolution layers which are themselves equivariant to translation. Further, the combination of multiple convolution and pooling layers means that in practice, the model is also approximately invariant to translation. However, until recently these models could only approach rotational invariance through data augmentation. Here we propose using a new model formulation which achieves rotational equaivariance without data augmentation for overhead imagery classification. We utilize the popular xView data set to compare the rotational equivariance formalization against a regular CNN and CNN with rotational data augmentation for the task of image classification.

Bynum, Lucius EJ↗

Block Island Seafloor Sediment and Geological Data

This dataset contains seafloor sediment composition and geological data for the Block Island region, including sediment sample locations, grain size distributions, and seafloor substrate classifications from multiple data sources.

17 WIND ENERGY↗

Photovoltaic System Health-State Architecture for Data-Driven Failure Detection

The timely detection of photovoltaic (PV) system failures is important for maintaining optimal performance and lifetime reliability. A main challenge remains the lack of a unified health-state architecture for the uninterrupted monitoring and predictive performance of PV systems. To this end, existing failure detection models are strongly dependent on the availability and quality of site-specific historic data. The scope of this work is to address these fundamental challenges by presenting a health-state architecture for advanced PV system monitoring. The proposed architecture comprises of a machine learning model for PV performance modeling and accurate failure diagnosis. The predictive model is optimally trained on low amounts of on-site data using minimal features and coupled to functional routines for data quality verification, whereas the classifier is trained under an enhanced supervised learning regime. The results demonstrated high accuracies for the implemented predictive model, exhibiting normalized root mean square errors lower than 3.40% even when trained with low data shares. The classification results provided evidence that fault conditions can be detected with a sensitivity of 83.91% for synthetic power-loss events (power reduction of 5%) and of 97.99% for field-emulated failures in the test-bench PV system. Finally, this work provides insights on how to construct an accurate PV system with predictive and classification models for the timely detection of faults and uninterrupted monitoring of PV systems, regardless of historic data availability and quality. Such guidelines and insights on the development of accurate health-state architectures for PV plants can have positive implications in operation and maintenance and monitoring strategies, thus improving the system’s performance.

photovoltaics↗

Flame stability analysis of flame spray pyrolysis by artificial intelligence

Flame spray pyrolysis (FSP) is a process used to synthesize nanoparticles through the combustion of an atomized precursor solution; this process has applications in catalysts, battery materials, and pigments. Current limitations revolve around understanding how to consistently achieve a stable flame and the reliable production of nanoparticles. Machine learning and artificial intelligence algorithms that detect unstable flame conditions in real time may be a means of streamlining the synthesis process and improving FSP efficiency. In this study, the FSP flame stability is first quantified by analyzing the brightness of the flame's anchor point. This analysis is then used to label data for both unsupervised and supervised machine learning approaches. The unsupervised learning approach allows for autonomous labeling and classification of new data by representing data in a reduced dimensional space and identifying combinations of features that most effectively cluster it. The supervised learning approach, on the other hand, requires human labeling of training and test data but is able to classify multiple objects of interest (such as the burner and pilot flames) within the video feed. The accuracy of each of these techniques is compared against the evaluations of human experts. Both the unsupervised and supervised approaches can track and classify FSP flame conditions in real time to alert users of unstable flame conditions. This research has the potential to autonomously track and manage flame spray pyrolysis as well as other flame technologies by monitoring and classifying the flame stability.

42 ENGINEERING↗