Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “classification models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Knowledge Graph Entity Linking using Graph Embeddings

Details the use of a custom embedding model on knowledge graphs to aid in downstream natural language processing (NLP) models for Derivative Classification Assist. Motivations, algorithms, and results were discussed.

Mahesh, Aarav [Sandia National Laboratories (SNL-N↗

A Hierarchical Feature-Based Methodology to Perform Cervical Cancer Classification

Prevention of cervical cancer could be performed using Pap smear image analysis. This test screens pre-neoplastic changes in the cervical epithelial cells; accurate screening can reduce deaths caused by the disease. Pap smear test analysis is exhaustive and repetitive work performed visually by a cytopathologist. This article proposes a workload-reducing algorithm for cervical cancer detection based on analysis of cell nuclei features within Pap smear images. We investigate eight traditional machine learning methods to perform a hierarchical classification. We propose a hierarchical classification methodology for computer-aided screening of cell lesions, which can recommend fields of view from the microscopy image based on the nuclei detection of cervical cells. We evaluate the performance of several algorithms against the Herlev and CRIC databases, using a varying number of classes during image classification. Results indicate that the hierarchical classification performed best when using Random Forest as the key classifier, particularly when compared with decision trees, k-NN, and the Ridge methods.

60 APPLIED LIFE SCIENCES↗

Predictive models enhance feedstock quality of corn stover via air classification

Feedstock heterogeneity is a fundamental obstacle to cost-competitive biobased products. Agricultural products like corn stover have anatomical components that vary in their chemical composition, mechanical properties, structure, and response to chemical and biological treatments. A technique that can enrich streams in select anatomical fractions would allow a tailored deconstruction approach to increase overall process efficiency. Air classification can be leveraged for such refining, however, fundamental characterization and understanding of the particle properties that underly the physics of air classification are only modestly documented. Here, we determine fundamental particle properties including mass-to-area ratio, drag coefficient, and partition velocity that describe how anatomical tissues of corn stover behave during air classification. In this work, mass-to-area ratios of anatomical tissues vary by nearly two orders of magnitude from 2.3 mg/mm 2 for cob to 0.04 mg/mm 2 for leaf. Drag coefficients of longer, fibrous materials (i.e., rind, husk, and sheath) are shown to correlate with particle area (p-value < 0.001) whereas granular tissues (i.e., cob, pith, and leaf) correlate better with mass-to-area ratio (p-values < 0.001). When compared to experimental observations, a simulated two-stage air classification and size reduction scenario predicts the overall partitioning of anatomical tissues within 15% for pith, husk, rind, and cob tissues. The model predicts an air-classified fraction preferentially enriched in cob (purity = 20%), rind (purity = 74%), and pith (purity = 4.5%) with a mass yield of 47%. Empirical relations for these properties can be used to predict the partitioning of corn stover during air classification based on anatomical type and size.

09 BIOMASS FUELS↗

Detecting Arsenic Contamination Using Satellite Imagery and Machine Learning

Arsenic, a potent carcinogen and neurotoxin, affects over 200 million people globally. Current detection methods are laborious, expensive, and unscalable, being difficult to implement in developing regions and during crises such as COVID-19. This study attempts to determine if a relationship exists between soil’s hyperspectral data and arsenic concentration using NASA’s Hyperion satellite. It is the first arsenic study to use satellite-based hyperspectral data and apply a classification approach. Four regression machine learning models are tested to determine this correlation in soil with bare land cover. Raw data are converted to reflectance, problematic atmospheric influences are removed, characteristic wavelengths are selected, and four noise reduction algorithms are tested. The combination of data augmentation, Genetic Algorithm, Second Derivative Transformation, and Random Forest regression (R 2 =0.840 and normalized root mean squared error (re-scaled to [0,1]) = 0.122) shows strong correlation, performing better than past models despite using noisier satellite data (versus lab-processed samples). Three binary classification machine learning models are then applied to identify high-risk shrub-covered regions in ten U.S. states, achieving strong accuracy (=0.693) and F1-score (=0.728). Overall, these results suggest that such a methodology is practical and can provide a sustainable alternative to arsenic contamination detection.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Evaluation of artificial neural network performance for classification of potato plants infected with potato virus Y using spectral data on multiple varieties and genotypes

Potato virus Y (Potyviridae, PVY) is a plant virus that poses a significant threat to potato producers on a global basis. The pathogen has disrupted seed potato supplies and negatively impacted yield and quality of commercial potato crops. The potato industry currently manages PVY infection levels via insecticide applications, regional seed certification programs that rely on field scouting to visually assess individual plants for infection status, and destructive and costly tissue sampling coupled with laboratory assays. Despite these efforts, PVY continues to confound potato industry stakeholders resulting in economic harm. Remote sensing and machine learning provide for the development of new tools to more accurately detect and spatially quantify PVY-infected plants versus the current state of the art. However, there is a need to understand how the occurrence of many different potato varieties impact the dynamics of developing models to detect potato plants impacted with PVY and their potential effectiveness. This study evaluates classification modelling outcomes using spectral datasets collected in different temporal and spatial environments (greenhouse and a production field) on multiple potato varieties consisting of labelled instances of plants infected with PVY and those not infected with the virus. A modelling framework was developed to support iterative modelling runs using artificial neural network (ANN) architectures configured as binary classifiers to develop sample populations to support statistical analysis on model performance using specific spectral subsets. When using spectral data to detect PVY-infected plants, ANN models achieved the highest mean accuracy of 0.894 on a single variety. Conversely, the same ANN model architecture only achieved a mean accuracy of 0.575 on a spectral data set representing 29 potato breeding lines. Additionally, statistical analysis indicates spectral regions including the red edge, near infrared and shortwave infrared contain more important spectral features for the ANN classifier introduced in this research.

60 APPLIED LIFE SCIENCES↗

Evaluating Limits of Machine Learning-Assisted Raman Spectroscopy in Classification of Biological Samples

Machine learning (ML)-assisted Raman spectroscopy has become a powerful analytical tool for the classification and identification of analytes; however, technical challenges impacting its detection accuracy have not been thoroughly investigated. This study explores experimental factors affecting classification performance. Among the evaluated ML models, ML algorithms show minimal impact on classification accuracy. Instead, experimental factors, including spectral similarity between tested samples and data quality, dominate detection performance. Increases in spectral noise and spectral similarity significantly reduce classification accuracy. In well-controlled samples with low experimental noise, ML-assisted Raman spectroscopy can discriminate lipid mixtures with a composition difference of 1.85 mol %. To assess the effect of biological heterogeneity, we analyzed single-cell Raman spectra from Saccharomyces cerevisiae strains carrying single, double, or triple gene mutations. Intrinsic cell-to-cell variability introduced substantial spectral differences, severely reducing the accuracy of multiclass classification of these genetically similar strains at the single-cell level. Averaging Raman spectra across multiple cells improved classification accuracy by reducing this spectral variability. We also assess the effectiveness of transfer learning across different Raman spectrometers, specifically by applying an ML model trained on one instrument to another Raman spectrometer. Transfer learning can be improved with proper instrument calibration, highlighting the importance of instrument standardization. Overall, our results demonstrate that data quality and spectral similarity are the primary bottlenecks in ML-assisted Raman spectroscopy. Careful attention to sample preparation, data acquisition, measurement conditions, and instrument calibration is critical to achieving robust and reliable classification performance.

Fungi↗

Jet classification using high-level features from anatomy of top jets

Recent advancements in deep learning models have significantly enhanced jet classification performance by analyzing low-level features (LLFs). However, this approach often leads to less interpretable models, emphasizing the need to understand the decision-making process and to identify the high-level features (HLFs) crucial for explaining jet classification. To address this, we consider the top jet tagging problems and introduce an analysis model (AM) that analyzes selected HLFs designed to capture important features of top jets. Our AM mainly consists of the following three modules: a relation network analyzing two-point energy correlations, mathematical morphology and Minkowski functionals for generalizing jet constituent multiplicities, and a recursive neural network analyzing subjet constituent multiplicity to enhance sensitivity to subjet color charges. We demonstrate that our AM achieves performance comparable to the Particle Transformer (ParT) while requiring fewer computational resources in a comparison of top jet tagging using jets simulated at the hadronic calorimeter angular resolution scale. Furthermore, as a more constrained architecture than ParT, the AM exhibits smaller training uncertainties because of the bias-variance tradeoff. We also compare the information content of AM and ParT by decorrelating the features already learned by AM. Lastly, we briefly comment on the results of AM with finer angular resolution inputs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Adaptive Discovery and Mixed-Variable Optimization of Next Generation Synthesizable Microelectronic Materials

Design of new microelectronic materials is characterized by several challenges such as high-dimensionality of the atomic structure-composition variable space, formidable cost of directly using high-fidelity simulations for design optimization, dispersity in literature-reported similar materials and synthesis methods, complex physical mechanisms, and mixed qualitative and quantitative design variables that lead to a disjointed design space. Even though machine learning (ML) techniques have been employed to expedite materials innovation, existing methods treat ML and design optimization as two separate processes, failing to resolve the fundamental challenges associated with high dimensionality and mixed-variable complexity. We have developed a ML enhanced mixed-variable material design optimization framework to efficiently extract useful information from existing data in literature and physics-based simulations to guide the autonomous search for optimal materials. Our proposed framework is composed of four computational modules: (1) a natural language processing (NLP) based virtual screening module, (2) classification based concept exploration module, (3) a density functional theory (DFT)-based high-fidelity evaluation model, and (4) a novel latent-variable Gaussian process (LVGP) ML model for mixed-variable problems with uncertainty quantification, which seamlessly integrates with Bayesian Optimization (BO) and achieves superb efficiency through embedded physics-based dimension reduction. Our approach is demonstrated and validated using the testbed of functional materials exhibiting metal-insulation transitions (MITs), with the targeted reversible resistivity changes (∼10^5) near room temperature. At the end of the 30-month project, we have developed a series of new ML techniques using NLP, conditional variational autoencoders, active learning, latent-variable Gaussian processes, integrated with Bayesian optimization. Our project has resulted in new predicted MITs compounds and improved understanding of MITs microscopic mechanisms, which in turn will revolutionize microelectronics science to provide energy-saving solutions. Our research has improved both creativity and efficiency in transforming rare-event discoveries of new functional materials to persistent innovations. In addition to open-sourcing the online MIT database and the classification model, the LVGP open source code has been downloaded more than 15,000 times within two years. More than 40 MIT compounds have been identified and many have been pursued experimentally via collaborators. The research results are published in close to 20 collaborative papers in high-impact journals, such as Chem. Mater., Appl. Phys. Rev., Sci. Rep., among others of design space.

36 MATERIALS SCIENCE↗

Path-BigBird: An AI-Driven Transformer Approach to Classification of Cancer Pathology Reports

PURPOSE Surgical pathology reports are critical for cancer diagnosis and management. To accurately extract information about tumor characteristics from pathology reports in near real time, we explore the impact of using domain-specific transformer models that understand cancer pathology reports. METHODS We built a pathology transformer model, Path-BigBird, by using 2.7 million pathology reports from six SEER cancer registries. We then compare different variations of Path-BigBird with two less computationally intensive methods: Hierarchical Self-Attention Network (HiSAN) classification model and an offthe-shelf clinical transformer model (Clinical BigBird). We use five pathology information extraction tasks for evaluation: site, subsite, laterality, histology, and behavior. Model performance is evaluated by using macro and micro F 1 scores. RESULTS We found that Path-BigBird and Clinical BigBird outperformed the HiSAN in all tasks. Clinical BigBird performed better on the site and laterality tasks. Versions of the Path-BigBird model performed best on the two most difficult tasks: subsite (micro F 1 score of 72.53, macro F 1 score of 35.76) and histology (micro F 1 score of 80.96, macro F 1 score of 37.94). The largest performance gains over the HiSAN model were for histology, for which a Path-BigBird model increased the micro F 1 score by 1.44 points and the macro F 1 score by 3.55 points. Overall, the results suggest that a Path-BigBird model with a vocabulary derived from wellcurated and deidentified data is the best-performing model. CONCLUSION The Path-BigBird pathology transformer model improves automated information extraction from pathology reports. Although Path-BigBird outperforms Clinical BigBird and HiSAN, these less computationally expensive models still have utility when resources are constrained.

60 APPLIED LIFE SCIENCES↗

Ensemble models for circuit topology estimation, fault detection and classification in distribution systems

This paper presents a methodology for simultaneous fault detection, classification, and topology estimation for adaptive protection of distribution systems. The methodology estimates the probability of the occurrence of each one of these events by using a hybrid structure that combines three sub-systems, a convolutional neural network for topology estimation, a fault detection based on predictive residual analysis, and a standard support vector machine with probabilistic output for fault classification. The input to all these sub-systems is the local voltage and current measurements. A convolutional neural network uses these local measurements in the form of sequential data to extract features and estimate the topology conditions. The fault detector is constructed with a Bayesian stage (a multitask Gaussian process) that computes a predictive distribution (assumed to be Gaussian) of the residuals using the input. Since the distribution is known, these residuals can be transformed into a Standard distribution, whose values are then introduced into a one-class support vector machine. The structure allows using a one-class support vector machine without parameter cross-validation, so the fault detector is fully unsupervised. Finally, a support vector machine uses the input to perform the classification of the fault types. All three sub-systems can work in a parallel setup for both performance and computation efficiency. In conclusion, we test all three sub-systems included in the structure on a modified IEEE123 bus system, and we compare and evaluate the results with standard approaches.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Masked Particle Modeling on Sets: Towards Self-Supervised High Energy Physics Foundation Models

Abstract We propose masked particle modeling (MPM) as a self-supervised method for learning generic, transferable, and reusable representations on unordered sets of inputs for use in high energy physics (HEP) scientific data. This work provides a novel scheme to perform masked modeling based pre-training to learn permutation invariant functions on sets. More generally, this work provides a step towards building large foundation models for HEP that can be generically pre-trained with self-supervised learning and later fine-tuned for a variety of down-stream tasks. In MPM, particles in a set are masked and the training objective is to recover their identity, as defined by a discretized token representation of a pre-trained vector quantized variational autoencoder. We study the efficacy of the method in samples of high energy jets at collider physics experiments, including studies on the impact of discretization, permutation invariance, and ordering. We also study the fine-tuning capability of the model, showing that it can be adapted to tasks such as supervised and weakly supervised jet classification, and that the model can transfer efficiently with small fine-tuning data sets to new classes and new data domains.

Heinrich, Lukas (ORCID:0000000240487584)↗

SNM Radiation Signature Classification Using Different Semi-Supervised Machine Learning Models

The timely detection of special nuclear material (SNM) transfers between nuclear facilities is an important monitoring objective in nuclear nonproliferation. Persistent monitoring enabled by successful detection and characterization of radiological material movements could greatly enhance the nuclear nonproliferation mission in a range of applications. Supervised machine learning can be used to signal detections when material is present if a model is trained on sufficient volumes of labeled measurements. However, the nuclear monitoring data needed to train robust machine learning models can be costly to label since radiation spectra may require strict scrutiny for characterization. Therefore, this work investigates the application of semi-supervised learning to utilize both labeled and unlabeled data. As a demonstration experiment, radiation measurements from sodium iodide (NaI) detectors are provided by the Multi-Informatics for Nuclear Operating Scenarios (MINOS) venture at Oak Ridge National Laboratory (ORNL) as sample data. Anomalous measurements are identified using a method of statistical hypothesis testing. After background estimation, an energy-dependent spectroscopic analysis is used to characterize an anomaly based on its radiation signatures. In the absence of ground-truth information, a labeling heuristic provides data necessary for training and testing machine learning models. Supervised logistic regression serves as a baseline to compare three semi-supervised machine learning models: co-training, label propagation, and a convolutional neural network (CNN). In each case, the semi-supervised models outperform logistic regression, suggesting that unlabeled data can be valuable when training and demonstrating value in semi-supervised nonproliferation implementations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Deep Learning for Fish Identification from Sonar Data: CRADA 481 [Abstract only]

To help solve the challenges of hydropower energy production related to the potential for eel injury and mortality from passage through hydropower turbines, we will develop a deep learning method for identifying migrating eels from imaging sonar. This project continues with a prior project conducted by the Pacific Northwest National Laboratory (PNNL) and the Electric Power Research Institute (EPRI) in FY2018-2019. The proposed method employs Convolution Neural Network (CNN), a powerful deep learning method for image classification, to distinguish between images of eels and non-eel moving objects. We propose to collect more laboratory data and add more existing field data to train a powerful deep learning model. In addition to eels and sticks as classified in previous studies, we will add images containing several non-eel fish species and macrophyte mats to the training data. A multi-class classification model will be developed to distinguish these objects. Object detection algorithm will be explored and developed to locate and identify multiple objects in each sonar frame. Motion analysis will be performed to track the movement of objects in sonar video clips. We will also improve the data conversion algorithm so that it can read in both DIDSON and ARIS (both are imaging sonars developed by Sound Metrics Corp) data files and convert them to images with comparably high resolution, regardless of the varying detection ranges in different environments. The developed algorithms will be packaged as a software with a graphic user interface. The software will be evaluated by external collaborators in the field. The developed framework can be generalized for automatic monitoring of fish passage and migration using other imaging sonars like ARIS and will benefit the design and operation of ecologically friendly hydroelectric projects. The developed wavelet and CNN model configuration parameters can potentially be transferred to lamprey detection in similar riverine environments.

13 HYDRO ENERGY↗

Characterizing Quantum Classifier Utility in Natural Language Processing Workflows

Quantum Natural Language Processing (QNLP) develops natural language processing (NLP) models for deployment on quantum computers. We explore feature and data prototype selection techniques to address challenges posed by encoding high dimensional features. Our study builds quantum circuit classifiers that includes classical feature pre-processing, quantum embedding and quantum model training. The quantum models are built on 4 or 6 qubits and the quantum neural network (QNN) uses the established bricklayer design. We compare the dependence of model performance (in terms of accuracy and F1 scores) on feature length, embedding gates and parameterized unitary design. We compare the performance of quantum machine learning models to classical convolution neural network model (CNN) on binary and multi-class classification tasks using two datasets of synthetic features and labels. The first is the ECP-CANDLE P3B3 dataset a corpus of synthetically generated cancer pathology reports. The second dataset is extracted from well-known benchmark dataset (MADELON) - features are generated with a combination of informative, repeated and uninformative features. Both datasets are used for binary classification and multi-class classification with 3 classes. We observe robust, accurate performance from all models on the binary classification tasks, but multiclass classification is a challenge for the quantum models-there is a notable decrease in accuracy when using 3 classes. Overall the performance is comparable in terms of recall and accuracy between QNNs and CNNs, even with large datasets. These results provide a point of comparison between quantum and classical models on real-world datasets.

Hamilton, Kathleen↗

Robust Machine Learning Inference from X-ray Absorption Near Edge Spectra through Featurization

X-ray absorption spectroscopy (XAS) is a commonly employed technique for characterizing functional materials. In particular, X-ray absorption near edge spectra (XANES) encode local coordination and electronic information, and machine learning approaches to extract this information are of significant interest. To date, most ML approaches for XANES have primarily focused on using the raw spectral intensities as input, overlooking the potential benefits of incorporating spectral transformations and dimensionality reduction techniques into ML predictions. Here, in this work, we focused on systematically comparing the impact of different featurization methods on the performance of ML models for XAS analysis. We evaluated the classification and regression capabilities of these models on computed data sets and validated their performance on previously unseen experimental data sets. Our analysis revealed an intriguing discovery: the cumulative distribution function feature achieves both high prediction accuracy and exceptional transferability. This remarkably robust performance can be attributed to its tolerance to horizontal shifts in the spectra, which is crucial when validating models using experimental data. While this work exclusively focuses on XANES analysis, we anticipate that the methodology presented here will hold promise as a versatile asset to the broader spectroscopy community.

36 MATERIALS SCIENCE↗

Finding predictive models for singlet fission by machine learning

Singlet fission (SF), the conversion of one singlet exciton into two triplet excitons, could significantly enhance solar cell efficiency. Molecular crystals that undergo SF are scarce. Computational exploration may accelerate the discovery of SF materials. However, many-body perturbation theory (MBPT) calculations of the excitonic properties of molecular crystals are impractical for large-scale materials screening. We use the sure-independence-screening-and-sparsifying-operator (SISSO) machine-learning algorithm to generate computationally efficient models that can predict the MBPT thermodynamic driving force for SF for a dataset of 101 polycyclic aromatic hydrocarbons (PAH101). SISSO generates models by iteratively combining physical primary features. The best models are selected by linear regression with cross-validation. The SISSO models successfully predict the SF driving force with errors below 0.2 eV. Based on the cost, accuracy, and classification performance of SISSO models, we propose a hierarchical materials screening workflow. Three potential SF candidates are found in the PAH101 set.

36 MATERIALS SCIENCE↗

Combinatorial Evaluation of Physical Feature Engineering, Classical Machine Learning, and Deep Learning Models for Synchrophasor Data at Scale

A major objective of the project was to train and evaluate the effectiveness of multiple event and anomaly detection, identification and classification deep temporal learning models for processing of real-time phasor measurement unit (PMU) data streams. A vast dataset, consisting of two years of phasor measurements from all three U.S. Interconnections, was curated and released by the Department of Energy (DOE) through Pacific Northwest National Laboratory (PNNL). The dataset also included an event log that provided event times and types (e.g. generator trips, line trips, planned service events, transformer operations, etc.). Our analysis of this dataset addressed six (6) of the eleven (11) research priorities identified in Funding Opportunity Announcement (FOA) DE-FOA-0001861 “Big Data Analysis of Synchrophasor Data” (FOA 1861). Rather than being limited to pre-determined specific algorithms, this project relied on the uniquely structured, highly performant underlying time series database capabilities of the PredictiveGrid platform to assess the vast dataset utilizing a wide variety of algorithms.

24 POWER TRANSMISSION AND DISTRIBUTION↗