Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

OpenCRUMS USA: An Open Machine Learning Framework for Characterizing Variability in Aerosol Reanalysis Data

Advances in artificial intelligence (AI) have called for exploring how these techniques can be used for exploring patterns in large climate datasets. To that regard, the U.S. Department of Energy AI for Earth System Predictability (AI4ESP) supported a pilot initiative called the Open Classification of Regimes in the Southeast USA (OpenCRUMS USA) project to explore how AI can be used to characterize modes of spatial variability in large climate datasets. For this study, we focus on comparing two methods for characterizing the modes of spatial variability of surface aerosol concentration over the Houston region: empirical orthogonal functions (EOFs) and layerwise relevance propagation (LRP) applied to a convolutional neural network (CNN) classifier. We show that EOF analysis typically attributes spatial variability modes that span all of southeast Texas, prohibiting the attribution of spatial variability to localized regions. However, using LRP on the CNN classifier resolves the explanatory parameters at a finer spatial resolution than EOFs. This allows for the attribution of the spatial variability of surface aerosols to local regions of organic carbon which was not possible using EOFs. In addition, the LRP analysis also suggests that synoptic-scale transport of dust is most prevalent during anticyclonic and pretrough synoptic conditions as categorized by self-organizing maps.

54 ENVIRONMENTAL SCIENCES↗

Chemical classification program synthesis using generative artificial intelligence

Accurately classifying chemical structures is essential for cheminformatics and bioinformatics, including tasks such as identifying bioactive compounds of interest, screening molecules for toxicity to humans, finding non-organic compounds with desirable material properties, or organizing large chemical libraries for drug discovery or environmental monitoring. However, manual classification is labor-intensive and difficult to scale to large chemical databases. Existing automated approaches either rely on manually constructed classification rules, or are deep learning methods that lack explainability. This work presents an approach that uses generative artificial intelligence to automatically write chemical classifier programs for classes in the Chemical Entities of Biological Interest (ChEBI) database. These programs can be used for efficient deterministic run-time classification of SMILES structures, with natural language explanations. The programs themselves constitute an explainable computable ontological model of chemical class nomenclature, which we call the ChEBI Chemical Class Program Ontology (C3PO). We validated our approach against the ChEBI database, and compared our results against deep learning models and a naive SMARTS pattern based classifier. C3PO outperforms the naive classifier, but does not reach the performance of state of the art deep learning methods. However, C3PO has a number of strengths that complement deep learning methods, including explainability and reduced data dependence. C3PO can be used alongside deep learning classifiers to provide an explanation of the classification, where both methods agree. The programs can be used as part of the ontology development process, and iteratively refined by expert human curators.

Artificial Intelligence↗

The Pixel Anomaly Detection Tool : a user-friendly GUI for classifying detector frames using machine-learning approaches

Data collection at X-ray free electron lasers has particular experimental challenges, such as continuous sample delivery or the use of novel ultrafast high-dynamic-range gain-switching X-ray detectors. This can result in a multitude of data artefacts, which can be detrimental to accurately determining structure-factor amplitudes for serial crystallography or single-particle imaging experiments. Here, a new data-classification tool is reported that offers a variety of machine-learning algorithms to sort data trained either on manual data sorting by the user or by profile fitting the intensity distribution on the detector based on the experiment. This is integrated into an easy-to-use graphical user interface, specifically designed to support the detectors, file formats and software available at most X-ray free electron laser facilities. The highly modular design makes the tool easily expandable to comply with other X-ray sources and detectors, and the supervised learning approach enables even the novice user to sort data containing unwanted artefacts or perform routine data-analysis tasks such as hit finding during an experiment, without needing to write code.

47 OTHER INSTRUMENTATION↗

A knowledge-informed large language model framework for U.S. nuclear power plant shutdown initiating event classification for probabilistic risk assessment

Identifying and classifying shutdown initiating events (SDIEs) is critical for developing shutdown probabilistic risk assessment for nuclear power plants. Existing computational approaches cannot achieve satisfactory performance due to the challenges of unavailable large, labeled datasets, imbalanced event types, and label noise. To address these challenges, we propose a hybrid pipeline that integrates a knowledge-informed machine learning model to prescreen non-SDIEs and a large language model (LLM) to classify SDIEs into four types. In the prescreening stage, we proposed a set of 44 SDIE text patterns that consist of the most salient keywords and phrases from six SDIE types. Text vectorization based on the SDIE patterns generates feature vectors that are highly separable by using a simple binary classifier. The second stage builds Bidirectional Encoder Representations from Transformers (BERT)-based LLM, which learns generic English language representations from self-supervised pretraining on a large dataset and adapts to SDIE classification by fine-tuning it on an SDIE dataset. The proposed approaches are evaluated on a dataset with 10,928 events using precision, recall ratio, F 1 score, and average accuracy. In conclusion, the results demonstrate that the prescreening stage can exclude more than 97% non-SDIEs, and the LLM achieves an average accuracy of 95.1% for SDIE classification.

99 - GENERAL AND MISCELLANEOUS↗

Structural- and Functional-Informed Machine Learning for Protein Function Prediction

In this project we aimed to extend methods for protein function prediction to include structural prediction data, and benchmark methods against existing tools. We proposed to apply the method to large metagenome datasets, and develop approaches to examine activity-based protein profiling results for protein function-structure patterns. Nitrogen cycle protein families were previously identified and are used here to provide a proof-of-principle for use of structure prediction in protein function classification.

59 BASIC BIOLOGICAL SCIENCES↗

A phenology- and trend-based approach for accurate mapping of sea-level driven coastal forest retreat

The rapid replacement of upland forest by encroaching marshland is a striking manifestation of global sea-level rise (SLR). Timely and high-resolution information on the location and extent of transition forest (the ecotone between upland forest and marsh where tree mortality due to seawater intrusion begins) is fundamental to understanding the processes and patterns of SLR-driven landscape reorganization. Despite its significance, accurate characterization of salt-impacted transition forest remains challenging due to the complexity of coastal environments, scarcity of ground-truth data, and the lack of effective mapping algorithms. Here we use the full archive of Landsat images between 1984 and 2021 to investigate the spectral, temporal, and phenological characteristics of transition forest, and develop a robust framework for monitoring coastal vegetation shifts in the mid-Atlantic U.S., a global SLR hotspot. Here, we found that transition forest exhibits strong negative NDVI trends and a deviation of land surface phenology from marsh and upland forest that distinguishes itself from surrounding vegetation. By integrating temporal trends and land surface phenology, our results demonstrate superior discrimination between marsh and coastal forests to existing map products (e.g. NOAA Coastal Change Analysis Program, National Land Cover Database) that allows a reliable identification of the coastal treeline. We applied the approach to map regional land cover in 1985, 2000 and 2020 (overall classification accuracy >92%) and found that the area of coastal forest decreased by 22.0% from 1985 to 2020, the majority of which transitioned to marshland (92.3%, 5.3 × 10 3 ha). Based upon fine-scale patterns of coastal transgression, we created a practical workflow for spatially explicit quantification of forest retreat rates. Concurrent with rising sea level, coastal forests migrated upslope from 0.63 (± 0.27) m above sea level in 1985 to 0.78 (± 0.32) m above sea level in 2020, and horizontal forest retreat rates accelerated from 3.1 (range of 0–36) m yr -1 during 1985–2000 to 4.7 (0–55) m yr -1 during 2001–2020. As SLR continues to accelerate, our study may serve as a scalable solution for consistent tracking of coastal landscape evolution that is urgently needed for sustainable forest and wetland management.

54 ENVIRONMENTAL SCIENCES↗

Quantum Transfer Learning to Boost Dementia Detection

Dementia is a devastating condition with profound implications for individuals, families, and healthcare systems. Early and accurate detection of dementia is critical for timely intervention and improved patient outcomes. While classical machine learning and deep learning approaches have been explored extensively for dementia prediction, these solutions often struggle with high-dimensional biomedical data and large-scale datasets, quickly reaching computational and performance limitations. To address this challenge, quantum machine learning (QML) has emerged as a promising paradigm, offering faster training and advanced pattern recognition capabilities. This work aims to demonstrate the potential of quantum transfer learning (QTL) to enhance the performance of a weak classical deep learning model applied to a binary classification task for dementia detection. Besides, we show the effect of noise on the QTL-based approach, investigating the reliability and robustness of this method. Using the OASIS 2 dataset, we show how quantum techniques can transform a suboptimal classical model into a more effective solution for biomedical image classification, highlighting their potential impact on advancing healthcare technology.

Bhowmik, Sounak [University of Tennessee, Knoxvill↗

Machine Learning-Based Classification of Lignocellulosic Biomass from Pyrolysis-Molecular Beam Mass Spectrometry Data

High-throughput analysis of biomass is necessary to ensure consistent and uniform feedstocks for agricultural and bioenergy applications and is needed to inform genomics and systems biology models. Pyrolysis followed by mass spectrometry such as molecular beam mass spectrometry (py-MBMS) analyses are becoming increasingly popular for the rapid analysis of biomass cell wall composition and typically require the use of different data analysis tools depending on the need and application. Here, the authors report the py-MBMS analysis of several types of lignocellulosic biomass to gain an understanding of spectral patterns and variation with associated biomass composition and use machine learning approaches to classify, differentiate, and predict biomass types on the basis of py-MBMS spectra. Py-MBMS spectra were also corrected for instrumental variance using generalized linear modeling (GLM) based on the use of select ions relative abundances as spike-in controls. Machine learning classification algorithms e.g., random forest, k-nearest neighbor, decision tree, Gaussian Naïve Bayes, gradient boosting, and multilayer perceptron classifiers were used. The k-nearest neighbors (k-NN) classifier generally performed the best for classifications using raw spectral data, and the decision tree classifier performed the worst. After normalization of spectra to account for instrumental variance, all the classifiers had comparable and generally acceptable performance for predicting the biomass types, although the k-NN and decision tree classifiers were not as accurate for prediction of specific sample types. Gaussian Naïve Bayes (GNB) and extreme gradient boosting (XGB) classifiers performed better than the k-NN and the decision tree classifiers for the prediction of biomass mixtures. The data analysis workflow reported here could be applied and extended for comparison of biomass samples of varying types, species, phenotypes, and/or genotypes or subjected to different treatments, environments, etc. to further elucidate the sources of spectral variance, patterns, and to infer compositional information based on spectral analysis, particularly for analysis of data without a priori knowledge of the feedstock composition or identity.

59 BASIC BIOLOGICAL SCIENCES↗

Wintertime synoptic patterns of midlatitude boundary layer clouds over the western North Atlantic: Climatology and insights from in-situ ACTIVATE observations

The winter synoptic evolution of the western North Atlantic and its influence on the atmospheric boundary layer is described by means of a regime classification based on Self Organizing Maps applied to 12 year of data (2009-2020). The regimes are classified into categories according to daily 600-hPa geopotential height: dominant ridge, trough to ridge eastward transition (trough-ridge), dominant trough, and ridge to trough eastward transition (ridge-trough). A fifth synoptic regime resembles the winter climatological mean. Coherent changes in sea-level pressure and large-scale winds are in concert with the synoptic regimes: 1) the ridge regime is associated with a well-developed anticyclone; 2) the trough-ridge gives rise to a low pressure center over the ocean, ascents, and northerly winds over the coastal zone; 3) trough is associated with the eastward displacement of a cyclone, coastal subsidence, and northerly winds, all representative characteristics of cold-air outbreaks; 4) the ridge-trough regime features the development of an anticyclone and weak coastal winds. Low clouds are characteristic of the trough regime, with both trough and trough-ridge featuring synoptic maxima in cloud droplet number concentration (N d ). The N d increase is primarily observed near the coast, concomitant with strong surface heat fluxes exceeding by more than 400 W m -2 compared to fluxes further east. Five consecutive days of aircraft observations collected during the ACTIVATE campaign corroborates the climatological characterization, confirming the occurrence of high N d for days identified as trough. This study emphasizes the role of boundary-layer dynamics and aerosol activation and their roles in modulating cloud microphysics.

54 ENVIRONMENTAL SCIENCES↗

Comparison of Expert Vocabulary Usage Patterns Between Mental Health and Nonmental Health Clinicians When Diagnosing Pediatric Anxiety Disorders

Objective: To compare the utilization patterns of expert vocabulary (EVo) in diagnosing pediatric anxiety between mental health and non-mental health clinical notes from electronic health records to understand the role of Evo in informing classification and decision-making in anxiety diagnoses. Study design: We conducted a retrospective study using a cohort less than age 25 from Cincinnati Children's Hospital including 897 685 patients with 61 586 446 notes. We analyzed EVo, collected from mental health clinicians, in both mental and nonmental health notes. We compared classification accuracy using EVo-based patient-level embedding from all clinical notes, mental-health notes, and nonmental health notes for 2 tasks: 1) pre-vs postdiagnosis anxiety patients, and 2) prediagnosis anxiety vs nonanxiety patients. Results: EVo usage was highest in prediagnosis anxiety, lower in nonanxiety, and lowest in post-diagnosis. Classification models using EVo features from all, mental-health, and non-mental health notes showed similar F1 scores for prediagnosis anxiety (0.70 ± 0.2 for 2 categories). For anxiety vs nonanxiety classification, all clinical and nonmental health notes had better F1 scores than mental-health notes (above 0.90 for 3 categories). There was a notable difference in class-wise performance across both tasks. Conclusions: There are significant differences in anxiety EVo use between mental health and nonmental health clinicians. Despite less anxiety-specific terminology, non-mental health notes still captured key aspects of patient presentations, emphasizing the importance of including all clinicians' notes in analysis. EVo's utility for anxiety classification is most effective in prediagnostic phases, suggesting the need for a dedicated diagnostic lexicon and further study before incorporating EVo into classification models.

feature engineering↗

Warm‐Season Afternoon Precipitation Peak in the Central Bay of Bengal: Process‐Oriented Diagnostics

Abstract Past studies have indicated that precipitation over tropical open oceans generally peaks in the early morning. However, an intriguing departure from this pattern is observed in the central Bay of Bengal (CBoB), where rainfall exhibits a distinct afternoon peak during the South Asian summer monsoon season. By using a novel satellite‐based cloud classification and tracking data set, we found that more than 75% of the afternoon rainfall (15–17 LST) over the CBoB comes from mesoscale convective systems (MCSs). Most of the MCSs contributing to the CBoB afternoon rainfall peak originate either locally over the CBoB or near the west and east coasts of the BoB, in contrast to the northern BoB as highlighted in previous studies. Analyses show that MCSs initiated near coastlines are primarily influenced by land‐sea breezes, whereas MCSs initiated over the BoB open ocean during early morning are strongly associated with diurnal radiative forcings. In addition, there are clear diurnal propagating MCS initiation signals from the west and north coastlines of the BoB to the CBoB, which are related to diurnal gravity waves emitted from the coastlines. The thermodynamic conditions conducive to MCS initiation over different sub‐regions of the BoB are also investigated. No systematic differences found in environmental convective available potential energy between days with and without MCS initiation. However, over most sub‐regions, days with MCS initiation generally have higher total column water vapor than days without MCS initiation. This difference suggests that the lower‐free‐tropospheric moisture content plays an important role in MCS initiation over the BoB.

54 ENVIRONMENTAL SCIENCES↗

Improving qubit readout with hidden Markov models

We demonstrate the application of pattern recognition algorithms via hidden Markov models (HMM) for qubit readout. This scheme provides a state-path trajectory approach capable of detecting qubit-state transitions and makes for a robust classification scheme with higher starting-state assignment fidelity than when compared to a multivariate Gaussian or a support vector machine scheme. Therefore, the method also eliminates the qubit-dependent readout time optimization requirement in current schemes. Using a HMM state discriminator we estimate fidelities reaching the ideal limit. Unsupervised learning gives access to transition matrix, priors, and IQ distributions, providing a toolbox for studying qubit-state dynamics during strong projective readout.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Power quality disturbances diagnosis: A 2D densely connected convolutional network framework

The fast and accurate diagnosis of power quality disturbances (PQD) aids in avoiding shutdowns and unnecessary procedures, concerning electric energy distribution systems. As such, a number of techniques have been tested and applied in order to reach this objective. Majority of the techniques applied are two-step based. On the first step, power quality disturbances features are extracted. Second step, considering features extracted, disturbance classification is implemented. Recently, relevant literature has presented data-driven signal processing-based approaches, as deep convolutional neural networks (DCNN), which can implement both processing steps while providing automated recognition of patterns and outliers in data. However, not considered by state-of-art, power quality disturbances are evolving in nature, while all possible regularities might not be represented in the dataset. In this work a 2 Dimension Densely Connected Convolutional Network (2D-DenseNet) framework is presented. Further, a case study with synthetic disturbance events are analyzed. Easy-to-implement formulation, built on the 2D-DenseNet, without hard-to-design parameters, highlight potential aspects for real-life implementation.

42 ENGINEERING↗

Unsupervised Power System Event Detection and Classification Using Unlabeled PMU Data

This paper proposes a novel data-driven power system event detection and classification method based on 5TB of actual PMU measurements collected from the US western interconnect. Firstly, a set of comprehensive power quality rules are proposed to pre-filter the raw data and extract the regions of interest (ROI). Six distinct event categories are defined and corresponding patterns are chosen as references. Meanwhile, detailed characteristics of patterns are summarized to enhance our understanding of the actual events. Then, the time-independent feature vectors are generated by extracting the statistical, temporal, and spectral features from the raw time-series data. Furthermore, an ensemble model is proposed to cluster the events by combining multiple K-means clustering models using a voting strategy. Besides, both system-level and PMU-level clustering models are developed. The accuracy and robustness of the event detection method are further improved through interactive evaluation of the two-level clustering results. This paper summarizes the actual characteristics of each event category and provides a reliable basis for accurate label generation. The experiments demonstrate the effectiveness of the proposed event detection and classification method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Unsupervised Power System Event Detection and Classification Using Unlabeled PMU Data

This paper proposes a novel data-driven power system event detection and classification method based on 5TB of actual PMU measurements collected from the US western interconnect. Firstly, a set of comprehensive power quality rules are proposed to pre-filter the raw data and extract the regions of interest (ROI). Six distinct event categories are defined and corresponding patterns are chosen as references. Meanwhile, detailed characteristics of patterns are summarized to enhance our understanding of the actual events. Then, the time-independent feature vectors are generated by extracting the statistical, temporal, and spectral features from the raw time-series data. Furthermore, an ensemble model is proposed to cluster the events by combining multiple K-means clustering models using a voting strategy. Besides, both system-level and PMU-level clustering models are developed. The accuracy and robustness of the event detection method are further improved through interactive evaluation of the two-level clustering results. This paper summarizes the actual characteristics of each event category and provides a reliable basis for accurate label generation. The experiments demonstrate the effectiveness of the proposed event detection and classification method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Siamese CNN + KNN-Based Classification Framework for Non-intrusive Load Monitoring

Through the development of smart grids, programs such as demand side response, have been presented as auxiliary services to the real-time operation of distributed networks. In order to provide consumers information on their energy consumption, so that a modulation in consumption is possible, non-intrusive load monitoring has been introduced as an solution to this pattern recognition problem. Non-intrusive load monitoring enables the modeling of electrical loads connected to the low-voltage system, considering only a single measurement point. Presented state-of-the-art solutions though, consider availability of data as well as representation of all possible classes of the environment. This is of course a most conservative hypothesis, since in real-life applications availability of such data is much difficult, as well as the dynamic behavior of models is implicitly evolving in time. Here, a framework that uses neural Siamese networks with k-nearest neighbor clustering is presented toward non-intrusive load monitoring. Online learning feature is implemented, which relaxes the hypothesis of data requirements as well addresses the evolving nature of load profile. k-nearest clustering allows nonlinear characteristic space modelling. Test results using synthetics and real-life data show that the solution, besides obtaining a good generalizability in the classification, also obtained results with an accuracy of 95.77%.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Predicting transcriptional responses to cold stress across plant species

Although genome-sequence assemblies are available for a growing number of plant species, gene-expression responses to stimuli have been cataloged for only a subset of these species. Many genes show altered transcription patterns in response to abiotic stresses. However, orthologous genes in related species often exhibit different responses to a given stress. Accordingly, data on the regulation of gene expression in one species are not reliable predictors of orthologous gene responses in a related species. Here, we trained a supervised classification model to identify genes that transcriptionally respond to cold stress. A model trained with only features calculated directly from genome assemblies exhibited only modest decreases in performance relative to models trained by using genomic, chromatin, and evolution/diversity features. Models trained with data from one species successfully predicted which genes would respond to cold stress in other related species. Cross-species predictions remained accurate when training was performed in cold-sensitive species and predictions were performed in cold-tolerant species and vice versa. Models trained with data on gene expression in multiple species provided at least equivalent performance to models trained and tested in a single species and outperformed single-species models in cross-species prediction. These results suggest that classifiers trained on stress data from well-studied species may suffice for predicting gene-expression patterns in related, less-studied species with sequenced genomes.

54 ENVIRONMENTAL SCIENCES↗

MetaPoL: Immersive VR based Indoor Patterns of Life (PoL) and Anomalies Data Generation for Insider Threat Modeling in Nuclear Security

Insider threats are perhaps the most serious challenges that nuclear and radiological security systems face. Insiders pose such a great threat due to their access, authority, and knowledge, granting them opportunities to bypass dedicated nuclear and radiological security elements. For example, in one of the latest major insider threat incidents to nuclear security, the Doel-4 nuclear powerplant in Belgium suffered a shutdown, the threat of nuclear materials diversion, and long-term loss of tens of millions of dollars. Seven years of investigation concluded that it was an inside job and attempted sabotage. In this regard, there is an immediate need for R&D and technology integration in the domain of modeling indoor Patterns-of-Life (PoL) and anomaly detection. This can be achieved by using datasets of facility users’ mobility and activity, which can support the design of algorithms for insider threat modeling and detection. However, due to classification, privacy, sensitivity, and safety protocols, such datasets from real physical nuclear reactor facilities are not only hard to share, but also not always feasible to deploy and collect. Aiming to find an alternate solution, our proposed demonstration work - MetaPoL, is the first-ever (for the application space) immersive VR (virtual reality) environment of a real-world secure facility and allows users to move-and-stay through the designed indoor physical layout and also encounter NPCs (non-player characters) that emulate other facility users. In the MetaPoL an interactive user performs realistic spatio-temporal movement, dwelling and activities using a Meta Quest Pro VR headset, and that generates high-frequency (in time) high-resolution (in space) indoor spatial-temporal datasets that are valuable for PoL modeling and anomaly detection research specifically for insider threat modeling and detection mission. Such generated realistic, rich in context, and mission specific datasets can boost AI/Machine Learning based research for modeling and detecting insider threats in nuclear security and nonproliferation.

Gunaratne, Chathika↗