Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

SpeckleNN: a unified embedding for real-time speckle pattern classification in X-ray single-particle imaging with limited labeled examples

With X-ray free-electron lasers (XFELs), it is possible to determine the three-dimensional structure of noncrystalline nanoscale particles using X-ray single-particle imaging (SPI) techniques at room temperature. Classifying SPI scattering patterns, or `speckles', to extract single-hits that are needed for real-time vetoing and three-dimensional reconstruction poses a challenge for high-data-rate facilities like the European XFEL and LCLS-II-HE. Here, we introduce SpeckleNN, a unified embedding model for real-time speckle pattern classification with limited labeled examples that can scale linearly with dataset size. Trained with twin neural networks, SpeckleNN maps speckle patterns to a unified embedding vector space, where similarity is measured by Euclidean distance. We highlight its few-shot classification capability on new never-seen samples and its robust performance despite having only tens of labels per classification category even in the presence of substantial missing detector areas. Without the need for excessive manual labeling or even a full detector image, our classification method offers a great solution for real-time high-throughput SPI experiments.

47 OTHER INSTRUMENTATION↗

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (V.2.0)

Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.

97 MATHEMATICS AND COMPUTING↗

Development of gas chromatographic pattern recognition and classification tools for compliance and forensic analyses of fuels: A review

Gas chromatography (GC) is undoubtedly the analytical technique of choice for analyzing the composition of petroleum-based fuels. Over the past twenty years, as comprehensive two-dimensional gas chromatography (GC×GC) has evolved, fuel analysis has often been highlighted in scientific reports, as their complexity allows for illustration of the impressive peak capacity gains afforded by GC×GC. Indeed, several research groups in recent years have applied GC×GC and chemometrics to demonstrate the potential of these analytical tools to address important compliance (tax evasion, tax credits, physical quality standards) and forensic (arson investigations, oil spills) applications involving fuels. None the less, routine use of GC×GC in forensic laboratories has been limited largely by (1) legal and regulatory guidelines, (2) lack of chemometrics training, and (3) concerns about the reproducibility of GC×GC. The goal of this review is to highlight recent advances in one-dimensional GC (1D-GC) and GC×GC analyses of fuels for compliance and forensic applications, in an effort to assist scientists in overcoming the aforementioned hindrances. An introduction to 1D-GC principles, GC×GC technology (column stationary phases and modulators) and several chemometric techniques will be provided. More specifically, chemometric techniques will be broken down into (1) signal pre-processing, (2) peak decomposition, identification and quantification, and (3) classification and pattern recognition. Examples of compliance and forensic applications will be discussed, with particular emphasis on the demonstrated success of the employed chemometric techniques. Overall, this review will hopefully make 1D-GC and GC×GC coupled with chemometric data analysis tools more accessible to the larger scientific community, and aid in eventual widespread standardization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

47 OTHER INSTRUMENTATION↗

Data-Driven Clustering and Classification of Outage Patterns with Insights into their Links to Extreme Events

At a global level extreme events have increased in both scale and impact. These events have the potential to affect the electrical grid infrastructure and cause a wide range of outages, which can lead to a disruption in daily patterns, cost millions of dollars and also the loss of life. Currently, to track these outage events there have been various approaches developed ranging from regional to national level quantifications for what defines an outage. However, this variation in methods can potentially lead to subjective decision-making and a lack of proper management in relation to the event. While previous work has made strides in determining spatio-temporal patterns, minimal attention has been given to the type and number of outages an area may be exposed to. The differences in incurred cost and the overall severity of an event between a transformer box malfunction and a hurricane are drastic, and by finding historical signals, we can allow for more efficient management, potentially saving lives and millions of dollars. Here, we leverage unsupervised machine learning techniques to delineate outage patterns among 22 counties within the United States and find that there are clear, segregated clusters (0.93 silhouette) of data which are related by event behavior and underlying cause. This finding will allow for energy stakeholders, policy makers, and researchers to gain a deeper understanding of the extent and severity of historic events and to better prepare for electrical grid infrastructure planning and management.

Koob, Benjamin [ORNL]↗

Deep learning model to detect various synchrophasor data anomalies

High-density synchrophasors provide valuable information for power grid situational awareness, operation and control. Unfortunately, due to factors including communication instability and hardware failure, their data quality can be greatly deteriorated by anomalies. Since the anomalies can impact the performance of the synchrophasor applications, it is of paramount significance to propose a model to detect anomalies in synchrophasor. In this study, a convolutional neural network model is established to detect and classify the anomalies in the synchrophasor measurements. Additionally, four types of anomalies observed in actual synchrophasors including erroneous patterns, random spikes, missing points and high-frequency interferences are considered in this study. The proposed model is extensively evaluated via field-collected measurements from the synchrophasor network in Jiangsu grid, China. The superior performance of the proposed model indicates the great potential of using deep learning for the detection of abnormal synchrophasor measurements.

42 ENGINEERING↗

INTERSECT Architecture Specification: Use Case Design Patterns (V.0.9)

Connecting scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing (HPC) center enables autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence (AI)-driven design, discovery and evaluation. The Self-driven Experiments for Science / Interconnected Science Ecosystem (INTERSECT) Open Architecture enables science breakthroughs using intelligent networked systems, instruments and facilities with a federated hardware/software architecture for the laboratory of the future. It relies on a novel approach, consisting of (1) science use case design patterns, (2) a system of systems architecture, and (3) a microservice architecture. This document introduces the science use case design patterns of the INTERSECT Architecture. It describes the overall background, the involved terminology and concepts, and the pattern format and classification. It further details the 12 defined patterns and provides insight into building solutions from these patterns. The document also describes the application of these patterns in the context of several INTERSECT autonomous laboratories. The target audience are computer, computational, instrument and domain science experts working in the field of autonomous experiments.

97 MATHEMATICS AND COMPUTING↗

Science Use Case Design Patterns for Autonomous Experiments

Connecting scientific instruments and robot-controlled laboratories with computing and data resources at the edge, the Cloud or the high-performance computing (HPC) center enables autonomous experiments, self-driving laboratories, smart manufacturing, and artificial intelligence (AI)-driven design, discovery and evaluation. The Self-driven Experiments for Science / Interconnected Science Ecosystem (INTERSECT) Open Architecture enables science breakthroughs using intelligent networked systems, instruments and facilities with a federated hardware/software architecture for the laboratory of the future. It relies on a novel approach, consisting of (1) science use case design patterns, (2) a system of systems architecture, and (3) a microservice architecture. This paper introduces the science use case design patterns of the INTERSECT Architecture. It describes the overall background, the involved terminology and concepts, and the pattern format and classification. It further offers an overview of the 12 defined patterns and 4 examples of patterns of 2 different pattern classes. It also provides insight into building solutions from these patterns. The target audience are computer, computational, instrument and domain science experts working in the field of autonomous experiments.

Engelmann, Christian↗

The agglomeration and dispersion dichotomy of human settlements on Earth

Human settlements on Earth are scattered in a multitude of shapes, sizes and spatial arrangements. These patterns are often not random but a result of complex geographical, cultural, economic and historical processes that have profound human and ecological impacts. However, little is known about the global distribution of these patterns and the spatial forces that creates them. This study analyses human settlements from high-resolution satellite imagery and provides a global classification of spatial patterns. We find two emerging classes, namely agglomeration and dispersion. In the former, settlements are fewer than expected based on the predictions of scaling theory, while an unexpectedly high number of settlements characterizes the latter. To explain the observed spatial patterns, we propose a model that combines two agglomeration forces and simulates human settlements’ historical growth. Our results show that our model accurately matches the observed global classification (F1: 0.73), helps to understand and estimate the growth of human settlements and, in turn, the distribution and physical dynamics of all human settlements on Earth, from small villages to cities.

54 ENVIRONMENTAL SCIENCES↗

Particle track classification using quantum associative memory

Pattern recognition algorithms are commonly employed to simplify the challenging and necessary step of track reconstruction in sub-atomic physics experiments. Aiding in the discrimination of relevant interactions, pattern recognition seeks to accelerate track reconstruction by isolating signals of interest. In high collision rate experiments, such algorithms can be particularly crucial for determining whether to retain or discard information from a given interaction even before the data is transferred to tape. As data rates, detector resolution, noise, and inefficiencies increase, pattern recognition becomes more computationally challenging, motivating the development of higher efficiency algorithms and techniques. Quantum associative memory is an approach that seeks to exploits quantum mechanical phenomena to gain advantage in learning capacity, or the number of patterns that can be stored and accurately recalled. Here, we study quantum associative memory based on quantum annealing and apply it to the particle track classification. We focus on discrimination models based on Ising formulations of quantum associative memory model (QAMM) recall and quantum content-addressable memory (QCAM) recall. We characterize classification performance of these approaches as a function detector resolution, pattern library size, and detector inefficiencies, using the D-Wave 2000Q processor as a testbed. Discrimination criteria is set using both solution-state energy and classification labels embedded in solution states. We find that energy-based QAMM classification performs well in regimes of small pattern density and low detector inefficiency. In contrast, state-based QCAM achieves reasonably high accuracy recall for large pattern density and the greatest recall accuracy robustness to a variety of detector noise sources.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Retrieval Augmented Generation for Robust Cyber Defense

In cybersecurity, the ability to efficiently analyze and respond to vulnerabilities, weaknesses, attack patterns, and threat tactics is critical for effective defense strategies. With the increasing complexity and volume of cybersecurity data, traditional methods of querying and retrieving information are often inadequate. To address this challenge, we implemented Retrieval-Augmented Generation (RAG) systems—CyRAG and GraphCyRAG—that integrate large language models (LLMs) with both structured data from relational databases and knowledge graphs such as Neo4j. CyRAG is designed to handle structured data, focusing on CVE (Common Vulnerabilities and Exposures) and CWE (Common Weakness Enumeration) entities to generate accurate and context-rich responses. In contrast, GraphCyRAG leverages Neo4j knowledge graphs to retrieve interconnected information from CVE, CWE, CAPEC (Common Attack Pattern Enumeration and Classification), and ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) datasets. By utilizing Neo4j’s graph-based framework, GraphCyRAG enables deeper traversal of relationships between vulnerabilities and attack patterns, providing cybersecurity analysts with more comprehensive insights into potential attack vectors and mitigation strategies. Our preliminary results demonstrate that integrating knowledge graphs with RAG significantly enhances both the accuracy and depth of threat analysis, allowing for the retrieval of dynamic, real-time data and the generation of contextually aware responses. This approach helps analysts uncover hidden relationships between cyber entities, predict exploit paths, and prioritize mitigation efforts effectively. The integration of RAG with cybersecurity knowledge graphs represents a significant advancement in cybersecurity threat intelligence, enabling more informed decision-making and stronger defense strategies.

97 MATHEMATICS AND COMPUTING↗

An advanced workflow for single-particle imaging with the limited data at an X-ray free-electron laser

An improved analysis for single-particle imaging (SPI) experiments, using the limited data, is presented here. Results are based on a study of bacteriophage PR772 performed at the Atomic, Molecular and Optical Science instrument at the Linac Coherent Light Source as part of the SPI initiative. Existing methods were modified to cope with the shortcomings of the experimental data: inaccessibility of information from half of the detector and a small fraction of single hits. The general SPI analysis workflow was upgraded with the expectation-maximization based classification of diffraction patterns and mode decomposition on the final virus-structure determination step. The presented processing pipeline allowed us to determine the 3D structure of bacteriophage PR772 without symmetry constraints with a spatial resolution of 6.9 nm. The obtained resolution was limited by the scattering intensity during the experiment and the relatively small number of single hits.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Radio Frequency Spectrum Audit to Inventory Private Cellular Base Station Infrastructure

The ever-changing cellular communication landscape makes it difficult to identify, map, and localize cellular base stations. Localizing cellular base stations provides various advantages, including information security, cybersecurity, spectrum management, and interference detection. For example, the MITRE ATT&CK® (Adversarial Tactics, Techniques, and Common Knowledge architecture) [1] and Common Attack Pattern Enumeration and Classification [2] emphasize the importance of being able to minimize the cyber security threat presented by unregulated private cellular base stations (PCBS). The majority of published research looks at the malicious use of PCBSs and focuses on using data retrieved from user equipment (UE), data obtained from an application on the UE, or data shared between the UE and a mobile network to locate it. This innovative strategy, however, focuses on the passively discovered uniqueness of radio frequency (RF) transmissions from commercial cellular infrastructure received in a designated monitoring position (DMP).

42 ENGINEERING↗

Particle Track Classification Using Quantum Associative Memory (Final Technical Report)

This project explored the use of quantum-assisted algorithms for pattern matching in sub-atomic physics experiments. Pattern matching algorithms are commonly employed to prune data of random noise and to help discriminate between signals generated by particle tracks of interest and signals generated by background events. The quantum-assisted algorithms explored in this project were based on an Ising formulation of quantum associative model (QAMM) recall and quantum content-addressable memory (QCAM) recall. The recall is performed by comparing a probe pattern with those stored in a library of patterns encoded in the QAMM/QCAM model. The classification accuracy of QAMM and QCAM recall was determined as a function of detector resolution, noise, and efficiency and pattern density, where pattern density is defined as the ratio of the number of reference signal patterns encoded in the library to each pattern’s length. We found that QAMM achieved high classification accuracy when applied to datasets with low pattern density. QCAM achieved high classification accuracy for datasets with high pattern density and was found to be more robust to detector noise. The project methodology and results are described in detail in our arXiv preprint (arXiv:2011.11848) . This project was conducted by scientists at the Johns Hopkins University Applied Physics Laboratory and Oak Ridge National Laboratory from August 2018 to August 2020 and was supported by DOE grant DE-SC0019497.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Classification of River Catchments in the Contiguous United States: Code, Dataset, Similarity Patterns, and Resulting Classes

This dataset serves as supplementary information for the paper by Ciulla F. and Varadharajan C. A Network Approach for Multiscale Catchment Classification using Traits (see reference 1). It contains environmental and physical catchment traits, such as temperatures, precipitation, land use and human interference, from 9067 sites across the contiguous United States (CONUS). The purpose of this dataset is to provide information for a better trait-based categorization of river catchments in the CONUS using networks as an analytical tool. The traits variables match the ones present in the GAGES-II dataset and the preprocessing steps are described in the Methods section (processed_dataset.csv). Additionally we include the topologies (nodes, edges and clusters, also referred as classes) of the catchment network and traits network generated by said dataset (csv and json files). A series of tables support the information carried by the network providing more detailed descriptions of cluster components (SI1.pdf). A summary of all the plots of clusters of catchments with at least 50 nodes is provided (SI2.pdf). The characteristic traits for each cluster of catchments is presented as z-score (traits_categories_zscores_per_catchment_class.csv). The link to the hydrological behavior of clusters of catchments is displayed by boxplots, each describing a particular river discharge index (SI3.pdf). Both csv and json files can be read by common text editors but the data contained into them can be better handled using programming languages like python and database oriented libraries like pandas. Pdf files can be read by any pdf reader software.[02-23-2024] Update: The code and datasets necessary to reproduce the results of the study are available as a zipped repository (code_datasets_catchments_similarity.zip).

54 ENVIRONMENTAL SCIENCES↗

Interpretable Models for Workflow Differentiation in High-Performance Scientific Networks

Scientific workflows in high-performance networks spawn hundreds of interdependent flows that must be managed collectively—yet existing network classifiers treat each flow in isolation, leading to fragmented QoS decisions and missed interflow patterns. We present a novel traffic classification solution that operates at the workflow level, distinguishing entire filetransfer operations from streaming analytics by capturing how concurrent flows interact and burst together. We introduce a workflow identification window (WIW) that ingests raw packet headers from parallel flows into unified tensors, preserving the spatial-temporal patterns that differentiate scientific workflows. This approach achieves 98.7% accuracy using CNN, LSTM, and hybrid architectures, while maintaining 84% accuracy on production traffic collected a week later—demonstrating robustness to temporal drift. By integrating SHAP and GradCAM explainability, we reveal that early-packet timing patterns and cross-flow correlations drive classification decisions, providing operators with interpretable insights. Our system enables coherent workflow-level QoS enforcement and dynamic bandwidth allocation in scientific networks, eliminating manual per-flow configuration while maintaining classification latency at millisecond level.

Giannakou, Anna [LBL, Berkeley]↗