Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data segmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Automated detection of photovoltaic cleaning events: A performance comparison of techniques as applied to a broad set of labeled photovoltaic data sets

Extracting accurate soiling loss information from photovoltaic (PV) production data first requires segmenting the time series data per natural or manually occurring cleaning events. Maintenance logs are often incomplete, rain data are often unavailable, and the debate on rain thresholds for cleaning and dew or wind cleanings is still ongoing. The present work aims to overtake these issues by improving automated methods to detect these cleaning events and therefore improve extraction of soiling loss information. Time series power production data from 22 PV inverters were labeled for natural or manually occurring cleaning events. The data sets were carefully selected to include varying degrees of soiling, cleaning events, and noise. Several algorithms, including filtering logic and change point detection, were examined for efficacy at detecting the labeled cleanings. All the methods introduced except for changepoint detection showed significant improvement at detecting the labeled cleaning events per the mean F 1 score. Furthermore, the highest performing cleaning detection algorithm achieved an absolute increase in the mean F 1 score of 43% over the default version of the RdTools stochastic rate and recovery (SRR) algorithm. The highest performing algorithm included irradiance filtering and a cleaning detection threshold, adjusted based on the 40-day centered rolling median of the absolute day-to-day deviations in the daily performance index (PI). Furthermore, these improvements are promising as cleaning detection is an essential step in the automated analysis of PV soiling.

14 SOLAR ENERGY↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improved Data Interpretation through Identification of Time Series Periodicity Changes

Analysis and interpretation of time series data is easiest when the data values occur at uniform intervals in time, but actual data may have differing data sampling frequencies, such as monthly and daily readings. Applying data analysis techniques, such as smoothing, to such a data set may not give a representative result between time segments. The ability to automatically distinguish time segments of differing data frequency would provide a means for applying data analysis independently to each segment, though a suitable blending at segment boundaries would be required. A method for detecting frequency changes was developed and applied to Gaussian and median smoothing of hydraulic head data from groundwater wells at the U.S. Department of Energy Hanford Site in southeastern Washington state. The process identifies time segments of high-frequency (daily) or low-frequency (greater than daily) data using adjusted-bandwidth Gaussian kernel density estimation and a threshold value, which are further refined to address small blocks of low-frequency data within larger blocks of high-frequency data. User-selectable levels of smoothing are then applied independently to the time segments prior to combining the segment results for a single smoothed data set. This time segment identification approach provides effective low- and high-frequency data separation, which provides a method to apply data analysis independently to each time segment.

97 MATHEMATICS AND COMPUTING↗

Statistical Performance of Forced Oscillation Detectors in the Presence of Missing Measurements

In bulk power systems, measurement-based monitoring for large oscillations can help maintain system reliability. One of the challenges encountered in a recent field demonstration was the unavailability of measurements due to underlying measurement quality or communication problems. During the demonstration, the oscillation detector ignored a measurement location if even 10 seconds of data was missing. To extend the detector's ability to operate in these conditions, this paper evaluates the impact of three methods for addressing missing data. The strengths and weaknesses of each approach are evaluated using theoretical expressions for the probability of detection along with results from simulated data and publicly available field measurements. Based on these results, a suitable approach is identified that can extend the oscillation detector's performance when large segments of data are missing.

Follum, James D.↗

SOC Microstructural Analyzer

This program was designed to analyze the 3-phase microstructure of the electrodes of a solid oxide fuel cell (SOFC) or electrolysis cell (SOEC), both referred to in combination as a solid oxide cell (SOC). It is agnostic to the exact system, so it could be repurposed to analyze any 3-phase microstructure. This tool directly analyzes segmented voxel-based data that has been segmented into phase IDs (1,2,3). The voxels will be analyzed directly for: - tortuosity factors - triple phase boundaries - 2-phase interfacial areas, using a meshed isosurface - mean diameters of each phase, using an inscribed sphere method - standard deviation of the diameters of each phase, from the same inscribed sphere data - connectivity information Comprehensive information is available in the readme file (within the zipped repository in Markdown language, and also available here as a rendered PDF). Please cite this page / DOI, as well as https://doi.org/10.1111/jace.14775, for usage.

3D microstructure↗

Improving microstructures segmentation via pretraining with synthetic data

Image analysis of material microstructures through microscopy is an integral capability in the field of materials science. The topological and chemical information obtained through microscopy allow us to draw vital connections between material microstructures, properties, and processing. While scanning electron microscopy (SEM) is able to yield a considerable wealth of information interpretable by the intuition of experts, there has been considerable interest in using machine learning, convolutional neural networks (CNNs) in particular, for such image analysis task. Training CNNs for an image analysis task requires a large annotated dataset. However, in many materials science applications, obtaining a large annotated dataset is cost and labor intensive. In this work, we study the use of synthetic data to enlarge the available annotated experimental data of uranium oxide. We utilize a modified Potts model to simulate uranium oxide particles with morphologies similar to those observed experimentally. We then leverage an image-to-image translation model to synthesize the simulated particles as if they are acquired with SEM. Through this process, we obtain pairs of particle images and their corresponding SEM representations, which corresponds to pairs of annotations and images. Unlike previous works, we leverage synthetic data for pretraining a CNN model prior, and finetune that model further with experimental data. We experimentally demonstrate that using synthetic data as incremental learning process benefits the overall performance compared to training a model on combined synthetic and experimental data.

36 MATERIALS SCIENCE↗

Pulse profile modelling of thermonuclear burst oscillations − I. The effect of neglecting variability

ABSTRACT We study the effects of the time-variable properties of thermonuclear X-ray bursts on modelling their millisecond-period burst oscillations. We apply the pulse profile modelling technique that is being used in the analysis of rotation-powered millisecond pulsars by the Neutron Star Interior Composition Explorer to infer masses, radii, and geometric parameters of neutron stars. By simulating and analysing a large set of models, we show that overlooking burst time-scale variability in temperatures and sizes of the hot emitting regions can result in substantial bias in the inferred mass and radius. To adequately infer neutron star properties, it is essential to develop a model for the time-variable properties or invest a substantial amount of computational time in segmenting the data into non-varying pieces. We discuss prospects for constraints from proposed future X-ray telescopes.

79 ASTRONOMY AND ASTROPHYSICS↗

Automated Energy-Dispersive X-ray Spectroscopy Analysis for Multi-Modal Few-Shot Learning

Scanning transmission electron microscopy (STEM) is a powerful tool that allows for the atomic-scale analysis of a materials’ structure, chemistry, and defect domains (Akers et al. 2021). The current generation of microscopes generate vast amounts of data, surpassing the limits of effective manual analysis traditionally performed by domain experts (Spurgeon et al. 2021). While recent strides in machine learning have significantly enhanced the processing of large and intricate datasets acquired through electron microscopy, the prevalent use of proprietary software packages for initial data collection poses a challenge. In many cases, these software packages act as a ‘black box’, constraining user functionality and hindering the output of data in a format that is conducive to seamless integration into machine learning models. This work addresses these challenges by adapting HyperSpy, an open-source Python library, for the analysis and quantification of raw energy dispersive spectroscopy (EDS) data acquired through STEM. The modified HyperSpy code successfully facilitates user-defined segmentation of the data, enabling the integration of atomic %, weight %, and raw EDS spectra for each segmented region into an existing few-shot machine learning model. While initial results reveal discrepancies in quantified atomic and weight percentages when compared to proprietary software, ongoing efforts aim to rectify this issue by refining the fit of the HyperSpy model to the EDS spectra. Overall, this research underscores the potential of open-source tools like HyperSpy to enhance the accessibility of analytical tools, fostering a transparent and user-friendly environment for seamlessly incorporating electron microscopy data into machine learning models.

36 MATERIALS SCIENCE↗

Constrained GAN-Generated X-Ray CT Data For Self-Supervised And Foundation-Model Segmentation Of Concrete Microstructures

Three-dimensional characterization of materials using X-ray computed tomography (XCT) is challenging due to the complexity of internal structures, noise, and variations in resolution. Traditional computer vision models often struggle to accurately segment these images, particularly in domain-specific applications like materials science. While supervised deep learning approaches have been developed to address the limitations of conventional algorithms, they typically require large amounts of labeled training data and often fail to generalize across different datasets. Self-supervised, few-and zero-shot learning methods have gained prominence in natural image processing and segmentation tasks, but their application to scientific imaging remains limited due to the unique structural complexity, noise, and textural artifacts present in materials science data. In this work, we investigate how domain adaptation, leveraging physics-based and GAN-generated synthetic data, impacts segmentation performance. We introduce a modified Contrastive Unpaired Translation (CUT) model designed to generate realistic labeled data, which can be used for training, pre-training, and fine-tuning segmentation models for real XCT microstructure data. We evaluate the performance of two segmentation approaches: a self-supervised network (SSL-ALPNet) and a foundation model (Segment Anything Model), assessing their improvements when pre-trained and/or fine-tuned on the synthesized data. Our results demonstrate that leveraging synthetic data significantly enhances segmentation performance, particularly in challenging materials science applications.

Ziabari, Amir [ORNL] (ORCID:000000034776457X)↗

Computing system operational methods and apparatus

Computing system operational methods and apparatus are described. According to one aspect, a computing system operational method includes accessing user information regarding a user logging onto a computing device of the computing system, processing the user information to determine if the user information is authentic, as a result of the processing determining that the user information is authentic, first enabling the computing device to execute an application segment, and as a result of the processing determining that the user information is authentic, second enabling the application segment to communicate data externally of the computing device via one of a plurality of network segments of the computing system.

Edgar, Thomas W.↗

TopoSZ: Preserving Topology in Error-Bounded Lossy Compression

Existing error-bounded lossy compression techniques control the pointwise error during compression to guarantee the integrity of the decompressed data. However, they typically do not explicitly preserve the topological features in data. When performing post hoc analysis with decompressed data using topological methods, preserving topology in the compression process to obtain topologically consistent and correct scientific insights is desirable. In this paper, we introduce TopoSZ, an error-bounded lossy compression method that preserves the topological features in 2D and 3D scalar fields. Specifically, we aim to preserve the types and locations of local extrema as well as the level set relations among critical points captured by contour trees in the decompressed data. The main idea is to derive topological constraints from contour-tree-induced segmentation from the data domain, and incorporate such constraints with a customized error-controlled quantization strategy from the SZ compressor (version 1.4). In conclusion, our method allows users to control the pointwise error and the loss of topological features during the compression process with a global error bound and a persistence threshold.

97 MATHEMATICS AND COMPUTING↗

Mappymatch FKA: YAMM (Yet Another Map-Matcher) [SWR 22-38]

A surprisingly non-trivial technical challenge is to associate points in space (e.g., GPS data) with specific segments of a road network or map. The software package ( allows users to match GPS point data to a road network (commonly known as "map matching"). The software is designed such that a user could match a set of GPS points to a variety of different road network representations using a variety of map matching algorithms. There are currently several "built-in" road networks and map matching algorithms but the software has been designed to enable new ones to be added with minimal overhead.

Reinicke, Nicholas↗

Model Assumptions and Data Characteristics: Impacts on Domain Adaptation in Building Segmentation

Studies on domain adaptation (DA) for remote sensing (RS) imagery analysis lack consistency in selection and description of evaluation scenarios. Without properly characterizing datasets, model assumptions, and evaluation scenarios, it is difficult to objectively compare DA methods and reach conclusions about their suitability across different applications. With this motivation, this work seeks to empirically assess to which extent the interaction between data characteristics and model assumptions influences the effectiveness of DA methods. Using the widely explored task of building footprint segmentation as a case study, we perform a large-scale study across over 200 DA scenarios that include variations across view angles, areas observed, and sensors used for data acquisition. Rather than adopting different model architectures or optimization criteria, we contrast the performances of two DA methods based on adversarial learning that differ only in their assumptions about source and target domains. Informed by metadata and data characteristics unveiled using traditional computer vision (CV) techniques as well as pretrained deep models, we provide a detailed meta-analysis of experiments highlighting the importance of accurately considering data assumptions for DA in RS segmentation tasks. As demonstrated by a “cherry-picking” exercise, different claims regarding which model is best could be made by selecting different subsets of evaluation scenarios. While well-calibrated assumptions can be beneficial, mismatching assumptions can lead to negative biases in DA applications. Furthermore, this study intends to motivate the community toward more consistent evaluation protocols while providing recommendations and insights toward creating novel benchmark datasets, documenting data characteristics, application-specific knowledge, and model assumptions.

42 ENGINEERING↗

Automated analysis of lattice structures using computed tomography

Systems, methods, and computer-readable media for evaluating a set of computed tomography data associated with a lattice structure. The lattice structure may be additively manufactured. The computed tomography data may be segmented using a filter for identifying blob-like structures to identify nodes present within the lattice structure. A three-dimensional path traversal is applied to volumetric data to identify a plurality of struts within the lattice structure that are compared to corresponding struts within a set if three-dimensional mesh data of the lattice structure to identify defective struts. Further, two-dimensional slices may be extracted from each of the computed tomography data and the mesh data and compared to identify one or more inconsistencies indicative of defects within the lattice structure.

Schiefelbein, Bryan E.↗

Automated analysis of lattice structures using computed tomography

Systems, methods, and computer-readable media for evaluating a set of computed tomography data associated with a lattice structure. The lattice structure may be additively manufactured. The computed tomography data may be segmented using a filter for identifying blob-like structures to identify nodes present within the lattice structure. A three-dimensional path traversal is applied to volumetric data to identify a plurality of struts within the lattice structure that are compared to corresponding struts within a set if three-dimensional mesh data of the lattice structure to identify defective struts. Further, two-dimensional slices may be extracted from each of the computed tomography data and the mesh data and compared to identify one or more inconsistencies indicative of defects within the lattice structure.

Schiefelbein, Bryan E.↗

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM↗