Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data segmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

ClimateNet: an expert-labeled open dataset and deep learning architecture for enabling high-precision analyses of extreme weather

Abstract. Identifying, detecting, and localizing extreme weather events is a crucial first step in understanding how they may vary under different climate change scenarios. Pattern recognition tasks such as classification, object detection, and segmentation (i.e., pixel-level classification) have remained challenging problems in the weather and climate sciences. While there exist many empirical heuristics for detecting extreme events, the disparities between the output of these different methods even for a single event are large and often difficult to reconcile. Given the success of deep learning (DL) in tackling similar problems in computer vision, we advocate a DL-based approach. DL, however, works best in the context of supervised learning – when labeled datasets are readily available. Reliable labeled training data for extreme weather and climate events is scarce. We create “ClimateNet” – an open, community-sourced human-expert-labeled curated dataset that captures tropical cyclones (TCs) and atmospheric rivers (ARs) in high-resolution climate model output from a simulation of a recent historical period. We use the curated ClimateNet dataset to train a state-of-the-art DL model for pixel-level identification – i.e., segmentation – of TCs and ARs. We then apply the trained DL model to historical and climate change scenarios simulated by the Community Atmospheric Model (CAM5.1) and show that the DL model accurately segments the data into TCs, ARs, or “the background” at a pixel level. Further, we show how the segmentation results can be used to conduct spatially and temporally precise analytics by quantifying distributions of extreme precipitation conditioned on event types (TC or AR) at regional scales. The key contribution of this work is that it paves the way for DL-based automated, high-fidelity, and highly precise analytics of climate data using a curated expert-labeled dataset – ClimateNet. ClimateNet and the DL-based segmentation method provide several unique capabilities: (i) they can be used to calculate a variety of TC and AR statistics at a fine-grained level; (ii) they can be applied to different climate scenarios and different datasets without tuning as they do not rely on threshold conditions; and (iii) the proposed DL method is suitable for rapidly analyzing large amounts of climate model output. While our study has been conducted for two important extreme weather patterns (TCs and ARs) in simulation datasets, we believe that this methodology can be applied to a much broader class of patterns and applied to observational and reanalysis data products via transfer learning.

54 ENVIRONMENTAL SCIENCES↗

Towards automating structural discovery in scanning transmission electron microscopy *

Abstract Scanning transmission electron microscopy is now the primary tool for exploring functional materials on the atomic level. Often, features of interest are highly localized in specific regions in the material, such as ferroelectric domain walls, extended defects, or second phase inclusions. Selecting regions to image for structural and chemical discovery via atomically resolved imaging has traditionally proceeded via human operators making semi-informed judgements on sampling locations and parameters. Recent efforts at automation for structural and physical discovery have pointed towards the use of ‘active learning’ methods that utilize Bayesian optimization with surrogate models to quickly find relevant regions of interest. Yet despite the potential importance of this direction, there is a general lack of certainty in selecting relevant control algorithms and how to balance a priori knowledge of the material system with knowledge derived during experimentation. Here we address this gap by developing the automated experiment workflows with several combinations to both illustrate the effects of these choices and demonstrate the tradeoffs associated with each in terms of accuracy, robustness, and susceptibility to hyperparameters for structural discovery. We discuss possible methods to build descriptors using the raw image data and deep learning based semantic segmentation, as well as the implementation of variational autoencoder based representation. Furthermore, each workflow is applied to a range of feature sizes including NiO pillars within a La:SrMnO 3 matrix, ferroelectric domains in BiFeO 3 , and topological defects in graphene. The code developed in this manuscript is open sourced and will be released at github.com/nccreang/AE_Workflows .

47 OTHER INSTRUMENTATION↗

High-Throughput Field Plant Phenotyping: A Self-Supervised Sequential CNN Method to Segment Overlapping Plants

High-throughput plant phenotyping—the use of imaging and remote sensing to record plant growth dynamics—is becoming more widely used. The first step in this process is typically plant segmentation, which requires a well-labeled training dataset to enable accurate segmentation of overlapping plants. However, preparing such training data is both time and labor intensive. To solve this problem, we propose a plant image processing pipeline using a self-supervised sequential convolutional neural network method for in-field phenotyping systems. This first step uses plant pixels from greenhouse images to segment nonoverlapping in-field plants in an early growth stage and then applies the segmentation results from those early-stage images as training data for the separation of plants at later growth stages. The proposed pipeline is efficient and self-supervising in the sense that no human-labeled data are needed. We then combine this approach with functional principal components analysis to reveal the relationship between the growth dynamics of plants and genotypes. We show that the proposed pipeline can accurately separate the pixels of foreground plants and estimate their heights when foreground and background plants overlap and can thus be used to efficiently assess the impact of treatments and genotypes on plant growth in a field environment by computer vision techniques. This approach should be useful for answering important scientific questions in the area of high-throughput phenotyping.

59 BASIC BIOLOGICAL SCIENCES↗

Image Processing Pipeline for Fluoroelastomer Crystallite Detection in Atomic Force Microscopy Images

Phase transformations in materials systems can be tracked using atomic force microscopy (AFM), enabling the examination of surface properties and macroscale morphologies. In situ measurements investigating phase transformations generate large datasets of time-lapse image sequences. The interpretation of the resulting image sequences, guided by domain-knowledge, requires manual image processing using handcrafted masks. Here this approach is time-consuming and restricts the number of images that can be processed. Her in this study, we developed an automated image processing pipeline which integrates image detection and segmentation methods. We examine five time-series AFM videos of various fluoroelastomer phase transformations. The number of image sequences per video ranges from a hundred to a thousand image sequences. The resulting image processing pipeline aims to automatically classify and analyze images to enable batch processing. Using this pipeline, the growth of each individual fluoroelastomer crystallite can be tracked through time. We incorporated statistical analysis into the pipeline to investigate trends in phase transformations between different fluoroelastomer batches. Understanding these phase transformations is crucial, as it can provide valuable insights into manufacturing processes, improve product quality, and possibly lead to the development of more advanced fluoroelastomer formulations.

36 MATERIALS SCIENCE↗

Automated 3D cytoplasm segmentation in soft X-ray tomography

Cells’ structure is key to understanding cellular function, diagnostics, and therapy development. Soft X-ray tomography (SXT) is a unique tool to image cellular structure without fixation or labeling at high spatial resolution and throughput. Fast acquisition times increase demand for accelerated image analysis, like segmentation. Currently, segmenting cellular structures is done manually and is a major bottleneck in the SXT data analysis. This paper introduces ACSeg, an automated 3D cytoplasm segmentation model. ACSeg is generated using semi-automated labels and 3D U-Net and is trained on 43 SXT tomograms of immune T cells, rapidly converging to high-accuracy segmentation, therefore reducing time and labor. Furthermore, adding only 6 SXT tomograms of other cell types diversifies the model, showing potential for optimal experimental design. ACSeg successfully segmented unseen tomograms and is published on Biomedisa, enabling high-throughput analysis of cell volume and structure of cytoplasm in diverse cell types.

59 BASIC BIOLOGICAL SCIENCES↗

Scalable probabilistic estimates of electric vehicle charging given observed driver behavior

To prepare for rapid growth in global electric vehicle adoption, grid and policy planners depend on detailed forecasts of future charging demand. In this paper we propose a novel holistic, scalable, probabilistic framework to produce large-scale estimates of electric vehicle charging load for long-term planning that capture real drivers’ charging patterns. Our framework captures the uncertainty and stochasticity in charging demand by taking a graphical modeling approach. It has three core elements: driver groups, charging segment choices, and charging session time and energy requirements. The framework uses hierarchical clustering to group drivers by their charging histories, capturing their heterogeneous behaviors and preferences across different segments or types of charging. The framework uses probabilistic mixture models for each driver group’s sessions to identify the unique charging behaviors observed within each segment. We illustrate its application with a large data set from California, profiling the charging patterns and unique driver clusters it identifies. Using the model knobs representing drivers’ battery capacities, behavior, and segment access we present scenarios for California’s charging demand in 2030 with 8 million passenger electric vehicles. Peak charging demand ranged from 3.3 to 8.7 GW across scenarios. Furthermore, each was calculated in under 45 s on a laptop computer.

33 ADVANCED PROPULSION SYSTEMS↗

Frosted Tracks

SAND2025-01893O Frosted Tracks is a software tool to group trajectories according to sequences of their behavior. The goal is to start with a very large number of trajectories and identify groups that exhibit similar behavior patterns. The application combines TICC and Metric DBSCAN clustering algorithms for behavioral segmentation and labeling of air/sea trajectory data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Dalbey, Keith↗

Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) using Machine Learning

Modern retrofit construction practices use 3D point cloud data of the building envelope to obtain the as-built dimensions. However, manual segmentation by a trained professional is required to identify and measure window openings, door openings, and other architectural features, making the use of 3D point clouds labor-intensive. In this study, the Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) algorithm is described, which can significantly reduce the time spent during point cloud segmentation. The Auto-CuBES algorithm inputs a 3D point cloud generated by commonly available surveying equipment and outputs a wire-frame model of the building envelope. Unsupervised machine learning methods were used to identify facades, windows, and doors while minimizing the number of calibration parameters. Additionally, Auto-CuBES generates a heat map of each facade indicating non-planar characteristics that are crucial for the optimization of connections used in overclad envelope retrofits. With a scan resolution of 3 mm, the resulting window dimensions showed a mean absolute error of 4.2 mm compared to manual laser measurements.

Maldonado Puente, Bryan↗

Deep learning-based segmentation of lithium-ion battery microstructures enhanced by artificially generated electrodes

Accurate 3D representations of lithium-ion battery electrodes, in which the active particles, binder and pore phases are distinguished and labeled, can assist in understanding and ultimately improving battery performance. Here, we demonstrate a methodology for using deep-learning tools to achieve reliable segmentations of volumetric images of electrodes on which standard segmentation approaches fail due to insufficient contrast. We implement the 3D U-Net architecture for segmentation, and, to overcome the limitations of training data obtained experimentally through imaging, we show how synthetic learning data, consisting of realistic artificial electrode structures and their tomographic reconstructions, can be generated and used to enhance network performance. We apply our method to segment x-ray tomographic microscopy images of graphite-silicon composite electrodes and show it is accurate across standard metrics. We then apply it to obtain a statistically meaningful analysis of the microstructural evolution of the carbon-black and binder domain during battery operation.

25 ENERGY STORAGE↗

Geothermal play fairway analysis, part 1: Example from the Snake River Plain, Idaho

The Snake River Plain (SRP) volcanic province overlies the track of the Yellowstone hotspot, a thermal anomaly that extends deep into the mantle. Most of the area is underlain by a basaltic volcanic province that overlies a mid-crustal intrusive complex, which in turn provides the long-term heat flux needed to sustain geothermal systems. Previous studies have identified several known geothermal resource areas within the SRP. For the geothermal study presented herein, our goals were to: (1) adapt the methodology of Play Fairway Analysis (PFA) for geothermal exploration to create a formal basis for its application to geothermal systems, (2) assemble relevant data for the SRP from publicly available and private sources, and (3) build a geothermal PFA model for the SRP and identify the most promising plays, using GIS-based software tools that are standard in the petroleum industry. The study focused on identifying three critical resource parameters for exploitable hydrothermal systems in the SRP: heat source, reservoir and recharge permeability, and cap or seal. Data included in the compilation for heat source were heat flow, distribution and ages of volcanic vents, groundwater temperatures, thermal springs and wells, helium isotope anomalies, and reservoir temperatures estimated using geothermometry. Reservoir and recharge permeability was inferred from the analysis of stress orientations and magnitudes, post-Miocene faults, and subsurface structural lineaments based on magnetics and gravity data. Data for cap or seal included the distribution of impermeable lake sediments and clay-seal associated with hydrothermal alteration below the regional aquifer. These data were used to compile Common Risk Segment maps for heat, permeability, and seal, which were combined to create a Composite Common Risk Segment map for all southern Idaho that reflects the risk associated with geothermal resource exploration and identifies favorable resource tracks. Our regional assessment indicated that undiscovered geothermal resources may be located in several areas of the SRP. Two of these areas, the western SRP and Camas Prairie, were selected for more detailed assessment, during which heat, permeability, and seal were evaluated using newly collected field data and smaller grid parameters to refine the location of potential resources. These higher resolution assessments illustrate the flexibility of our approach over a range of scales.

54 ENVIRONMENTAL SCIENCES↗

Enhancing segmentation fairness through curriculum learning and progressive loss: a centralized and federated perspective on radiograph analysis

Bias in medical image segmentation can lead to unequal performance across demographic subgroups, raising concerns about fairness and reliability in clinical AI systems. While deep learning models have achieved high segmentation accuracy, ensuring equitable performance across race and gender remains a significant challenge, particularly in privacy-sensitive healthcare environments. This study investigates fairness-aware medical image segmentation for hip and knee radiographs using deep learning models evaluated in both centralized and Federated Learning (FL) settings. We introduce Curriculum Learning (CL) strategies and Progressive Loss (PL) functions to regulate sample difficulty during training. In addition, we propose two novel fairness-oriented federated learning algorithms, Federated Intersection over Union (FedIoU) and Federated Intersection over Union with Outlier Analysis (FedIoUoutlier). Experiments are conducted using multiple segmentation backbones and simulated multi-site data partitions derived from the Osteoarthritis Initiative dataset. Model performance is evaluated using Intersection over Union (IoU), IoU standard deviation, Skewed Error Ratio (SER), and Min-Max Disparity across race and gender subgroups. Statistical significance was verified using paired t-tests to compare per-sample IoU performance against baseline configurations. Across both hip and knee segmentation tasks, curriculum learning and progressive loss strategies consistently improved segmentation accuracy and reduced demographic performance disparities in centralized training. In federated settings, fairness-aware aggregation further enhanced performance. Notably, FedIoUoutlier combined with balanced curriculum learning and tiered progressive loss achieved the highest mean IoU while yielding the lowest SER and Min-Max Disparity, indicating improved fairness without sacrificing accuracy. In several configurations, federated models matched or exceeded the performance of optimized centralized models, with statistically significant improvements in per-sample IoU over baseline configurations. The results demonstrate that structured training strategies and fairness-aware federated aggregation can jointly improve accuracy, stability, and demographic fairness in medical image segmentation. By integrating curriculum learning, progressive loss, and novel FL algorithms, this work provides a practical pathway toward equitable and privacy-preserving AI systems for medical imaging.

97 MATHEMATICS AND COMPUTING↗

Imaging and Segmenting Grains and Subgrains Using Backscattered Electron Techniques

We present two new methods of processing data from backscattered electron signals in a scanning electron microscope to image grains and subgrains. The first combines data from multiple backscattered electron images acquired at different specimen geometries to (1) better reveal grain boundaries in recrystallized microstructures and (2) distinguish between recrystallized and unrecrystallized regions in partially recrystallized microstructures. The second utilizes spherical harmonic transform indexing of electron backscatter diffraction patterns to produce high angular resolution orientation data that enable the characterization of subgrains. Subgrains are produced during high-temperature plastic deformation and have boundary misorientation angles ranging from a few degrees down to a few hundredths of a degree. Here, we also present an algorithm to automatically segment grains from combined backscattered electron image data or grains and subgrains from high angular resolution electron backscatter diffraction data. Together, these new techniques enable rapid measurements of individual grains and subgrains from large populations.

36 MATERIALS SCIENCE↗

Mauka Energy FEVER Tool Dataset

Mauka Energy’s dataset, developed under the Forestry Electric Vehicle Energy Routing (FEVER) project and funded by the U.S. Department of Energy’s Small Business Innovation Research program, is a high-resolution geospatial resource designed to support energy modeling for electric log trucks in complex forestry environments. The dataset integrates detailed spatial and road network data to enable accurate simulation of vehicle performance across varied terrain. At its core, the dataset incorporates lidar-derived elevation models, road alignments, and surface classifications from Oregon State University’s McDonald-Dunn Research Forest. These data capture fine-scale variations in slope, curvature, and surface conditions across forest road systems, allowing for vehicle-level analysis of energy consumption and recovery. The dataset also includes data collected on the surrounding public and private road networks in Benton County, Oregon, used in real-world haul routes. These connecting segments provide critical context for modeling transitions between forest operations and regional transportation infrastructure, incorporating attributes such as grade profiles, elevation change, and speed constraints. This combined dataset underpins the development of Mauka Energy’s rolldown tool, which quantifies energy use and regenerative braking potential on downhill and variable-grade segments. By leveraging high-resolution terrain and road data, the FEVER project enables more accurate assessment of electric vehicle feasibility and performance in forestry applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The Number and Pattern of Viral Genomic Reassortments are not Necessarily Identifiable from Segment Trees

Reassortment is an evolutionary process common in viruses with segmented genomes. These viruses can swap whole genomic segments during cellular co-infection, giving rise to novel progeny formed from the mixture of parental segments. Since large-scale genome rearrangements have the potential to generate new phenotypes, reassortment is important to both evolutionary biology and public health research. However, statistical inference of the pattern of reassortment events from phylogenetic data is exceptionally difficult, potentially involving inference of general graphs in which individual segment trees are embedded. In this paper, we argue that, in general, the number and pattern of reassortment events are not identifiable from segment trees alone, even with theoretically ideal data. We call this fact the fundamental problem of reassortment, which we illustrate using the concept of the “first-infection tree,” a potentially counterfactual genealogy that would have been observed in the segment trees had no reassortment occurred. Further, we illustrate four additional problems that can arise logically in the inference of reassortment events and show, using simulated data, that these problems are not rare and can potentially distort our observation of reassortment even in small data sets. Finally, we discuss how existing methods can be augmented or adapted to account for not only the fundamental problem of reassortment, but also the four additional situations that can complicate the inference of reassortment.

59 BASIC BIOLOGICAL SCIENCES↗

RU-net for automatic characterization of TRISO fuel cross sections

During irradiation, phenomena such as kernel swelling and buffer densification may impact the performance of tristructural isotropic (TRISO) particle fuel. Post-irradiation microscopy is often used to identify these irradiation-induced morphologic changes. However, each fuel compact generally contains thousands of TRISO particles. Manually performing the work to get statistical information on these phenomena is cumbersome and subjective. Here, to reduce the subjectivity inherent in that process and to accelerate data analysis, we used convolutional neural networks (CNNs) to automatically segment cross-sectional images of microscopic TRISO layers. CNNs are a class of machine-learning algorithms specifically designed for processing structured grid data. They have gained popularity in recent years due to their remarkable performance in various computer vision tasks, including image classification, object detection, and image segmentation. In this research, we generated a large irradiated TRISO layer dataset with more than 2,000 microscopic images of cross-sectional TRISO particles and the corresponding annotated images. Based on these annotated images, we used different CNNs to automatically segment different TRISO layers. These CNNs include RU-Net (developed in this study), as well as three existing architectures: U-Net, Residual Network (ResNet), and Attention U-Net. The preliminary results show that the model based on RU-Net performs best in terms of Intersection over Union (IoU). Using CNN models, we can expedite the analysis of TRISO particle cross sections, significantly reducing the manual labor involved and improving the objectivity of the segmentation results.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Application of unsupervised deep learning to image segmentation and in-situ contact angle measurements in a CO 2 -water-rock system

Rock surface wettability is a critical property that regulates multiphase flows in porous media, which can be quantified using the surface contact angle (CA). X-ray micro-computed tomography (μCT) provides an effective approach to in-situ measurements of surface CAs. However, the CA measurement accuracy depends significantly on the quality of CT image segmentation, which is the clustering of CT pixels into separate phases. Inspired by this, we developed a deep learning (DL)-based CA measurement workflow. Motivated by the recent tremendous progress in unsupervised learning techniques and aiming to avoid expensive manual data annotations, an unsupervised DL pipeline for CT image segmentation was proposed and implemented, which includes unsupervised model training and post-processing. The unsupervised model training was driven by a novel loss function constrained with feature similarity and spatial continuity and implemented by iterative forward and backward paths; the former clustered the pixel-wise feature vectors extracted by convolution neural networks, whereas the latter updated the parameters using gradient descent. An over-segmentation strategy was adopted for model training. The post-processing steps based on agglomerative hierarchical clustering (AHC) were implemented to further merge the over-segmented model output to the desired cluster number, which is intended to improve the efficiency of image segmentation. The developed unsupervised DL pipeline was compared with other commonly-used image segmentation methods using pixel-wise and physics-based evaluation metrics on a synthetic raw-image dataset, which had a known ground truth. The unsupervised DL pipeline showed the best performance. Next, the segmented images were input to an automatic CA measurement tool, and the results were validated by comparisons with manual measurements. The CA values from the manual and automatic measurements showed similar distributions and statistical properties. The automatic measurement demonstrated a wider spectrum because of the much larger number of measurement data points. The primary novelty of the unsupervised DL pipeline developed in this study lies in the novel loss function and the over-segmentation strategy associated with AHC post-processing. Finally, the workflow has been proven an efficient tool for pore-scale wettability characterization, which has a wide range of applications in fundamental studies of multiphase flows in natural porous media, which have critical implications to geological carbon sequestration, hydrocarbon energy recovery, and contaminant transport in groundwater.

42 ENGINEERING↗

Statistical characterization of experimental magnetized liner inertial fusion stagnation images using deep-learning-based fuel–background segmentation

Significant variety is observed in spherical crystal x-ray imager (SCXI) data for the stagnated fuel–liner system created in Magnetized Liner Inertial Fusion (MagLIF) experiments conducted at the Sandia National Laboratories Z-facility. As a result, image analysis tasks involving, e.g., region-of-interest selection (i.e. segmentation), background subtraction and image registration have generally required tedious manual treatment leading to increased risk of irreproducibility, lack of uncertainty quantification and smaller-scale studies using only a fraction of available data. We present a convolutional neural network (CNN)-based pipeline to automate much of the image processing workflow. This tool enabled batch preprocessing of an ensemble of N scans = 139 SCXI images across N exp = 67 different experiments for subsequent study. The pipeline begins by segmenting images into the stagnated fuel and background using a CNN trained on synthetic images generated from a geometric model of a physical three-dimensional plasma. The resulting segmentation allows for a rules-based registration. Our approach flexibly handles rarely occurring artifacts through minimal user input and avoids the need for extensive hand labelling and augmentation of our experimental dataset that would be needed to train an end-to-end pipeline. Here we also fit background pixels using low-degree polynomials, and perform a statistical assessment of the background and noise properties over the entire image database. Our results provide a guide for choices made in statistical inference models using stagnation image data and can be applied in the generation of synthetic datasets with realistic choices of noise statistics and background models used for machine learning tasks in MagLIF data analysis. We anticipate that the method may be readily extended to automate other MagLIF stagnation imaging applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A workflow for segmenting soil and plant X-ray computed tomography images with deep learning in Google’s Colaboratory

X-ray micro-computed tomography (X-ray μCT) has enabled the characterization of the properties and processes that take place in plants and soils at the micron scale. Despite the widespread use of this advanced technique, major limitations in both hardware and software limit the speed and accuracy of image processing and data analysis. Recent advances in machine learning, specifically the application of convolutional neural networks to image analysis, have enabled rapid and accurate segmentation of image data. Yet, challenges remain in applying convolutional neural networks to the analysis of environmentally and agriculturally relevant images. Specifically, there is a disconnect between the computer scientists and engineers, who build these AI/ML tools, and the potential end users in agricultural research, who may be unsure of how to apply these tools in their work. Additionally, the computing resources required for training and applying deep learning models are unique, more common to computer gaming systems or graphics design work, than to traditional computational systems. To navigate these challenges, we developed a modular workflow for applying convolutional neural networks to X-ray μCT images, using low-cost resources in Google’s Colaboratory web application. Here we present the results of the workflow, illustrating how parameters can be optimized to achieve best results using example scans from walnut leaves, almond flower buds, and a soil aggregate. We expect that this framework will accelerate the adoption and use of emerging deep learning techniques within the plant and soil sciences.

59 BASIC BIOLOGICAL SCIENCES↗