Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data segmentation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A mesoscale 3D model of irradiated concrete informed via a 2.5 U-Net semantic segmentation

The concrete biological shield in light-water reactors is exposed to neutron and gamma irradiation, which deteriorates the concrete’s mechanical properties in the long term. To assess the irradiation-induced damage, predictive mechanical models are developed and used in parallel with the characterization of irradiated concrete samples. Realistic 3D simulation domains can drastically improve a model’s prediction. In this work, we utilized x-ray computed tomography (XCT) data of a concrete specimen to reconstruct its 3D microstructure. The XCT data shows low contrast between the concrete’s aggregates and cement paste, resulting in poor image segmentation when using traditional unsupervised techniques. To address this issue, we developed and trained a 2.5D U-Net model on only 24 pre-labeled XCT layers to segment 651 layers of the XCT data. The overall F1-score of the model is approximately 96%. Then, we created a 3D finite element (FE) mesh based on the stack of segmented images. The FE model contains radiation-induced expansion, damage, and creep. The constitutive equations are adapted to each phase (aggregates and cement paste). Here, we simulated the effects of neutron irradiation in the concrete specimen as well as the specimen’s mechanical response to uniaxial compression. Finally, model validation was performed using experimental data on similar concrete specimens in the literature.

2.5D U-Net↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Timeseries Unlabeled and Labeled Photos, Modeled Stream Elevation, and (Meta)Data of Variably Inundated Streams Across The Yakima River Basin, Washington, United States (v2)

This dataset is associated with the “River Monitoring Photos” (RMP) study and subsequent manuscript (Bao et al. 2025. Monitoring river flow status using low-cost wildlife camera and image segmentation artificial intelligence doi: 10.1016/j.envsoft.2025.106715). Game camera timeseries photos were collected to evaluate stream variable inundation via changes in width. A subset of photos was labeled for training the YOLOv8 and Mask2Former models and used to segment water surface fractions from all the game camera photos.This data package was originally published in March 2024. It was updated in October 2025 (v2) to add additional photos and files associated with the manuscript (i.e., processed data, labeled photos, and trained models). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to a readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; and (5) folders containing game camera photos and manuscript-associated files. Each Yakima River Basin site has a folder that contains subfolders for each month photos were collected. There is also a folder for files associated with the manuscript which has subfolders for labeled data, trained models, Yakima River Basin site water surface fractions, and USGS site water surface fractions. All files are .csv, .json, .txt, .yaml, .pth, .pt, or .pdf. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Prong Segmentation using Point Set Transformers in Multiple View Neutrino Detectors

NOvA is a long-baseline neutrino experiment studying neutrino oscillations by detecting neutrinos from the NuMI beam at Fermilab. Its physics analysis relies on accurate prong segmentation, which involves matching each hit to its source particle and identifying the particle type. This task has commonly been addressed using a combination of traditional clustering algorithms and convolutional neural networks (CNNs). However, NOvA’s detector design presents data as two sparse and decoupled 2D images (XZ and YZ views) rather than a native 3D representation, posing a significant challenge for traditional CNN-based models. In this talk, we propose a novel neural network based on the Point Set Transformer. By treating detector hits as sparse point clouds and implementing a cross-view attention mechanism, our model enables efficient information mixing between both views. Evaluated on NOvA simulated data, our model achieves superior accuracy while requiring significantly fewer computational resources compared to other models. Furthermore, the model demonstrates great performance when applied to Liquid Argon Time Projection Chamber (LArTPC) data, which shows its potential as a universal prong segmentation algorithm for multiple view neutrino detectors.

Liu, Jiaxi [UC, Irvine]↗

Machine-Learning-based Algorithms for Automated Image Segmentation Techniques of Transmission X-ray Microscopy (TXM)

Four state-of-the-art Deep Learning-based Convolutional Neural Networks (DCNN) were applied to automate the semantic segmentation of a 3D Transmission x-ray Microscopy (TXM) nanotomography image data. The standard U-Net architecture as baseline along with UNet++, PSPNet, and DeepLab v3+ networks were trained to segment the microstructural features of an AA7075 micropillar. A workflow was established to evaluate and compare the DCNN prediction dataset with the manually segmented features using the Intersection of Union (IoU) scores, time of training, confusion matrix, and visual assessment. Comparing all model segmentation accuracy metrics, it was found that using pre-trained models as a backbone along with appropriate training encoder-decoder architecture of the Unet++ can robustly handle large volumes of x-ray radiographic images in a reasonable amount of time. This opens a new window for handling accurate and efficient image segmentation of in situ time-dependent 4D x-ray microscopy experimental datasets.

36 MATERIALS SCIENCE↗

Dynamically Collected Local Density using Low-Cost Lidar and its Application to Traffic Models

This article demonstrates the use of traffic density observations collected dynamically in the vicinity of probe vehicles. Fixed position sensors cannot capture the longitudinal evolution of local traffic density in the corridor. In this research, dynamic traffic density observations were collected in a naturalistic driving setting that was free of any controlled experiment biases. Speed from global positioning system and space headway from a light detection and ranging module was collected on one arterial and one freeway segment, 2 and 4mi long, respectively. The combined data frequency was approximately 3Hz. Space headway was used to estimate the local density and consequently to identify the density of a specific location in a corridor. Besides, driver behavior was characterized using the relationship between instantaneous speed and local density under different regimes of the Wiedemann car-following model. Macroscopic traffic stream models were used to investigate the relationship between dynamically collected instantaneous speed and local density. Using the longitudinal evolution of density, precise local density across the corridor can be obtained along with the leader and follower trajectories. A method to identify driver behavior across density ranges was developed for different facility types using a microscopic relationship between instantaneous speed and local density. Overall driving behavior on the freeway segment can be represented by translating the instantaneous speed and local density relationship to macroscopic stream models.

Engineering↗

NanoPSD: A software for automatic detection of Nano-Particle Shape Distribution in electron microscopy images

Accurate quantification of the size and morphology of nanoparticles from electron microscopy (EM) images is essential to understand growth mechanisms, surface reactivity, and functional behavior in nanoscale materials. Manual analysis remains slow, subjective, and difficult to reproduce in large datasets. We introduce NanoPSD (Nano-Particle Shape Distribution), an open-source and fully automated framework for quantitative particle detection and morphology analysis from EM images. NanoPSD integrates adaptive contrast enhancement, polarity-agnostic scale-bar detection, Optical Character Recognition (OCR)-based calibration, and classical segmentation via Otsu thresholding with morphological refinement. Particle contours are used to extract geometric descriptors, including equivalent circular diameter, aspect ratio, circularity, and solidity, enabling automated classification into spherical, rod-like, and aggregate morphologies. The framework supports both single-image and batch processing, generating publication-quality visualizations, LaTeX-ready tables, and structured comma-separated values (CSV) datasets. As a demonstration, we applied NanoPSD to plasma-synthesized nanoparticle samples diagnosed via transmission electron microscopy (TEM). The code produced statistically robust size and morphology distributions spanning a few to tens of nanometers with minimal user supervision. The pipeline demonstrates high reproducibility and scalability, processing large image collections with consistent calibration and output formatting. Its modular design enables seamless integration of future deep-learning-based segmentation models, providing a pathway toward intelligent, data-driven electron microscopy analysis.

36 MATERIALS SCIENCE↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Systems and methods for vehicle cruise speed recommendation

A method for providing a cruising speed recommendation to an operator of a vehicle includes determining a projected route; receiving route characteristic data including route elevation data; determining a sampling resolution; sampling the route elevation data at the sampling resolution to generate sampled route elevation data; determining at least one start of uphill position and at least one start of downhill position; determining at least one cruise speed route segment based at least in part on the at least one start of uphill position and the at least one start of downhill position; determining a corresponding cruising speed for the at least one cruise speed route segment based at least in part on one or more of the route elevation data and the sampled route elevation data; and communicating the corresponding cruising speed for the at least one cruise speed route segment.

Aggoune, Karim↗

High-dimensional Data-driven Energy optimization for Multi-Modal Transit Agencies (HD-EMMA) (Final Technical Report)

Public bus transit services in the U.S. are responsible for at least 19.7 million metric tons of CO 2 emission annually. Electric vehicles (EVs) can have a much lower environmental impact than comparable internal combustion engine vehicles (ICEVs), especially in urban areas. Unfortunately, EVs are also much more expensive than ICEVs. As a result, many public transit agencies can afford only mixed fleets of transit vehicles, consisting of EVs, hybrids (HEVs), and ICEVs. Transit agencies that operate such mixed fleets of vehicles face a challenging optimization problem: these agencies need to decide which vehicles are assigned to serving which transit trips. Since the advantage of EVs over ICEVs varies depending on the route and time of day (e.g., the benefit of EVs is higher in slower traffic with frequent stops and lower on highways), the assignment can have a significant effect on energy use and, hence, environmental impact. Through this project, we have developed reference data about energy collections and constructed a set of machine learning models that can accurately predict the energy consumption for the whole fleet at the level of each trip. We have used these models to develop a scheduling and assignment strategy that can rotate the different vehicle types across the transit agencies’ routes. The optimization algorithm ensures that the vehicles are matched to trips considering weather patterns, expected congestion, and road gradients to minimize the overall energy usage. We list the key observations from our project for other practitioners below. Details are available in the report, and the list of source code and our publications are included in the appendix. 1. We have demonstrated the feasibility of collecting, merging and analyzing large volumes of high-resolution real-world telemetry data from a mixed vehicle fleet. To mitigate the inherent noise of the recorded GPS points, the team developed an algorithm that filters data and maps the points onto a street. The algorithm considers previous and subsequent location measurements and different characteristics of nearby streets to determine how likely the vehicle travels on them. Then, the team segmented the time series into disjoint contiguous samples based on adjacent road segments and repeated the outlier detection and removal. For each data point, the team added features corresponding to elevation changes within the samples, weather features, such as temperature, and traffic data, such as speed ratio between actual speed and free-flow speed. 2. We have developed two forms of machine learning models that be used to understand and analyze the energy operations of a mixed vehicle transit fleet. The micro prediction model provides estimates of instantaneous energy prediction for all types of buses (diesel, hybrid, and electric). Such a model is important in evaluating the energy impacts of real-time bus operation strategies, but it is challenging due to diversified driving cycles of transit buses. The model can help the drivers understand the impact of their driving behaviors and short-term congestions. The macro prediction models estimate average energy consumption across the whole trip considering the features: distance traveled, various road-type features, elevation change, day of the week, time of day, various weather features (temperature, humidity, etc.), and traffic features (speed ratio and jam factor). 3. We have demonstrated that it is possible to transfer the machine learning models we have developed in this project to other teams and cities by using inductive transfer learning. We also showed that the performance of the macro energy prediction models can be improved using a multi-task learning approach where the learning parameters are shared between the models being developed for different vehicle types. The advantage of this approach is improved learning performance as the models can exploit common spatio-temporal and environmental characteristics. 4. Finally, we have developed trip and vehicle assignment and scheduling algorithms that use the energy prediction models and develop a trip to vehicle type (diesel, electric, hybrid) assignment for the whole operation to reduce overall emissions and cost. We have shown through simulations that the proposed algorithms can save $\$$ 48,910 in energy costs and 175 metric tons of CO 2 emission annually for CARTA.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Developing Novel Performance Measures for Traffic Congestion Management and Operational Planning Based on Connected Vehicle Data

In this study, the authors present their efforts in exploring a new type of traffic data, referred to as internet-connected vehicle (ICV) data, for traffic congestion management and operational planning. Most currently manufactured vehicles contain onboard GPS and cellular modules, and they constantly connect to automobile manufacturers' clouds via cellular networks and upload their status. Some automobile manufacturers have recently redistributed the nonpersonal part of such data, such as geolocation, to third-party organizations for innovative applications. Compared with the traditional vehicle GPS data, the ICV data contain high-resolution GPS waypoints accompanied with the vehicles' abnormal moving events (e.g., hard braking). The ICV data also have huge potential in congestion management and operational planning. They explore to identify and analyze traffic congestion on both freeways and arterials using the ICV data. The ICV data adopted for this research are redistributed by Wejo Data Service, representing 10%-15% of all moving vehicles in the Dallas-Fort Worth (DFW) area in Texas. Through one case study for a freeway segment and one for an arterial segment, new traffic performance metrics based on the characteristics of ICV data have been presented. The highlights of these efforts are as follows: (I) queue length and propagation at freeway bottlenecks can be directly measured based on where and when most internet-connected vehicles slow down and join the queue; (II) an internet-connected vehicle's actual delay time on arterials can be directly measured according to its slow movement percentage, without assuming the nondelay travel speed; and (III) the ICV data set are also combined with the high-resolution traffic signal events to generate a ground-truth time-space diagram (TSD) on arterials - a common visualization of arterial signal performance for transportation planning and operations.

33 ADVANCED PROPULSION SYSTEMS↗

Multimodal Few-Shot Segmentation of Electron Micrographs

Scanning transmission electron microscopy (STEM) is one of the most used methods of analyzing the chemistry and composition of materials. By analyzing microstructures, these microscopes can help scientists better understand the molecular underpinnings of microelectronics, batteries, and more. However, STEM data can be difficult to interpret, so recent developments have been made in applications of machine learning to analyze these images. The PNNL-developed pyCHIP Classifier has achieved results in segmenting STEM these images via few-shot learning, a method which requires little data and human input, perfect for quickly analysis. In my internship I (Eli Meyers) investigated a multimodal improvement of this classifier by incorporating energy dispersive x-ray spectroscopy (EDS) data into the classification process for a more accurate segmentation. Furthermore, I encoded the spectral data by training a mass spectrometry encoder on the EDS data to extract a more meaningful representation of the data.

36 MATERIALS SCIENCE↗

Orchestration of materials science workflows for heterogeneous resources at large scale

In the era of big data, materials science workflows need to handle large-scale data distribution, storage, and computation. Any of these areas can become a performance bottleneck. We present a framework for analyzing internal material structures (e.g., cracks) to mitigate these bottlenecks. We demonstrate the effectiveness of our framework for a workflow performing synchrotron X-ray computed tomography reconstruction and segmentation of a silica-based structure. Our framework provides a cloud-based, cutting-edge solution to challenges such as growing intermediate and output data and heavy resource demands during image reconstruction and segmentation. Specifically, our framework efficiently manages data storage, scaling up compute resources on the cloud. The multi-layer software structure of our framework includes three layers. A top layer uses Jupyter notebooks and serves as the user interface. A middle layer uses Ansible for resource deployment and managing the execution environment. A low layer is dedicated to resource management and provides resource management and job scheduling on heterogeneous nodes (i.e., GPU and CPU). At the core of this layer, Kubernetes supports resource management, and Dask enables large-scale job scheduling for heterogeneous resources. The broader impact of our work is four-fold: through our framework, we hide the complexity of the cloud’s software stack to the user who otherwise is required to have expertise in cloud technologies; we manage job scheduling efficiently and in a scalable manner; we enable resource elasticity and workflow orchestration at a large scale; and we facilitate moving the study of nonporous structures, which has wide applications in engineering and scientific fields, to the cloud. While we demonstrate the capability of our framework for a specific materials science application, it can be adapted for other applications and domains because of its modular, multi-layer architecture.

97 MATHEMATICS AND COMPUTING↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

Eco-PiNN: A Physics-informed Neural Network for Eco-toll Estimation

The eco-toll estimation problem quantifies the expected environmental cost (e.g., energy consumption, exhaust emissions) for a vehicle to travel along a path. This problem is important for societal applications such as eco-routing, which aims to find paths with the lowest exhaust emissions or energy need. The challenges of this problem are threefold: (1) the dependence of a vehicle's eco-toll on its physical parameters; (2) the lack of access to data with eco-toll information; and (3) the influence of contextual information (i.e. the connections of adjacent segments in the path) on the eco-toll of road segments. Prior work on eco-toll estimation has mostly relied on pure data-driven approaches and has high estimation errors given the limited training data. To address these limitations, we propose a novel Eco-toll estimation Physics-informed Neural Network framework (Eco-PiNN) using three novel ideas, namely, (1) a physics-informed decoder that integrates the physical laws governing vehicle dynamics into the network, (2) an attention-based contextual information encoder, and (3) a physics-informed regularization to reduce overfitting. Experiments on real-world heavy-duty truck data show that the proposed method can greatly improve the accuracy of eco-toll estimation compared with state-of-the-art methods.

97 MATHEMATICS AND COMPUTING↗

Using porous random fields to predict the elastic modulus of unoxidized and oxidized superfine graphite

Nuclear graphite is a candidate material for Generation IV nuclear power plants. Porous materials such as graphite can contain complex networks of pores that influence the material's mechanical and irradiation response. A methodology known as the random finite element method (RFEM) was adapted to create synthetic microstructures and predict the influence of porosity on the elastic properties of graphite during oxidation. RFEM combines random field theory and the finite element method in a Monte Carlo framework to estimate the mechanical response of a given grade of graphite. In this research, the random fields were verified through experimental characterization to predict the elastic response of three nuclear graphite grades, ETU-10, IG-110, and 2114. Finite element models (FEM) were generated using segmentations of x-ray computed tomography (XCT) data known as image-based models (IBMs) to validate and compare with the RFEM results and better understand the effects of uniform oxidation in these graphite grades. The RFEM predictions appear to correlate well with the experimental values of the measured Young’s modulus of the three graphite grades and display the same trends as IBMs.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Combining Force Fields and Neural Networks for an Accurate Representation of Chemically Diverse Molecular Interactions

A key goal of molecular modeling is the accurate reproduction of the true quantum mechanical potential energy of arbitrary molecular ensembles with a tractable classical approximation. The challenges are that analytical expressions found in general purpose force fields struggle to faithfully represent the intermolecular quantum potential energy surface at close distances and in strong interaction regimes; that the more accurate neural network approximations do not capture crucial physics concepts, e.g., nonadditive inductive contributions and application of electric fields; and that the ultra-accurate narrowly targeted models have difficulty generalizing to the entire chemical space. We therefore designed a hybrid wide-coverage intermolecular interaction model consisting of an analytically polarizable force field combined with a short-range neural network correction for the total intermolecular interaction energy. Here, we describe the methodology and apply the model to accurately determine the properties of water, the free energy of solvation of neutral and charged molecules, and the binding free energy of ligands to proteins. The correction is subtyped for distinct chemical species to match the underlying force field, to segment and reduce the amount of quantum training data, and to increase accuracy and computational speed. For the systems considered, the hybrid ab initio parametrized Hamiltonian reproduces the two-body dimer quantum mechanics (QM) energies to within 0.03 kcal/mol and the nonadditive many-molecule contributions to within 2%. Simulations of molecular systems using this interaction model run at speeds of several nanoseconds per day.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗