Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Introduction to Special Section: Machine Learning for Image-based Geologic Interpretation

Image-based geological interpretation has been a labor-intensive and time-consuming process because it requires well-trained geoscientists to identify geological structures, features, and textures from various types of images. These images include scanning electron microscopic images, optical microscopic images, optical photos, resistivity images, seismic volumes, remote-sensing images, etc. With fast-evolving machine learning (ML) technology and computing power in recent decades, computers can achieve nearhuman-level to super-human-level performance with scalable high efficiency in the computer vision field. These technological revolutions facilitated image-based geological interpretation in petroleum exploration and production. For example, a fault picking method applied to 3-D seismic volume data using deep learning can achieve superior performance in comparison to conventional auto-picking methods. In addition, under the new normal of low oil prices, the petroleum industry seeks cost-effective strategies such as automating traditionally labor-intensive processes. Nevertheless, the potential of applying ML to geological image interpretation is still facing a few key challenges including data scarcity, data distribution, poor data and/or label quality, data leakage, learning algorithms, model architecture, training methodologies, testing and evaluation metrics, hyper-parameters optimization, model drift, production deployment, and the like.

58 GEOSCIENCES↗

Replace Human Intelligence with Fast and Smart Geometric Reasoning and Graph Neural Network to Accelerate Next Gen ModSim Workflows

We present an agent-guided approach to CAD geometry decomposition that automates hex/hybrid meshing with graph neural networks (GNNs) to accelerate next-generation ModSim workflows. Our end-to-end pipeline (i) reduces 3D boundary-representation (B-Rep) models to a 2D chordal axis skeleton (CAT) and then to a 1D bipartite graph of surface and curve nodes, (ii) assigns per node labels as Cubit® WebCut actions, (iii) trains a multi-action GNN under supervised learning, and (iv) predicts five surface-node and three curve-node actions on out-of-distribution test geometries. Each graph node carries geometric, topological, and meshing attributes drawn from the B-Rep “skin” and CAT “skeleton,” with two-way mappings across 3D↔2D↔1D representations to maintain traceability back to 3D CAD. The supervised learning model exhibits stable convergence of the binary cross-entropy loss and achieves 98.7% accuracy on unseen lattice models. To operationalize decision-making, we rank predicted commands by geometric significance and prototyped the agent-guided workflow through the Cubit® Meshing PowerTool GUI. As a stretch goal, we explore reinforcement learning (RL) to reduce or remove label requirements and to learn policies for action sequences that maximize total reward (e.g., size of hex-meshable regions and resulting hex mesh quality). When all-hex meshing is not feasible, the agent assists in producing hybrid meshes—prioritizing hex in critical regions and transitioning to tetrahedral elements (tets) elsewhere—maintaining fidelity while ensuring robustness. The overarching objective is to replace manual, heuristics-based decomposition with data-driven, reproducible automation, cutting meshing turnaround time by orders of magnitude. We anticipate direct impact on simulation workflows through intelligent, scalable decomposition of complex CAD models into hex-meshable subdomains.

97 MATHEMATICS AND COMPUTING↗

Closing the Loop between In Situ Stress Complexity and EGS Fracture Complexity

We present an agent-guided approach to CAD geometry decomposition that automates hex/hybrid meshing with graph neural networks (GNNs) to accelerate next-generation ModSim workflows. Our end-to-end pipeline (i) reduces 3D boundary-representation (B-Rep) models to a 2D chordal axis skeleton (CAT) and then to a 1D bipartite graph of surface and curve nodes, (ii) assigns per node labels as Cubit® WebCut actions, (iii) trains a multi-action GNN under supervised learning, and (iv) predicts five surface-node and three curve-node actions on out-of-distribution test geometries. Each graph node carries geometric, topological, and meshing attributes drawn from the B-Rep “skin” and CAT “skeleton,” with two-way mappings across 3D↔2D↔1D representations to maintain traceability back to 3D CAD. The supervised learning model exhibits stable convergence of the binary cross-entropy loss and achieves 98.7% accuracy on unseen lattice models. To operationalize decision-making, we rank predicted commands by geometric significance and prototyped the agent-guided workflow through the Cubit® Meshing PowerTool GUI. As a stretch goal, we explore reinforcement learning (RL) to reduce or remove label requirements and to learn policies for action sequences that maximize total reward (e.g., size of hex-meshable regions and resulting hex mesh quality). When all-hex meshing is not feasible, the agent assists in producing hybrid meshes—prioritizing hex in critical regions and transitioning to tetrahedral elements (tets) elsewhere—maintaining fidelity while ensuring robustness. The overarching objective is to replace manual, heuristics-based decomposition with data-driven, reproducible automation, cutting meshing turnaround time by orders of magnitude. We anticipate direct impact on simulation workflows through intelligent, scalable decomposition of complex CAD models into hex-meshable subdomains.

42 ENGINEERING↗

Comparison of automated chemical-guided segmentation and human annotation of soil organic matter in X-ray microcomputed tomography imaging in contrasted soil types

Soil organic matter (OM) formation and persistence is strongly influenced by the spatial distribution of organic substrates and microscale soil heterogeneity by dictating OM accessibility to microorganisms. However, traditional size and/or density fractionation techniques disrupt aggregate architecture, eliminating spatial information needed to fully understand intra-aggregate OM distribution. To quantify three-dimensional OM spatial distribution and automate segmentation in X-ray microcomputed tomography (µCT) imaging without human annotation bias, we developed an iodine gas vapor (I2) based staining workflow that eliminates labor-intensive manual annotation while maintaining segmentation accuracy, using aggregates from four taxonomically diverse soils (Xerofluvent, Haploxeroll Sphagnofibrist, Palehumult) with an 8-fold range of soil organic carbon. Human annotation of 10 µCT slices by the experienced and inexperienced annotators resulted in variations up to 3% in the Dice similarity coefficient (DSC), reflecting a degree of inherent subjectivity of manual labeling. Such inconsistencies are expected to compound as the number of manually annotated slices increases. Dual-energy µCT imaging at 33.1 keV (below the iodine (I) K-edge) and 33.2 keV (above the I K-edge) was used to resolve aggregate microstructure following I2 staining. The automated image subtraction pipeline identified OM regions by the I Kedge induced brightness increases, achieving DSC values of 0.58–0.83 relative to an experienced annotator. Sensitivity analyses revealed that the reconstruction alpha value—optimized via the open-source tool TomocuPy—and the 3D registration slice count were the primary determinants of accuracy, providing a novel benchmark for dual-energy soil imaging. The pipeline without GPU acceleration achieved 9.6 to 43.2 times faster than manual annotation. Using GPU-accelerated image post-processing and affine transformation matrices, the pipeline successfully segmented OM elements for large-scale datasets (3232×3232 pixel, 2048 slices) within ~5200 s from raw file acquisition to segmented output. The high-throughput approach enables the quantification of OM spatial distribution across diverse and heterogeneous soil.

Soil microbial biomass↗

Detecting damaged buildings using real-time crowdsourced images and transfer learning

After significant earthquakes, we can see images posted on social media platforms by individuals and media agencies owing to the mass usage of smartphones these days. These images can be utilized to provide information about the shaking damage in the earthquake region both to the public and research community, and potentially to guide rescue work. This paper presents an automated way to extract the damaged buildings images after earthquakes from social media platforms such as Twitter and thus identify the particular user posts containing such images. Using transfer learning and ~ 6500 manually labelled images, we trained a deep learning model to recognize images with damaged buildings in the scene. The trained model achieved good performance when tested on newly acquired images of earthquakes at different locations and when ran in near real-time on Twitter feed after the 2020 M7.0 earthquake in Turkey. Furthermore, to better understand how the model makes decisions, we also implemented the Grad-CAM method to visualize the important regions on the images that facilitate the decision.

58 GEOSCIENCES↗

Machine Learning for Improved Availability of the SNS Klystron High Voltage Converter Modulators

Beam availability has increased at the SNS, however, the targeted availability is greater than 95 %, while the SNS has failed to meet lower targets in the past. The HVCM used to power the linac klystrons have been one source of lost beam time and was chosen to explore using AI/ML techniques to improve reliability. Among the possibilities being explored are automating the tuning of HVCMs and predicting component failures such as capacitor aging, rectifier assemblies containing hundreds of diodes, and insulating oil degradation. The methodology pursued includes data cleaning, de-noising, post-analysis data labeling, and machine learning model development. We explore using Long Short-Term Memory and autoencoders for anomaly detection and prognostication used to schedule maintenance. We evaluate the use of model regularizers and constraints to improve the performance of the model and investigate methods to estimate the uncertainty of the models to provide a robust prediction with statistical interoperability. This paper describes the operational experience and known failures of the HVCMs and the proposed ML methodology and the preliminary results of training the AI/ML algorithms.

Pappas, G. C.↗

Machine Learning for Improved Availability of the SNS Klystron High Voltage Converter Modulators

Beam availability has increased at the SNS, however, the targeted availability is greater than 95 %, while the SNS has failed to meet lower targets in the past. The HVCM used to power the linac klystrons have been one source of lost beam time and was chosen to explore using AI/ML techniques to improve reliability. Among the possibilities being explored are automating the tuning of HVCMs and predicting component failures such as capacitor aging, rectifier assemblies containing hundreds of diodes, and insulating oil degradation. The methodology pursued includes data cleaning, de-noising, post-analysis data labeling, and machine learning model development. We explore using Long Short-Term Memory and autoencoders for anomaly detection and prognostication used to schedule maintenance. We evaluate the use of model regularizers and constraints to improve the performance of the model and investigate methods to estimate the uncertainty of the models to provide a robust prediction with statistical interoperability. This paper describes the operational experience and known failures of the HVCMs and the proposed ML methodology and the preliminary results of training the AI/ML algorithms.

Pappas, Chris↗

Machine Learning for Improved Availability of the SNS Klystron High Voltage Converter Modulators

Beam availability has increased at the SNS, however, the targeted availability is greater than 95 %, while the SNS has failed to meet lower targets in the past. The HVCM used to power the linac klystrons have been one source of lost beam time and was chosen to explore using AI/ML techniques to improve reliability. Among the possibilities being explored are automating the tuning of HVCMs and predicting component failures such as capacitor aging, rectifier assemblies containing hundreds of diodes, and insulating oil degradation. The methodology pursued includes data cleaning, de-noising, post-analysis data labeling, and machine learning model development. We explore using Long Short-Term Memory and autoencoders for anomaly detection and prognostication used to schedule maintenance. We evaluate the use of model regularizers and constraints to improve the performance of the model and investigate methods to estimate the uncertainty of the models to provide a robust prediction with statistical interoperability. This paper describes the operational experience and known failures of the HVCMs and the proposed ML methodology and the preliminary results of training the AI/ML algorithms.

Pappas, Chris↗

Massive all-atom analysis of 2D materials with quantum properties (Final report)

Improvements in microscopy have enabled the acquisition of data at a scale that is difficult to process manually, making automated machine learning approaches to analyzing experimental images essential. In this project, we developed and applied machine learning (ML) workflows for atomic resolution scanning transmission electron microscopy (STEM) images. This development included improving both methodology as well as generating user-friendly codes. We developed machine learning architectures which, after training, automatically identify the location and types of defects throughout a material. We used these data to produce class-averaged images of 2D atomic coordinates with up to 0.3 pm precision, uncovering the structure and oscillations of long-range strain fields around point defects in WSe 2-2x Te 2x . We also resolved a long-standing problem in this field in the training of ML models, a lack of labeled experimental data, by developing a cycle-GAN that transformed simulated-generated labeled data into labeled data indistinguishable from experiment and therefore suitable for training. This removed the remaining parts of the ML data processing workflow where human intervention was still critical and therefore a bottleneck to working at scale. Codes have been developed and released for this full machine learning workflow. ML approaches to partially automate STEM acquisition were also developed. Finally we applied ML and other advanced data processing methods to several materials science problems in two-dimensional materials, including studying the evolution of hyperuniformity with defect concentration in WSe2, understanding phase transformations in transition metal dichalcogenides during in-situ heating in the STEM, and exploring how 2D interfaces transform from twisted into aligned structures.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

DONUT: physics-aware machine learning for real-time X-ray nanodiffraction analysis

Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Materials science↗

DONUT: Physics-aware Machine Learning for Real-time X-ray Nanodiffraction Analysis

SF-25-088 Coherent X-ray scattering techniques are critical for investigating the fundamental structural properties of materials at the nanoscale. While advancements have made these experiments more accessible, real-time analysis remains a significant bottleneck, often hindered by artifacts and computational demands. In scanning X-ray nanodiffraction microscopy, which is widely used to spatially resolve structural heterogeneities, this challenge is compounded by the convolution of the divergent beam with the sample’s local structure. To address this, we introduce DONUT (Diffraction with Optics for Nanobeam by Unsupervised Training), a physics-aware neural network designed for the rapid and automated analysis of nanobeam diffraction data. By incorporating a differentiable geometric diffraction model directly into its architecture, DONUT learns to predict crystal lattice strain and orientation in real-time. Crucially, this is achieved without reliance on labeled datasets or pre-training, overcoming a fundamental limitation for supervised machine learning in X-ray science. We demonstrate experimentally that DONUT accurately extracts all features within the data over 200 times more efficiently than conventional fitting methods.

Zhou, Tao [Argonne National Laboratory (ANL), Argo↗

Network Architecture Verification & Validation Tool

The NAVV Tool is an automation of Linux commands run Zeek IDS software on a packet capture to create a Microsoft Excel spreadsheet table breaking down network traffic observed. The tool automates the Zeek software analysis, the collation of logs, and then the dissection of the Conn.log and DNS.logs to create a summary table within a Excel. This spreadsheet can then be updated with network segments using CIDR formatting and labels along with inventory information including name and IP address. Using the tool again will integrate these label and color coding into the existing analysis table to aid in conducting an evaluation of the network traffic.

Nichols, DonovanW↗

Dynamic Nanopeptide Assemblies for Trans-Tympanic Drug Delivery

Aim: Otitis media is a common otolaryngologic diagnosis worldwide. Invasive methods to curtail and treat frequent occurrences are undesirable, thus necessitating the identification and production of a non-invasive approach to treating the disease. Due to tympanic membrane thickness, ototopical drug delivery is challenging. In this preliminary study, formulations integrating nanopeptides and thermoresponsive polymeric hydrogels are utilized to improve the efficiency of trans-tympanic membrane drug delivery. Methods: Peptides were synthesized using standard Fmoc (fluorenylmethoxycarbonyl protecting group) based solid state peptide synthesis on an automated peptide synthesizer. Ciprofloxacin release was simulated using multiwell microplates with porous inserts. Rate of Ciprofloxacin release was measured over a 48-hour period using a 200 uL solution of peptide fibers and Ciprofloxacin at 1 wt% each, and the labeled peptide at 0.1 wt% in PBS at pH of 7.4. The cytotoxicity of the PA (peptide amphiphile, specifically c16-AHL 3 K 3 -CO 2 H) micelle and fiber with and without ciprofloxacin was investigated by examining epidermal keratinocyte viability in the presence of the material at various concentrations. Laser scanning confocal microscopy was performed with excitation of the calcein dye at 485 nm and the PA-TAMRA (rhodamine labeled peptide) at 515 nm. Results: We have demonstrated the potential viability of a self-assembled peptide amphiphile hydrogel capable of transitioning from a network of 1D nanoscale fibers to 0D micelles. This dissociative mechanism of action yields a peptide that is an effective cell penetrating peptide (CPP) while temporally controlling the release of the antibiotic ciprofloxacin. Conclusion: This work highlights the potential utility of the dynamic process of an engineered peptide hydrogel capable of dissociating into CPPs capable of facilitating drug delivery across the tympanic membrane.

cell culture↗

NanoSIP: NanoSIMS Applications for Microbial Biology

High-resolution imaging with secondary ion mass spectrometry (nanoSIMS) has become a standard method in systems biology and environmental biogeochemistry and is broadly used to decipher ecophysiological traits of environmental microorganisms, metabolic processes in plant and animal tissues, and cross-kingdom symbioses. When combined with stable isotope-labeling—an approach we refer to as nanoSIP—nanoSIMS imaging offers a distinctive means to quantify net assimilation rates and stoichiometry of individual cell-sized particles in both low- and high-complexity environments. While the majority of nanoSIP studies in environmental and microbial biology have focused on nitrogen and carbon metabolism (using 15 N and 13 C tracers), multiple advances have pushed the capabilities of this approach in the past decade. The development of a high-brightness oxygen ion source has enabled high-resolution metal analyses that are easier to perform, allowing quantification of metal distribution in cells and environmental particles. New preparation methods, tools for automated data extraction from large data sets, and analytical approaches that push the limits of sensitivity and spatial resolution have allowed for more robust characterization of populations ranging from marine archaea to fungi and viruses. Further, NanoSIMS studies continue to be enhanced by correlation with orthogonal imaging and ‘omics approaches; when linked to molecular visualization methods, such as in situ hybridization and antibody labeling, these techniques enable in situ function to be linked to microbial identity and gene expression. Here we present an updated description of the primary materials, methods, and calculations used for nanoSIP, with an emphasis on recent advances in nanoSIMS applications, key methodological steps, and potential pitfalls.

07 ISOTOPE AND RADIATION SOURCES↗

The System for Classification of Low-Pressure Systems (SyCLoPS): An All-In-One Objective Framework for Large-Scale Data Sets

We propose the first unified objective framework (SyCLoPS) for detecting and classifying all types of low-pressure systems (LPSs) in a given data set. We use the state-of-the-art automated feature tracking software TempestExtremes (TE) to detect and track LPS features globally in ERA5 and compute 16 parameters from commonly found atmospheric variables for classification. A Python classifier is implemented to classify all LPSs at once. The framework assigns 16 different labels (classes) to each LPS data point and designates four different types of high-impact LPS tracks, including tracks of tropical cyclone (TC), monsoonal system, subtropical storm and polar low. The classification process involves disentangling high-altitude and drier LPSs, differentiating tropical and non-tropical LPSs using novel criteria, and optimizing for the detection of the four types of high-impact LPS. A comparison of our labels with those in the International Best Track Archive for Climate Stewardship (IBTrACS) revealed an overall accuracy of 95% in distinguishing between tropical systems, extratropical cyclones, and disturbances. SyCLoPS produces a better TC detection skill compared to the previous algorithms, highlighted by an approximately 6% reduction in the false alarm rate compared to the previous TE algorithm. The vertical cross section composite of the four types of high-impact LPS we detect each shows distinct structural characteristics. Finally, we demonstrate that SyCLoPS is valuable for investigating various aspects of LPSs in climate data, such as the evolution of a single LPS track, patterns of LPS frequencies, and precipitation or wind influence associated with a particular LPS class.

54 ENVIRONMENTAL SCIENCES↗

A Data-Driven Framework for Automated Detection of Aircraft-Generated Signals in Seismic Array Data Using Machine Learning

Abstract Ground motions associated with aircraft overflights can cover a significant portion of the seismic data collected by shallowly emplaced seismometers, such as new nodal and Distributed Acoustic Sensing systems. This article describes the first published framework for automated detection of aircraft on single channel and multichannel seismic data. The seismic data are converted to spectrograms in a sliding time window and classified as aircraft or nonaircraft in each window using a deep convolutional neural network trained with analyst-labeled data. A majority voting scheme is used to convert the output from the sequence of sliding time windows onto a decision time sequence for each channel and to combine the binary classifications on the decision time sequences across multiple channels. Precision, recall, and F-score are used to quantify the detection performance of the algorithm on nodal data using fourfold time-series cross validation. By applying our framework to data from the Sage Brush Flats nodal array in Southern California, we provide a benchmark performance and demonstrate the advantage of using an array of sensors.

Geochemistry & Geophysics↗

Beyond Binary: Automated PLC Memory Forensics through RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and the evolution of industrial control systems (ICS) to adopt Internet-based technologies enhanced productivity, but have inadvertently increased their vulnerability to cyber-based malicious attacks. When an ICS system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies to safeguard against future instances. Memory forensics is critical in the analysis process to ascertain what occurred. To date, approaches to analyze the persistent memory in ICS devices are limited, and almost nonexistent for volatile memory. This paper proposes an automated methodology, COMA, for PLC memory dump analysis using computer vision and deep learning techniques. Specifically, COMA converts the sequences of bytes in a PLC memory dump to RGB pixels and creates a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. COMA then uses the trained model to automatically segment new memory images and extract forensic artifacts. We evaluate COMA on a Schneider Electric Modicon M221 PLC involving two cyber-based attack scenarios: (i) code injection and (ii) code modification. The empirical results show that COMA can successfully detect attack artifacts in memory dumps in both scenarios.

Asmar Awad, Rima↗

Oak Ridge National Laboratory Building Envelope Library (ORNOBEL)

The Oak Ridge National Laboratory Building Envelope Library (ORNOBEL) is a collection of dense exterior building-facade point clouds acquired using a survey-grade terrestrial laser scanner. Each file represents an individual facade from a building on the Oak Ridge National Laboratory (ORNL) campus or in Knoxville, Tennessee, with an average point-cloud resolution of approximately 3 mm. The points in each facade are semantically labeled into three classes: (1) window/door, representing openings in the building envelope; (2) wall, representing planar opaque envelope surfaces; and (3) other, representing the remaining facade-adjacent elements, architectural features, and protrusions. ORNOBEL supports the development, training, and evaluation of advanced deep-learning methods for automated building-envelope segmentation, geometric reconstruction, and building information modeling (BIM).

Maldonado Puente, Bryan [ORNL] (ORCID:000000033880↗