Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Beyond Binary: Automated PLC Memory Forensics through RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and the evolution of industrial control systems (ICS) to adopt Internet-based technologies enhanced productivity, but have inadvertently increased their vulnerability to cyber-based malicious attacks. When an ICS system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies to safeguard against future instances. Memory forensics is critical in the analysis process to ascertain what occurred. To date, approaches to analyze the persistent memory in ICS devices are limited, and almost nonexistent for volatile memory. This paper proposes an automated methodology, COMA, for PLC memory dump analysis using computer vision and deep learning techniques. Specifically, COMA converts the sequences of bytes in a PLC memory dump to RGB pixels and creates a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. COMA then uses the trained model to automatically segment new memory images and extract forensic artifacts. We evaluate COMA on a Schneider Electric Modicon M221 PLC involving two cyber-based attack scenarios: (i) code injection and (ii) code modification. The empirical results show that COMA can successfully detect attack artifacts in memory dumps in both scenarios.

Asmar Awad, Rima↗

Oak Ridge National Laboratory Building Envelope Library (ORNOBEL)

The Oak Ridge National Laboratory Building Envelope Library (ORNOBEL) is a collection of dense exterior building-facade point clouds acquired using a survey-grade terrestrial laser scanner. Each file represents an individual facade from a building on the Oak Ridge National Laboratory (ORNL) campus or in Knoxville, Tennessee, with an average point-cloud resolution of approximately 3 mm. The points in each facade are semantically labeled into three classes: (1) window/door, representing openings in the building envelope; (2) wall, representing planar opaque envelope surfaces; and (3) other, representing the remaining facade-adjacent elements, architectural features, and protrusions. ORNOBEL supports the development, training, and evaluation of advanced deep-learning methods for automated building-envelope segmentation, geometric reconstruction, and building information modeling (BIM).

Maldonado Puente, Bryan [ORNL] (ORCID:000000033880↗

Statistical characterization of experimental magnetized liner inertial fusion stagnation images using deep-learning-based fuel–background segmentation

Significant variety is observed in spherical crystal x-ray imager (SCXI) data for the stagnated fuel–liner system created in Magnetized Liner Inertial Fusion (MagLIF) experiments conducted at the Sandia National Laboratories Z-facility. As a result, image analysis tasks involving, e.g., region-of-interest selection (i.e. segmentation), background subtraction and image registration have generally required tedious manual treatment leading to increased risk of irreproducibility, lack of uncertainty quantification and smaller-scale studies using only a fraction of available data. We present a convolutional neural network (CNN)-based pipeline to automate much of the image processing workflow. This tool enabled batch preprocessing of an ensemble of N scans = 139 SCXI images across N exp = 67 different experiments for subsequent study. The pipeline begins by segmenting images into the stagnated fuel and background using a CNN trained on synthetic images generated from a geometric model of a physical three-dimensional plasma. The resulting segmentation allows for a rules-based registration. Our approach flexibly handles rarely occurring artifacts through minimal user input and avoids the need for extensive hand labelling and augmentation of our experimental dataset that would be needed to train an end-to-end pipeline. Here we also fit background pixels using low-degree polynomials, and perform a statistical assessment of the background and noise properties over the entire image database. Our results provide a guide for choices made in statistical inference models using stagnation image data and can be applied in the generation of synthetic datasets with realistic choices of noise statistics and background models used for machine learning tasks in MagLIF data analysis. We anticipate that the method may be readily extended to automate other MagLIF stagnation imaging applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Automated Shift Detection in Sensor-Based PV Power and Irradiance Time Series

PV power and irradiance sensor-based measurements are prone to error, resulting in issues such as time series data shifts. In this research, a changepoint detection (CPD) algorithm that automatically detects data shifts in sensor-based time series is introduced. Data shift periods in 101 daily PV power and irradiance time series were labeled manually by two solar experts. These data streams represent sensor-based measurements, and display a variety of data shift behaviors. A changepoint detection algorithm was tuned using the 101 labeled data streams, with each model configuration's ability to detect labeled changepoints benchmarked using metrics such as F1-score, recall, and Rand Index. Best performing models on seasonality-corrected data streams include the Pruned Exact Linear (PELT) method, the Binary Segmentation method, and the Bottom-Up method, all scoring an average F1-score of 0.76 or greater at detecting labeled changepoints within a 30-day window across the labeled data sets. Pending approval, we plan to release the labeled data sets for this research on NREL's DuraMAT Data Hub, and the associated algorithm in the Python PVAnalytics package. By supplying the training sets and algorithm, we hope to encourage further development in this research space.

data shift↗

ROI-Finder : machine learning to guide region-of-interest scanning for X-ray fluorescence microscopy

The microscopy research at the Bionanoprobe (currently at beamline 9-ID and later 2-ID after APS-U) of Argonne National Laboratory focuses on applying synchrotron X-ray fluorescence (XRF) techniques to obtain trace elemental mappings of cryogenic biological samples to gain insights about their role in critical biological activities. The elemental mappings and the morphological aspects of the biological samples, in this instance, the bacterium Escherichia coli ( E. Coli ), also serve as label-free biological fingerprints to identify E. coli cells that have been treated differently. The key limitations of achieving good identification performance are the extraction of cells from raw XRF measurements via binary conversion, definition of features, noise floor and proportion of cells treated differently in the measurement. Automating cell extraction from raw XRF measurements across different types of chemical treatment and the implementation of machine-learning models to distinguish cells from the background and their differing treatments are described. Principal components are calculated from domain knowledge specific features and clustered to distinguish healthy and poisoned cells from the background without manual annotation. The cells are ranked via fuzzy clustering to recommend regions of interest for automated experimentation. The effects of dwell time and the amount of data required on the usability of the software are also discussed.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Deep Point Cloud Building Envelope Segmentation (DeeP-CuBES) using Deep Learning

Building Information Modeling (BIM) plays an important role in building design and construction, particularly for achieving energy-efficient retrofits. Building envelope retrofits using panelized prefabricated system, such as those popularized by the Energiesprong program, need accurate as-built dimensions of facade features (windows, doors, etc.) to achieve the desired thermal and air tightness. Traditionally, building surveying is done manually, resulting in a time-consuming and labor-intensive process. Recently, 3D point clouds from terrestrial LiDAR have been used to automate the generation of as-built dimensions of existing buildings. However, automated BIM using LiDAR relies on solving the point cloud semantic segmentation (PCSS) problem. In this work, we propose a robust pipeline for solving the PCSS problem using deep neural networks, focusing on overcoming challenges posed by imbalanced datasets and complex architectural features. We introduce the first high-density, labeled, and validated building envelope point cloud dataset derived from multiple building scans, specifically curated to tackle challenges in facade-level segmentation. Results from the trained neural networks show that advanced attention-based architectures and incorporating radiometry (light intensity and RGB) features significantly boost segmentation accuracy for windows and doors.

Selvakumar, Balaji [ORNL]↗

Automated segmentation of soft X-ray tomography: Native cellular structure with submicron resolution at high-throughput for whole-cell quantitative imaging in yeast

Soft X-ray tomography (SXT) is an invaluable tool for quantitatively analyzing cellular structures at suboptical isotropic resolution. However, it has traditionally depended on manual segmentation, limiting its scalability for large datasets. Here, we leverage a deep learning-based autosegmentation pipeline to segment and label cellular structures in hundreds of cells across three Saccharomyces cerevisiae strains. This task-based pipeline uses manual iterative refinement to improve segmentation accuracy for key structures, including the cell body, nucleus, vacuole, and lipid droplets, enabling high-throughput and precise phenotypic analysis. Using this approach, we quantitatively compared the three-dimensional (3D) whole-cell morphometric characteristics of wild-type, VPH1-GFP, and vac14 strains, uncovering detailed strain-specific cell and organelle size and shape variations. We show the utility of SXT data for precise 3D curvature analysis of entire organelles and cells and detection of fine morphological features using surface meshes. Our approach facilitates comparative analyses with high spatial precision and statistical throughput, uncovering subtle morphological features at the single-cell and population level. This workflow significantly enhances our ability to characterize cell anatomy and supports scalable studies on the mesoscale, with applications in investigating cellular architecture, organelle biology, and genetic research across diverse biological contexts.

Chen, Jianhua [Lawrence Berkeley National Laborato↗

Automated Programmable Logic Controller Memory Forensics Using RGB Image Analysis and Deep Learning

The introduction of Industry 4.0 and Internet-based technologies has enhanced industrial control system operations but have inadvertently increased their vulnerabilities to cyber attacks. When an industrial control system is compromised, security analysts need to identify the root cause quickly to start the recovery process and develop mitigation strategies. Memory forensics is critical in the incident analysis process to ascertain what occurred. Approaches for analyzing the persistent memory in industrial control devices are limited and almost nonexistent for volatile memory. This chapter proposes an automated methodology for programmable logic controller memory dump analysis using computer vision and deep learning techniques. The methodology converts the sequences of bytes in a programmable logic controller memory dump to red-green-blue pixels and employs a deep learning model that learns the underlying patterns and features of pre-labeled forensic artifacts in images and segments them into distinct regions. The trained model is employed to automatically segment new memory images and identify forensic artifacts. Evaluation of the methodology on a Schneider Electric Modicon M221 programmable logic controller under code injection and code modification attacks demonstrates its ability to detect attack artifacts in memory dumps.

Asmar Awad, Rima [ORNL] (ORCID:0000000233407742)↗

Automated algorithms to build active galactic nucleus classifiers

ABSTRACT We present a machine learning model to classify active galactic nuclei (AGNs) and galaxies (AGN-galaxy classifier) and a model to identify type 1 (optically unabsorbed) and type 2 (optically absorbed) AGN (type 1/2 classifier). We test tree-based algorithms, using training samples built from the X-ray Multi-Mirror Mission–Newton (XMM–Newton) catalogue and the Sloan Digital Sky Survey (SDSS), with labels derived from the SDSS survey. The performance was tested making use of simulations and of cross-validation techniques. With a set of features including spectroscopic redshifts and X-ray parameters connected to source properties (e.g. fluxes and extension), as well as features related to X-ray instrumental conditions, the precision and recall for AGN identification are 94 and 93 per cent, while the type 1/2 classifier has a precision of 74 per cent and a recall of 80 per cent for type 2 AGNs. The performance obtained with photometric redshifts is very similar to that achieved with spectroscopic redshifts in both test cases, while there is a decrease in performance when excluding redshifts. Our machine learning model trained on X-ray features can accurately identify AGN in extragalactic surveys. The type 1/2 classifier has a valuable performance for type 2 AGNs, but its ability to generalize without redshifts is hampered by the limited census of absorbed AGN at high redshift.

Falocco, S. (ORCID:0000000299841103)↗

Clinical Natural Language Processing for Radiation Oncology: A Review and Practical Primer

Natural language processing (NLP), which aims to convert human language into expressions that can be analyzed by computers, is one of the most rapidly developing and widely used technologies in the field of artificial intelligence. Natural language processing algorithms convert unstructured free text data into structured data that can be extracted and analyzed at scale. In medicine, this unlocking of the rich, expressive data within clinical free text in electronic medical records will help untap the full potential of big data for research and clinical purposes. Recent major NLP algorithmic advances have significantly improved the performance of these algorithms, leading to a surge in academic and industry interest in developing tools to automate information extraction and phenotyping from clinical texts. Thus, these technologies are poised to transform medical research and alter clinical practices in the future. Radiation oncology stands to benefit from NLP algorithms if they are appropriately developed and deployed, as they may enable advances such as automated inclusion of radiation therapy details into cancer registries, discovery of novel insights about cancer care, and improved patient data curation and presentation at the point of care. However, challenges remain before the full value of NLP is realized, such as the plethora of jargon specific to radiation oncology, nonstandard nomenclature, a lack of publicly available labeled data for model development, and interoperability limitations between radiation oncology data silos. Successful development and implementation of high quality and high value NLP models for radiation oncology will require close collaboration between computer scientists and the radiation oncology community. Here, we present a primer on artificial intelligence algorithms in general and NLP algorithms in particular; provide guidance on how to assess the performance of such algorithms; review prior research on NLP algorithms for oncology; and describe future avenues for NLP in radiation oncology research and clinics.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Production of GMP-Compliant Clinical Amounts of Copper-61 Radiopharmaceuticals from Liquid Targets

PET imaging has gained significant momentum in the last few years, especially in the area of oncology, with an increasing focus on metal radioisotopes owing to their versatile chemistry and favourable physical properties. Copper-61 (t1/2 = 3.33 h, 61% β+, Emax = 1.216 MeV) provides unique advantages versus the current clinical standard (i.e., gallium-68) even though, until now, no clinical amounts of 61Cu-based radiopharmaceuticals, other than thiosemicarbazone-based molecules, have been produced. This study aimed to establish a routine production, using a standard medical cyclotron, for a series of widely used somatostatin analogues, currently labelled with gallium-68, that could benefit from the improved characteristics of copper-61. We describe two possible routes to produce the radiopharmaceutical precursor, either from natural zinc or enriched zinc-64 liquid targets and further synthesis of [61Cu]Cu-DOTA-NOC, [61Cu]Cu-DOTA-TOC and [61Cu]Cu-DOTA-TATE with a fully automated GMP-compliant process. The production from enriched targets leads to twice the amount of activity (3.28 ± 0.41 GBq vs. 1.84 ± 0.24 GBq at EOB) and higher radionuclidic purity (99.97% vs. 98.49% at EOB). Our results demonstrate, for the first time, that clinical doses of 61Cu-based radiopharmaceuticals can easily be obtained in centres with a typical biomedical cyclotron optimised to produce 18F-based radiopharmaceuticals.

Fonseca, Alexandra I. (ORCID:000000031924178X)↗

Real-time tracking of structural evolution in 2D MXenes using theory-enhanced machine learning

In situ Electron Energy Loss Spectroscopy (EELS) combined with Transmission Electron Microscopy (TEM) has traditionally been pivotal for understanding how material processing choices affect local structure and composition. However, the ability to monitor and respond to ultrafast transient changes, now achievable with EELS and TEM, necessitates innovative analytical frameworks. Here, we introduce a machine learning (ML) framework tailored for the real-time assessment and characterization of in operando EELS Spectrum Images (EELS-SI). We focus on 2D MXenes as the sample material system, specifically targeting the understanding and control of their atomic-scale structural transformations that critically influence their electronic and optical properties. This approach requires fewer labeled training data points than typical deep learning classification methods. By integrating computationally generated structures of MXenes and experimental datasets into a unified latent space using Variational Autoencoders (VAE) in a unique training method, our framework accurately predicts structural evolutions at latencies pertinent to closed-loop processing within the TEM. This study presents a critical advancement in enabling automated, on-the-fly synthesis and characterization, significantly enhancing capabilities for materials discovery and the precision engineering of functional materials at the atomic scale.

47 OTHER INSTRUMENTATION↗

Real-time biomass feedstock particle quality detection using image analysis and machine vision

Abstract A common and costly challenge in the nascent biorefinery industry is the consistent handling and conveyance of biomass feedstock materials, which can vary widely in their chemical, physical, and mechanical properties. Solutions to cope with varying feedstock qualities will be required, including advanced process controls to adjust equipment and reject feedstocks that do not meet a quality standard. In this work, we present and evaluate methods to autonomously assess corn stover feedstock quality in real time and provide data to process controls with low-cost camera hardware. We explore the use of neural networks to classify feedstocks based on actual processing behavior and pixel matrix feature parameterization to further assess particle attributes that may explain the variable processing behavior. We used the pretrained ResNet neural network coupled with a gated recurrent unit (GRU) time-series classifier trained on our image data, resulting in binary classification of feedstock anomalies with favorable performance. The textural aspects of the image data were statistically analyzed to determine if the textural features were predictive of operational disruptions. The significant textural features were angular second moment, prominence, mean height of surface profile, mean resultant vector, shade, skewness, variation of the polar facet orientation, and direction of azimuthal facets. Expansion of these models is recommended across a wider variety of labeled feedstock images of different qualities and species to develop a more robust tool that may be deployed using low-cost cameras within biorefineries.

09 BIOMASS FUELS↗

Instance Segmentation for Direct Measurements of Satellites in Metal Powders and Automated Microstructural Characterization from Image Data

In this work, we propose instance segmentation as a useful tool for image analysis in materials science. Instance segmentation is an advanced technique in computer vision which generates individual segmentation masks for every object of interest that is recognized in an image. Using an out-of-the-box implementation of Mask R-CNN, instance segmentation is applied to images of metal powder particles produced through gas atomization. Leveraging transfer learning allows for the analysis to be conducted with a very small training set of labeled images. As well as providing another method for measuring the particle size distribution, we demonstrate the first direct measurements of the satellite content in powder samples. After analyzing the results for the labeled data dataset, the trained model was used to generate measurements for a much larger set of unlabeled images. The resulting particle size measurements showed reasonable agreement with laser scattering measurements. The satellite measurements were self-consistent and showed good agreement with the expected trends for different samples. Finally, we present a small case study showing how instance segmentation can be used to measure spheroidite content in the UltraHigh Carbon Steel DataBase, demonstrating the flexibility of the technique.

36 MATERIALS SCIENCE↗

Accurate and Data‐Efficient Micro X‐ray Diffraction Phase Identification Using Multitask Learning: Application to Hydrothermal Fluids

Traditional analysis of highly distorted micro X‐ray diffraction (μ‐XRD) patterns from hydrothermal fluid environments is a time‐consuming process, often requiring substantial data preprocessing and labeled experimental data. Herein, the potential of deep learning with a multitask learning (MTL) architecture to overcome these limitations is demonstrated. MTL models are trained to identify phase information in μ‐XRD patterns, minimizing the need for labeled experimental data and masking preprocessing steps. Notably, MTL models show superior accuracy compared to binary classification convolutional neural networks. Additionally, introducing a tailored cross‐entropy loss function improves MTL model performance. Most significantly, MTL models tuned to analyze raw and unmasked XRD patterns achieve close performance to models analyzing preprocessed data, with minimal accuracy differences. This work indicates that advanced deep learning architectures like MTL can automate arduous data handling tasks, streamline the analysis of distorted XRD patterns, and reduce the reliance on labor‐intensive experimental datasets.

97 MATHEMATICS AND COMPUTING↗

Graph neural network for neutrino physics event reconstruction

Liquid argon time projection chamber (LArTPC) detector technology offers a wealth of high-resolution information on particle interactions, and leveraging that information to its full potential requires sophisticated automated reconstruction techniques. Here, this article describes NUGRAPH 2, a graph neural network for low-level reconstruction of simulated neutrino interactions in a LArTPC detector. Simulated neutrino interactions in the MicroBooNE detector geometry are described as heterogeneous graphs, with energy depositions on each detector plane forming nodes on planar subgraphs. The network utilizes a multihead attention message-passing mechanism to perform background filtering and semantic labeling on these graph nodes, identifying those associated with the primary physics interaction with 98.0% efficiency and labeling them according to particle type with 94.9% efficiency. The network operates directly on detector observables across multiple two-dimensional representations but utilizes a three-dimensional-context-aware mechanism to encourage consistency between these representations. Model inference takes 0.12 s / event on a CPU and 0.005 s / event batched on a GPU. This architecture is designed to be a general-purpose solution for particle reconstruction in neutrino physics, with the potential for deployment across a broad range of detector technologies, and offers a core convolution engine that can be leveraged for a variety of tasks beyond the two described in this paper.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Automated theoretical chemical kinetics: Predicting the kinetics for the initial stages of pyrolysis

Large scale implementation of high level computational theoretical chemical kinetics offers the prospect for dramatically improving the fidelity of combustion chemical modeling. To facilitate such efforts, we have developed a suite of codes, collectively referred to as AutoMech, that allow for the automatic prediction of the kinetics for large sets of reactions via ab initio transition-state-theory based master-equation calculations. The primary input is simply the mechanism, a dictionary relating chemically identifiable species descriptors (e.g., SMILES or InChIs) to species labels in the mechanism, and a specification of the electronic structure and transition state theory models to be implemented. Here we illustrate the current utility of AutoMech through a study of the initial stages of pyrolysis for 3 sets of fuels: sequences of alkanes, alcohols, and aldehydes. For simplicity, the analysis focuses on abstractions from the fuel by H, CH 3 , and OH, and the decomposition of the resulting radicals. Altogether, there are a total of 166 input channels in these sets (more than 363 forward reactions when expanded to the full set of elementary reactions). The code successfully produces high quality rate estimates (with apparent uncertainties less than a factor of two in limited comparisons with experiment) for > 95% of these. For the radical decomposition reactions, the analysis includes predictions for the pressure dependence of the kinetics. This wide-ranging exploration illustrates (i) the effect of different levels of prediction on the expected accuracy, (ii) the branching between abstractions at different sites for different abstractors, (iii) the dependence of the rates on the chemical structure, and (iv) the variation in radical stabilities across chemical families. These results, as well as the demonstrated feasibility of the methodology, should find further utility in the development of accurate rate expressions for arbitrary fuels. (c) 2020 The Combustion Institute. Published by Elsevier Inc. All rights reserved.

Elliott, Sarah N.↗

Multitask Machine Learning of Collective Variables for Enhanced Sampling of Rare Events

Computing accurate reaction rates is a central challenge in computational chemistry and biology because of the high cost of free energy estimation with unbiased molecular dynamics. In this work, a data-driven machine learning algorithm is devised to learn collective variables with a multitask neural network, where a common upstream part reduces the high dimensionality of atomic configurations to a low dimensional latent space and separate downstream parts map the latent space to predictions of basin class labels and potential energies. Here, the resulting latent space is shown to be an effective low-dimensional representation, capturing the reaction progress and guiding effective umbrella sampling to obtain accurate free energy landscapes. This approach is successfully applied to model systems including a 5D Müller Brown model, a 5D three-well model, the alanine dipeptide in vacuum, and an Au(110) surface reconstruction unit reaction. It enables automated dimensionality reduction for energy controlled reactions in complex systems, offers a unified and data-efficient framework that can be trained with limited data, and outperforms single-task learning approaches, including autoencoders.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗