Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Reinforcement learning based automated history matching for improved hydrocarbon production forecast

History matching aims to find a numerical reservoir model that can be used to predict the reservoir performance. An engineer and model calibration (data inversion) method are required to adjust various parameters/properties of the numerical model in order to match the reservoir production history. In this study, we develop deep neural networks within the reinforcement learning framework to achieve automated history matching that will reduce engineers’ efforts, human bias, automatically and intelligently explore the parameter space, and remove the need of large set of labeled training data. To that end, a fast-marching-based reservoir simulator is encapsulated as an environment for the proposed reinforcement learning. The deep neural-network-based learning agent interacts with the reservoir simulator within reinforcement learning framework to achieve the automated history matching. Reinforcement learning techniques, such as discrete Deep Q Network and continuous Deep Deterministic Policy Gradients, are used toth, used to train the learning agents. The continuous actions enable the Deep Deterministic Policy Gradients to explore more states at each iteration in a a learning episode; consequently, a better history matching is achieved using this algorithm as compared to Deep Q Network. For simplified dual-target composite reservoir models, the best history-matching performances of the discrete and continuous learning methods in terms of normalized root mean square errors are 0.0447 and 0.0038, respectively. Furthermore, our study shows that continuous action space achieved by the deep deterministic policy gradient drastically outperforms deep Q network.

42 ENGINEERING↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

CyBERT: Cybersecurity Claim Classification by Fine-Tuning the BERT Language Model

We introduce CyBERT, a cybersecurity feature claims classifier based on bidirectional encoder representations from transformers and a key component in our semi-automated cybersecurity vetting for industrial control systems (ICS). To train CyBERT, we created a corpus of labeled sequences from ICS device documentation collected across a wide range of vendors and devices. This corpus provides the foundation for fine-tuning BERT’s language model, including a prediction-guided relabeling process. We propose an approach to obtain optimal hyperparameters, including the learning rate, the number of dense layers, and their configuration, to increase the accuracy of our classifier. Fine-tuning all hyperparameters of the resulting model led to an increase in classification accuracy from 76% obtained with BertForSequenceClassification’s original architecture to 94.4% obtained with CyBERT. Furthermore, we evaluated CyBERT for the impact of randomness in the initialization, training, and data-sampling phases. CyBERT demonstrated a standard deviation of ±0.6% during validation across 100 random seed values. Finally, we also compared the performance of CyBERT to other well-established language models including GPT2, ULMFiT, and ELMo, as well as neural network models such as CNN, LSTM, and BiLSTM. The results showed that CyBERT outperforms these models on the validation accuracy and the F1 score, validating CyBERT’s robustness and accuracy as a cybersecurity feature claims classifier.

97 MATHEMATICS AND COMPUTING↗

Automated Identification of Molecular Crystals’ Packing Motifs

Packing motifs—patterns in how molecules orient relative to one another in a crystal structure—are an important concept in many subdisciplines of materials science because of correlations observed between specific packing motifs and properties of interest. That said, packing motif data sets have remained small and noisy due to intensive manual labeling processes and insufficient labeling schemes. The most prominent labeling algorithms calculate relative interplanar angles of nearest neighbor molecules to determine the packing motif of a molecular crystal, but this simple approach can fail when neighbors are naively sampled isotropically around the crystal structure. To remedy this issue, here we propose an optimization algorithm, which rotates the molecular crystal structure to find representative molecules that inform the packing motif. We package this algorithm into an automated framework—Autopack—which both optimally rotates the crystal structure and labels the packing motif based on the appropriate neighboring molecules. In this work, we detail the Autopack framework and its performance, which shows improvements compared to previous state-of-the-art labeling methods, providing the first quantitative point of comparison for packing motif labeling algorithms. Furthermore, using Autopack (available at https://ipo.llnl.gov/technologies/software/autopack), we perform the first large-scale study of potential relationships between chemicals’ compositions and packing motifs, which shows that these relationships are more complex than previously hypothesized from studies that used only tens of polycyclic aromatic hydrocarbon molecules. Autopack’s capabilities help pose next steps for crystal engineering research focusing not only on a molecule’s adoption of a specific packing motif but also on new structure–property relationships.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Label Assist: Personalized Travel Models for Longitudinal Data Collection

Understanding travel behavior is crucial to transportation decarbonization. OpenPATH is an open-source mobility platform which collects and analyzes human travel behavior at the individual level. The mobile application passively senses trips and prompts users to label them. However, users find the labeling process burdensome; less than half the trips are typically labeled, making much of the data unusable in aggregate analyses of mobility patterns. Prior work has addressed the response fatigue challenge through automated mode inference using sensor data, but sensors cannot capture all aspects of travel behavior. We explore an alternative approach in which we leverage prior user input to predict travel choices in novel trips. We first explore trip clustering methods and develop a novel two-step pipeline using DBSCAN and SVMs to extract realistic geospatial clusters. We then propose two strategies to predict trip labels: (i) clustering trips and extrapolating labels for similar trips, and (ii) random forest classification. The random forest approach is able to achieve - $70-80% accuracy (purpose: 72%, mode: 79%, replaced mode: 81%). These novel approaches to trip classification allow us to increase the rate of user labeling by suggesting predicted labels to be verified by the user. Unlabeled trips can also contribute to aggregate analyses, using label predictions and their associated confidences as a substitute. While there exist other travel survey apps with the ability to infer travel choices, to our knowledge, this is the first paper to describe such a supervised system and rigorously evaluate it.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Sample Preparation Methods for Targeted Single-Cell Proteomics

We compared three cell isolation and two proteomic sample preparation methods for single-cell and near-single-cell analysis. Whole blood was used to quantify hemoglobin (Hb) and glycated-Hb (gly-Hb) in erythrocytes using targeted mass spectrometry and stable isotope-labeled standard peptides. Each method differed in cell isolation and sample preparation as follows: 1) FACS and automated preparation in one-pot for trace samples (autoPOTS); 2) limited dilution via microscopy and a novel rapid one-pot sample preparation method that circumvented the need for the solid-phase extraction, low-volume liquid handling instrumentation and humidified incubation chamber; and 3) CellenONE-based cell isolation and the same one-pot sample preparation method used for limited dilution. Only the CellenONE device routinely isolated single-cells from which Hb was measured to be 540–660 amol per red blood cell (RBC), which was comparable to the calculated SI reference range for mean corpuscular hemoglobin (390–540 amol/RBC). FACSAria sorter and limited dilution could routinely isolate single-digit cell numbers, to reliably quantify CMV-Hb heterogeneity. Finally, we observed that repeated measures, using 5–25 RBCs obtained from N = 10 blood donors, could be used as an alternative and more efficient strategy than single RBC analysis to measure protein heterogeneity, which revealed multimodal distribution, unique for each individual.

59 BASIC BIOLOGICAL SCIENCES↗

Self-supervised Representation Learning for Astronomical Images

Sky surveys are the largest data generators in astronomy, making automated tools for extracting meaningful scientific information an absolute necessity. We show that, without the need for labels, self-supervised learning recovers representations of sky survey images that are semantically useful for a variety of scientific tasks. These representations can be directly used as features, or fine-tuned, to outperform supervised methods trained only on labeled data. We apply a contrastive learning framework on multiband galaxy photometry from the Sloan Digital Sky Survey (SDSS), to learn image representations. We then use them for galaxy morphology classification and fine-tune them for photometric redshift estimation, using labels from the Galaxy Zoo 2 data set and SDSS spectroscopy. In both downstream tasks, using the same learned representations, we outperform the supervised state-of-the-art results, and we show that our approach can achieve the accuracy of supervised models while using 2-4 times fewer labels for training. The codes, trained models, and data can be found at https://portal.nersc.gov/project/dasrepo/self-supervised-learning-sdss.

79 ASTRONOMY AND ASTROPHYSICS↗

ODI notebook additive_manufacturing_video_2022

Additive Manufacturing, video dataset Two-photon lithography (TPL) is a widely used 3D nanoprinting technique that uses laser light to create objects. Challenges to large-scale adoption of this additive manufacturing method include identifying light dosage parameters and monitoring during fabrication. A research team from LLNL, Iowa State University, and Georgia Tech is applying machine learning models to tackle these challenges-i.e., accelerate the process of identifying optimal light dosage parameters and automate the detection of part quality. Funded by LLNL's Laboratory Directed Research and Development Program, the project team has curated a video dataset of TPL processes for parameters such as light dosages, photo-curable resins, and structures. Both raw and labeled versions of the datasets are available on the links in the Open Data Initiative page. The code uses the labeled dataset. Notebook compiled by Nisha Mulakken (mulakken1@llnl.gov) for LLNL Open Data Initiative, Summer 2022. Original code provided by research team. Publications: X.Y. Lee, S.K. Saha, S. Sarkar, B. Giera. "Automated detection of part quality during two-photon lithography via deep learning." Additive Manufacturing 36, December 2020: doi.org/10.1016/j.addma.2020.101444 X.Y. Lee, S.K. Saha, S. Sarkar, B. Giera. "wo Photon lithography additive manufacturing: Video dataset of parameter sweep of light dosages, photo-curable resins, and structures." Data in Brief 32, October 2020. doi.org/10.1016/j.dib.2020.106119.

Mulakken, NishaJ↗

Automated Grain Boundary (GB) Segmentation and Microstructural Analysis in 347H Stainless Steel Using Deep Learning and Multimodal Microscopy

Austenitic 347H stainless steel offers superior mechanical properties and corrosion resistance required for extreme operating conditions such as high temperature. The change in microstructure due to composition and process variations is expected to impact material properties. Identifying microstructural features such as grain boundaries thus becomes an important task in the process-microstructure-properties loop. Applying convolutional neural network (CNN)-based deep learning models is a powerful technique to detect features from material micrographs in an automated manner. In contrast to microstructural classification, supervised CNN models for segmentation tasks require pixel-wise annotation labels. However, manual labeling of the images for the segmentation task poses a major bottleneck for generating training data and labels in a reliable and reproducible way within a reasonable timeframe. Microstructural characterization especially needs to be expedited for faster material discovery by changing alloy compositions. Here, in this study, we attempt to overcome such limitations by utilizing multimodal microscopy to generate labels directly instead of manual labeling. We combine scanning electron microscopy images of 347H stainless steel as training data and electron backscatter diffraction micrographs as pixel-wise labels for grain boundary detection as a semantic segmentation task. The viability of our method is evaluated by considering a set of deep CNN architectures. We demonstrate that despite producing instrumentation drift during data collection between two modes of microscopy, this method performs comparably to similar segmentation tasks that used manual labeling. Additionally, we find that naïve pixel-wise segmentation results in small gaps and missing boundaries in the predicted grain boundary map. By incorporating topological information during model training, the connectivity of the grain boundary network and segmentation performance is improved. Finally, our approach is validated by accurate computation on downstream tasks of predicting the underlying grain morphology distributions which are the ultimate quantities of interest for microstructural characterization.

36 MATERIALS SCIENCE↗

Automated Registration of Vector Data to Overhead Imagery

The availability of open source, remote sensing-derived vector data has increased exponentially in recent years. Unfortunately, these vector data are rarely made available with the corresponding source images from which objects were extracted. As such, these derived datasets are commonly used in combination with target images which differ from the source. Satellite viewing geometry can cause objects, extracted from a source image, to appear shifted when overlaid with a target image. Whether for purely cartographic purposes, reusability of preexisting training labels, or any spatial analysis where spatial correspondence between vector and image are required, the following paragraphs outline a method to address this challenge through the automated registration of vector data to overhead imagery, including existing literature regarding this challenge, followed by a case study of Sioux Falls, South Dakota.

McKee, Jacob↗

An automated workflow that generates atom mappings for large‐scale metabolic models and its application to Arabidopsis thaliana

SUMMARY Quantification of reaction fluxes of metabolic networks can help us understand how the integration of different metabolic pathways determines cellular functions. Yet, intracellular fluxes cannot be measured directly but are estimated with metabolic flux analysis (MFA), which relies on the patterns of isotope labeling of metabolites in the network. The application of MFA also requires a stoichiometric model with atom mappings that are currently not available for the majority of large‐scale metabolic network models, particularly of plants. While automated approaches such as the Reaction Decoder Toolkit (RDT) can produce atom mappings for individual reactions, tracing the flow of individual atoms of the entire reactions across a metabolic model remains challenging. Here we establish an automated workflow to obtain reliable atom mappings for large‐scale metabolic models by refining the outcome of RDT, and apply the workflow to metabolic models of Arabidopsis thaliana . We demonstrate the accuracy of RDT through a comparative analysis with atom mappings from a large database of biochemical reactions, MetaCyc. We further show the utility of our automated workflow by simulating 15 N isotope enrichment and identifying nitrogen (N)‐containing metabolites which show enrichment patterns that are informative for flux estimation in future 15 N‐MFA studies of A. thaliana . The automated workflow established in this study can be readily expanded to other species for which metabolic models have been established and the resulting atom mappings will facilitate MFA and graph‐theoretic structural analyses with large‐scale metabolic networks.

59 BASIC BIOLOGICAL SCIENCES↗

Ionic Liquid Aided [ 11 C]CO Fixation for Synthesis of 11 C‐carbonyls

Tributyl(ethyl)phosphonium oxopentenolate ([P 4442 ][Pen]) is an ionic liquid developed to capture CO and has shown ability to catalyze carbonylation reactions in organic chemistry. Carbon-11 ( 11 C, t 1/2 =20.4 min) labeled CO is a highly versatile building block for the synthesis of positron emission tomography (PET) radiotracers that are applied for medical imaging. The use of [ 11 C]CO is limited by its low solubility in organic solvents. Herein, we report a proof-of-concept study evaluating a new method to prepare 11 C-labeled amides, ureas and carbamates via reaction of [ 11 C]CO in [P 4442 ][Pen] and applied for fully automated radiosyntheses of Bruton's tyrosine kinase inhibitors, [ 11 C]evobrutinib and [ 11 C]ibrutinib.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting solid state material platforms for quantum technologies

Semiconductor materials provide a compelling platform for quantum technologies (QT). However, identifying promising material hosts among the plethora of candidates is a major challenge. Therefore, we have developed a framework for the automated discovery of semiconductor platforms for QT using material informatics and machine learning methods. Different approaches were implemented to label data for training the supervised machine learning (ML) algorithms logistic regression, decision trees, random forests and gradient boosting. We find that an empirical approach relying exclusively on findings from the literature yields a clear separation between predicted suitable and unsuitable candidates. In contrast to expectations from the literature focusing on band gap and ionic character as important properties for QT compatibility, the ML methods highlight features related to symmetry and crystal structure, including bond length, orientation and radial distribution, as influential when predicting a material as suitable for QT.

36 MATERIALS SCIENCE↗

Panel-Segmentation: A Python Package for Automated Solar Array Metadata Extraction Using Satellite Imagery

The NREL Python Panel-Segmentation package is a toolkit that automates the process of extracting accurate and valuable metadata related to solar array installations, using publicly available Google Maps satellite imagery. Previously published work includes automated azimuth estimation for individual solar installations in satellite images. Our continued research focuses on automated detection and classification of solar installation mounting configuration (tracking or fixed-tilt; rooftop, ground, or carport). Specifically, a Faster-RCNN Resnet-50 feature pyramid network (FPN) model was trained and validated on 862 manually labeled satellite images. This model was used to perform object detection on satellite imagery, locating and classifying individual solar installations' mounting configuration and type. Model results showed a mean average precision score (mAP) of 77.79%, with the model strongest at detecting fixed-tilt ground mount and fixed-tilt carport installations. The object detection model and its outputs have been incorporated into the Panel-Segmentation package's automated metadata extraction pipeline, which returns the mounting configuration and azimuth for individual solar arrays in satellite imagery. The complete image data set with labels has been released on the U.S. Department of Energy (DOE) DuraMAT DataHub, to encourage further research in this area.

deep learning↗

Computational Framework for Machine-Learning-Enabled 13 C Fluxomics

13 C metabolic flux analysis (MFA) has emerged as a powerful tool for synthetic biology. This optimization-based approach suffers long computation time and unstable solutions depending on the initial guess. Here, we develop a machine-learning-based framework for 13 C fluxomics. Specifically, training and test data sets are generated by metabolic network decomposition and flux sampling, in which flux ratios at metabolic nodes and simulated labeling patterns of metabolites are used as training targets and features, respectively. To improve prediction accuracy and simplify the model, automated processes are developed for flux ratio selection based on solvability and feature screening based on importance. We found that predictive performance can be significantly improved using both amino acids and central carbon metabolites in comparison with amino acids alone. Together with measured external fluxes, the predicted flux ratios determine the mass balance system, yielding global flux distributions. This approach is validated by flux estimation using both simulated and experimental data in comparison with canonical 13 C MFA. The approach represents a reliable fluxomics method readily applicable to high-throughput metabolic phenotyping, which highlights the advances of intelligent learning algorithms in synthetic biology, specifically in the Test and Learn stage of the Design-Build-Test-Learn cycle.

13C metabolic flux analysis↗

The CanBikeCO Mini Pilot: Preliminary Results and Lessons Learned

In fall 2020, the Colorado Energy Office, as part of the State of Colorado’s “Can Do Colorado” initiative, initiated a project aimed at encouraging energy-efficient transportation during the COVID-19 pandemic. The initial mini-pilot provided e-bikes to 13 low-income households under an individual ownership model. This report assesses the impact of providing this additional mobility option on the travel behavior of participants. It also outlines the lessons learned from deploying a continuous monitoring platform to track the travel behavior. These lessons will influence the evaluation component for the full pilot, which will cover multiple geographic regions, starting in summer 2021, and run for 2 years. The continuous data collection was enabled by a customized version of the open-source e-mission platform, called CanBikeCO, configured with a behavioral gamification feature. The Colorado Energy Office used this system to collect a unique data set consisting of 3 months of partially automated travel diaries, combining sensed and surveyed data and linked with demographic information, from 12 participants. The data collection process worked well overall: users generally liked the app, appreciated the game, and did not complain about battery life. The long tracking period introduced behavioral challenges in user engagement, which we plan to address using repeated patterns and automated status checks for the full pilot. The analysis results, based on the subset of trips with user-reported labels (68%), indicate that the e-bike was the dominant commute mode share (31%), in sharp contrast to the census bicycle commute mode share (<1%). E-bike trips primarily replaced single-occupancy vehicle (SOV) trips (28%), followed closely by walking (24%) and regular bike (20%). The nonmotorized mode replacement corresponds to lower travel time and increased productivity enabled by the program. The emissions impact analysis of the program, computed using trip-level energy intensity factors, indicates savings of 1,367 lbs. of CO 2 . Although the results are strongly positive, the narrow demographic profile of study participants, their limited mobility alternatives, and nonuniform labeling indicate caution in broader interpretation. These preliminary results do suggest that such programs, supported by real-time education and support from program managers, can simultaneously meet equity and sustainability goals. The planned full pilot, addressing the data collection challenges and broadening the geographic scope, will provide additional insights into the generality of this approach.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Challenges in Automated Detection of COVID-19 Misinformation

The COVID-19 pandemic has made the dangers of the spread of misinformation obvious but despite much global effort to curbing its spread, fake information about the pandemic keeps proliferating. In this paper, we address the development of automated methods for verification of claims about COVID-19 and discuss the challenges associated with this task. We focus on labeled data collection, limitations of existing models, and difficulties of applying misinformation detection models in practical applications. Our initial analysis indicates label imbalance may be a particular challenge for developing claim verification models and we discuss options for alleviating this issue.

Herrmannova, Dasha↗

The Clean Energy Mortgage

Explore the source record for details and available documents.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗