Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1 %, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "hits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. Preliminary experiments show that the presence of the fluorescent probe does not affect the nucleation process or the quality of the X-ray data obtained.

Pusey, Marc L.↗

Automated Identification of Molecular Crystals’ Packing Motifs

Packing motifs—patterns in how molecules orient relative to one another in a crystal structure—are an important concept in many subdisciplines of materials science because of correlations observed between specific packing motifs and properties of interest. That said, packing motif data sets have remained small and noisy due to intensive manual labeling processes and insufficient labeling schemes. The most prominent labeling algorithms calculate relative interplanar angles of nearest neighbor molecules to determine the packing motif of a molecular crystal, but this simple approach can fail when neighbors are naively sampled isotropically around the crystal structure. To remedy this issue, here we propose an optimization algorithm, which rotates the molecular crystal structure to find representative molecules that inform the packing motif. We package this algorithm into an automated framework—Autopack—which both optimally rotates the crystal structure and labels the packing motif based on the appropriate neighboring molecules. In this work, we detail the Autopack framework and its performance, which shows improvements compared to previous state-of-the-art labeling methods, providing the first quantitative point of comparison for packing motif labeling algorithms. Furthermore, using Autopack (available at https://ipo.llnl.gov/technologies/software/autopack), we perform the first large-scale study of potential relationships between chemicals’ compositions and packing motifs, which shows that these relationships are more complex than previously hypothesized from studies that used only tens of polycyclic aromatic hydrocarbon molecules. Autopack’s capabilities help pose next steps for crystal engineering research focusing not only on a molecule’s adoption of a specific packing motif but also on new structure–property relationships.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Label Assist: Personalized Travel Models for Longitudinal Data Collection

Understanding travel behavior is crucial to transportation decarbonization. OpenPATH is an open-source mobility platform which collects and analyzes human travel behavior at the individual level. The mobile application passively senses trips and prompts users to label them. However, users find the labeling process burdensome; less than half the trips are typically labeled, making much of the data unusable in aggregate analyses of mobility patterns. Prior work has addressed the response fatigue challenge through automated mode inference using sensor data, but sensors cannot capture all aspects of travel behavior. We explore an alternative approach in which we leverage prior user input to predict travel choices in novel trips. We first explore trip clustering methods and develop a novel two-step pipeline using DBSCAN and SVMs to extract realistic geospatial clusters. We then propose two strategies to predict trip labels: (i) clustering trips and extrapolating labels for similar trips, and (ii) random forest classification. The random forest approach is able to achieve - $70-80% accuracy (purpose: 72%, mode: 79%, replaced mode: 81%). These novel approaches to trip classification allow us to increase the rate of user labeling by suggesting predicted labels to be verified by the user. Unlabeled trips can also contribute to aggregate analyses, using label predictions and their associated confidences as a substitute. While there exist other travel survey apps with the ability to infer travel choices, to our knowledge, this is the first paper to describe such a supervised system and rigorously evaluate it.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Sample Preparation Methods for Targeted Single-Cell Proteomics

We compared three cell isolation and two proteomic sample preparation methods for single-cell and near-single-cell analysis. Whole blood was used to quantify hemoglobin (Hb) and glycated-Hb (gly-Hb) in erythrocytes using targeted mass spectrometry and stable isotope-labeled standard peptides. Each method differed in cell isolation and sample preparation as follows: 1) FACS and automated preparation in one-pot for trace samples (autoPOTS); 2) limited dilution via microscopy and a novel rapid one-pot sample preparation method that circumvented the need for the solid-phase extraction, low-volume liquid handling instrumentation and humidified incubation chamber; and 3) CellenONE-based cell isolation and the same one-pot sample preparation method used for limited dilution. Only the CellenONE device routinely isolated single-cells from which Hb was measured to be 540–660 amol per red blood cell (RBC), which was comparable to the calculated SI reference range for mean corpuscular hemoglobin (390–540 amol/RBC). FACSAria sorter and limited dilution could routinely isolate single-digit cell numbers, to reliably quantify CMV-Hb heterogeneity. Finally, we observed that repeated measures, using 5–25 RBCs obtained from N = 10 blood donors, could be used as an alternative and more efficient strategy than single RBC analysis to measure protein heterogeneity, which revealed multimodal distribution, unique for each individual.

59 BASIC BIOLOGICAL SCIENCES↗

Self-supervised Representation Learning for Astronomical Images

Sky surveys are the largest data generators in astronomy, making automated tools for extracting meaningful scientific information an absolute necessity. We show that, without the need for labels, self-supervised learning recovers representations of sky survey images that are semantically useful for a variety of scientific tasks. These representations can be directly used as features, or fine-tuned, to outperform supervised methods trained only on labeled data. We apply a contrastive learning framework on multiband galaxy photometry from the Sloan Digital Sky Survey (SDSS), to learn image representations. We then use them for galaxy morphology classification and fine-tune them for photometric redshift estimation, using labels from the Galaxy Zoo 2 data set and SDSS spectroscopy. In both downstream tasks, using the same learned representations, we outperform the supervised state-of-the-art results, and we show that our approach can achieve the accuracy of supervised models while using 2-4 times fewer labels for training. The codes, trained models, and data can be found at https://portal.nersc.gov/project/dasrepo/self-supervised-learning-sdss.

79 ASTRONOMY AND ASTROPHYSICS↗

Automatically Finding Ship-Tracks to Enable Large-Scale Analysis of Aerosol-Cloud Interactions

Ship tracks appear as long winding linear features in satellite images and are produced by aerosols from ship exhausts changing low cloud properties. They are one of the best examples of aerosol‐cloud interaction experiments. However, manually finding ship tracks from satellite data on a large scale is prohibitively costly while a large number of samples are required to improve our understanding. Here we train a deep neural network to automate finding ship tracks. The neural network model generalizes well as it not only finds ship tracks labeled by human experts but also detects those that are occasionally missed by humans. It finds more ship tracks than all previous studies combined and produces a map of ship track distributions off the California coast that matches well with known shipping traffic. Our technique will enable studying aerosol effects on low clouds using ship tracks on a large scale, which will potentially narrow the uncertainty of the aerosol‐cloud interactions.

aerosol cloud interactions↗

ODI notebook additive_manufacturing_video_2022

Additive Manufacturing, video dataset Two-photon lithography (TPL) is a widely used 3D nanoprinting technique that uses laser light to create objects. Challenges to large-scale adoption of this additive manufacturing method include identifying light dosage parameters and monitoring during fabrication. A research team from LLNL, Iowa State University, and Georgia Tech is applying machine learning models to tackle these challenges-i.e., accelerate the process of identifying optimal light dosage parameters and automate the detection of part quality. Funded by LLNL's Laboratory Directed Research and Development Program, the project team has curated a video dataset of TPL processes for parameters such as light dosages, photo-curable resins, and structures. Both raw and labeled versions of the datasets are available on the links in the Open Data Initiative page. The code uses the labeled dataset. Notebook compiled by Nisha Mulakken (mulakken1@llnl.gov) for LLNL Open Data Initiative, Summer 2022. Original code provided by research team. Publications: X.Y. Lee, S.K. Saha, S. Sarkar, B. Giera. "Automated detection of part quality during two-photon lithography via deep learning." Additive Manufacturing 36, December 2020: doi.org/10.1016/j.addma.2020.101444 X.Y. Lee, S.K. Saha, S. Sarkar, B. Giera. "wo Photon lithography additive manufacturing: Video dataset of parameter sweep of light dosages, photo-curable resins, and structures." Data in Brief 32, October 2020. doi.org/10.1016/j.dib.2020.106119.

Mulakken, NishaJ↗

Automated Grain Boundary (GB) Segmentation and Microstructural Analysis in 347H Stainless Steel Using Deep Learning and Multimodal Microscopy

Austenitic 347H stainless steel offers superior mechanical properties and corrosion resistance required for extreme operating conditions such as high temperature. The change in microstructure due to composition and process variations is expected to impact material properties. Identifying microstructural features such as grain boundaries thus becomes an important task in the process-microstructure-properties loop. Applying convolutional neural network (CNN)-based deep learning models is a powerful technique to detect features from material micrographs in an automated manner. In contrast to microstructural classification, supervised CNN models for segmentation tasks require pixel-wise annotation labels. However, manual labeling of the images for the segmentation task poses a major bottleneck for generating training data and labels in a reliable and reproducible way within a reasonable timeframe. Microstructural characterization especially needs to be expedited for faster material discovery by changing alloy compositions. Here, in this study, we attempt to overcome such limitations by utilizing multimodal microscopy to generate labels directly instead of manual labeling. We combine scanning electron microscopy images of 347H stainless steel as training data and electron backscatter diffraction micrographs as pixel-wise labels for grain boundary detection as a semantic segmentation task. The viability of our method is evaluated by considering a set of deep CNN architectures. We demonstrate that despite producing instrumentation drift during data collection between two modes of microscopy, this method performs comparably to similar segmentation tasks that used manual labeling. Additionally, we find that naïve pixel-wise segmentation results in small gaps and missing boundaries in the predicted grain boundary map. By incorporating topological information during model training, the connectivity of the grain boundary network and segmentation performance is improved. Finally, our approach is validated by accurate computation on downstream tasks of predicting the underlying grain morphology distributions which are the ultimate quantities of interest for microstructural characterization.

36 MATERIALS SCIENCE↗

Automated Microbial Metabolism Laboratory

Development of the automated microbial metabolism laboratory (AMML) concept is reported. The focus of effort of AMML was on the advanced labeled release experiment. Labeled substrates, inhibitors, and temperatures were investigated to establish a comparative biochemical profile. Profiles at three time intervals on soil and pure cultures of bacteria isolated from soil were prepared to establish a complete library. The development of a strategy for the return of a soil sample from Mars is also reported.

Source record↗

Automated Registration of Vector Data to Overhead Imagery

The availability of open source, remote sensing-derived vector data has increased exponentially in recent years. Unfortunately, these vector data are rarely made available with the corresponding source images from which objects were extracted. As such, these derived datasets are commonly used in combination with target images which differ from the source. Satellite viewing geometry can cause objects, extracted from a source image, to appear shifted when overlaid with a target image. Whether for purely cartographic purposes, reusability of preexisting training labels, or any spatial analysis where spatial correspondence between vector and image are required, the following paragraphs outline a method to address this challenge through the automated registration of vector data to overhead imagery, including existing literature regarding this challenge, followed by a case study of Sioux Falls, South Dakota.

McKee, Jacob↗

An automated workflow that generates atom mappings for large‐scale metabolic models and its application to Arabidopsis thaliana

SUMMARY Quantification of reaction fluxes of metabolic networks can help us understand how the integration of different metabolic pathways determines cellular functions. Yet, intracellular fluxes cannot be measured directly but are estimated with metabolic flux analysis (MFA), which relies on the patterns of isotope labeling of metabolites in the network. The application of MFA also requires a stoichiometric model with atom mappings that are currently not available for the majority of large‐scale metabolic network models, particularly of plants. While automated approaches such as the Reaction Decoder Toolkit (RDT) can produce atom mappings for individual reactions, tracing the flow of individual atoms of the entire reactions across a metabolic model remains challenging. Here we establish an automated workflow to obtain reliable atom mappings for large‐scale metabolic models by refining the outcome of RDT, and apply the workflow to metabolic models of Arabidopsis thaliana . We demonstrate the accuracy of RDT through a comparative analysis with atom mappings from a large database of biochemical reactions, MetaCyc. We further show the utility of our automated workflow by simulating 15 N isotope enrichment and identifying nitrogen (N)‐containing metabolites which show enrichment patterns that are informative for flux estimation in future 15 N‐MFA studies of A. thaliana . The automated workflow established in this study can be readily expanded to other species for which metabolic models have been established and the resulting atom mappings will facilitate MFA and graph‐theoretic structural analyses with large‐scale metabolic networks.

59 BASIC BIOLOGICAL SCIENCES↗

Overcoming small minirhizotron datasets using transfer learning

Minirhizotron technology is widely used to study root growth and development. Yet, standard approaches for tracing roots in minirhiztron imagery is extremely tedious and time consuming. Machine learning approaches can help to automate this task. However, lack of enough annotated training data is a major limitation for the application of machine learning methods. Transfer learning is a useful technique to help with training when available datasets are limited. In this paper, we investigated the effect of pre-trained features from the massives-cale, irrelevant ImageNet dataset and a relatively moderate-scale, but relevant peanut root dataset on switchgrass root imagery segmentation applications. We compiled two minirhizotron image datasets to accomplish this study: one with 17,550 peanut root images and another with 28 switchgrass root images. Both datasets were paired with manually labeled ground truth masks. Deep neural networks based on the U-net architecture were used with different pre-trained features as initialization for automated, precise pixel-wise root segmentation in minirhizotron imagery. We observed that features pre-trained on a closely related but relatively moderate size dataset like our peanut dataset were more effective than features pre-trained on the large but unrelated ImageNet dataset. Here, we achieved high quality segmentation on peanut root dataset with 99.04% accuracy at the pixel-level and overcame errors in human-labeled ground truth masks. By applying transfer learning technique on limited switchgrass dataset with features pre-trained on peanut dataset, we obtained 99% segmentation accuracy in switchgrass imagery using only 21 images for training (fine tuning). Furthermore, the peanut pre-trained features can help the model converge faster and have much more stable performance.

59 BASIC BIOLOGICAL SCIENCES↗

Ionic Liquid Aided [ 11 C]CO Fixation for Synthesis of 11 C‐carbonyls

Tributyl(ethyl)phosphonium oxopentenolate ([P 4442 ][Pen]) is an ionic liquid developed to capture CO and has shown ability to catalyze carbonylation reactions in organic chemistry. Carbon-11 ( 11 C, t 1/2 =20.4 min) labeled CO is a highly versatile building block for the synthesis of positron emission tomography (PET) radiotracers that are applied for medical imaging. The use of [ 11 C]CO is limited by its low solubility in organic solvents. Herein, we report a proof-of-concept study evaluating a new method to prepare 11 C-labeled amides, ureas and carbamates via reaction of [ 11 C]CO in [P 4442 ][Pen] and applied for fully automated radiosyntheses of Bruton's tyrosine kinase inhibitors, [ 11 C]evobrutinib and [ 11 C]ibrutinib.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predicting solid state material platforms for quantum technologies

Semiconductor materials provide a compelling platform for quantum technologies (QT). However, identifying promising material hosts among the plethora of candidates is a major challenge. Therefore, we have developed a framework for the automated discovery of semiconductor platforms for QT using material informatics and machine learning methods. Different approaches were implemented to label data for training the supervised machine learning (ML) algorithms logistic regression, decision trees, random forests and gradient boosting. We find that an empirical approach relying exclusively on findings from the literature yields a clear separation between predicted suitable and unsuitable candidates. In contrast to expectations from the literature focusing on band gap and ionic character as important properties for QT compatibility, the ML methods highlight features related to symmetry and crystal structure, including bond length, orientation and radial distribution, as influential when predicting a material as suitable for QT.

36 MATERIALS SCIENCE↗

Crop identification studies using Landsat data Separation of barley from other spring small grains and corn and soybean decision logic

Two labeling procedures were developed which identify various agricultural crops through the use of Landsat data. One procedure separates barley from other spring small grains, and the other identifies corn and soybeans. For both procedures, a minimum data set (critical acquisition time) has been designated. Landsat data in both image format and various graphic displays were used along with ancillary data to obtain information which aided in labeling the spectral signatures. The corn and soybean procedure also employed a structured decision logic. Test results for the barley separation procedure emphasized the importance of obtaining a critical acquisition and showed some success especially in areas where spring crops followed the expected growth patterns. Two tests of the corn and soybean procedure produced good labeling accuracies. Problems with the procedure were easy to identify, and some solutions were implemented for the second test. Automation of various parts of the procedure and extension to other crops and regions were recommended.

Dailey, C. L.↗

Panel-Segmentation: A Python Package for Automated Solar Array Metadata Extraction Using Satellite Imagery

The NREL Python Panel-Segmentation package is a toolkit that automates the process of extracting accurate and valuable metadata related to solar array installations, using publicly available Google Maps satellite imagery. Previously published work includes automated azimuth estimation for individual solar installations in satellite images. Our continued research focuses on automated detection and classification of solar installation mounting configuration (tracking or fixed-tilt; rooftop, ground, or carport). Specifically, a Faster-RCNN Resnet-50 feature pyramid network (FPN) model was trained and validated on 862 manually labeled satellite images. This model was used to perform object detection on satellite imagery, locating and classifying individual solar installations' mounting configuration and type. Model results showed a mean average precision score (mAP) of 77.79%, with the model strongest at detecting fixed-tilt ground mount and fixed-tilt carport installations. The object detection model and its outputs have been incorporated into the Panel-Segmentation package's automated metadata extraction pipeline, which returns the mounting configuration and azimuth for individual solar arrays in satellite imagery. The complete image data set with labels has been released on the U.S. Department of Energy (DOE) DuraMAT DataHub, to encourage further research in this area.

deep learning↗

Computational Framework for Machine-Learning-Enabled 13 C Fluxomics

13 C metabolic flux analysis (MFA) has emerged as a powerful tool for synthetic biology. This optimization-based approach suffers long computation time and unstable solutions depending on the initial guess. Here, we develop a machine-learning-based framework for 13 C fluxomics. Specifically, training and test data sets are generated by metabolic network decomposition and flux sampling, in which flux ratios at metabolic nodes and simulated labeling patterns of metabolites are used as training targets and features, respectively. To improve prediction accuracy and simplify the model, automated processes are developed for flux ratio selection based on solvability and feature screening based on importance. We found that predictive performance can be significantly improved using both amino acids and central carbon metabolites in comparison with amino acids alone. Together with measured external fluxes, the predicted flux ratios determine the mass balance system, yielding global flux distributions. This approach is validated by flux estimation using both simulated and experimental data in comparison with canonical 13 C MFA. The approach represents a reliable fluxomics method readily applicable to high-throughput metabolic phenotyping, which highlights the advances of intelligent learning algorithms in synthetic biology, specifically in the Test and Learn stage of the Design-Build-Test-Learn cycle.

13C metabolic flux analysis↗