Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Image Labeler: A Web Interface to Catalog Earth Science Events

Advances in machine learning (ML) have made it possible to automatically detect Earth science phenomena from satellite imagery. While useful, ML algorithms typically require an extensive dataset containing labeled images for training. Systematic labeling and management of such datasets is quite cumbersome. With this in mind, we present the Image Labeler. Image Labeler is a fast and scalable cloud-based tool that facilitates the rapid development of Earth science event databases, in order to aid automated ML-based image classification.

Case Study↗

Characterizing Interference in Radio Astronomy Observations through Active and Unsupervised Learning

In the process of observing signals from astronomical sources, radio astronomers must mitigate the effects of man-made radio sources such as cell phones, satellites, aircraft, and observatory equipment. Radio frequency interference (RFI) often occurs as short bursts (< 1 ms) across a broad range of frequencies, and can be confused with signals from sources of interest such as pulsars. With ever-increasing volumes of data being produced by observatories, automated strategies are required to detect, classify, and characterize these short “transient” RFI events. We investigate an active learning approach in which an astronomer labels events that are most confusing to a classifier, minimizing the human effort required for classification. We also explore the use of unsupervised clustering techniques, which automatically group events into classes without user input. We apply these techniques to data from the Parkes Multibeam Pulsar Survey to characterize several million detected RFI events from over a thousand hours of observation

Doran, G.↗

Iterative self-organizing SCEne-LEvel sampling (ISOSCELES) for large-scale building extraction

Convolutional neural networks (CNN) provide state-of-the-art performance in many computer vision tasks, including those related to remote-sensing image analysis. Successfully training a CNN to generalize well to unseen data, however, requires training on samples that represent the full distribution of variation of both the target classes and their surrounding contexts. With remote sensing data, acquiring a sufficiently representative training set is a challenge due to both the inherent multi-modal variability of satellite or aerial imagery and the general high cost of labeling data. To address this challenge, we have developed ISOSCELES, an Iterative Self-Organizing SCEne LEvel Sampling method for hierarchical sampling of large image sets. Using affinity propagation, ISOSCELES automates the selection of highly representative training images. Compared to random sampling or using available reference data, the distribution of the training is principally data driven, reducing the chance of oversampling uninformative areas or undersampling informative ones. In comparison to manual sample selection by an analyst, ISOSCELES exploits descriptive features, spectral and/or textural, and eliminates human bias in sample selection. Using a hierarchical sampling approach, ISOSCELES can obtain a training set that reflects both between-scene variability, such as in viewing angle and time of day, and within-scene variability at the level of individual training samples. We verify the method by demonstrating its superiority to stratified random sampling in the challenging task of adapting a pre-trained model to a new image and spatial domain for country-scale building extraction. Using a pair of hand-labeled training sets comprising 1,987 sample image chips, a total of 496,000,000 individually labeled pixels, we show, across three distinct model architectures, an increase in accuracy, as measured by F1-score, of 2.2–4.2%.

42 ENGINEERING↗

SafeDNN: Understanding and Verifying Neural Networks

The SafeDNN project at NASA Ames explores analysis techniques and tools to ensure that systems that use Deep Neural Networks (DNN) are safe, robust and interpretable. Research directions we are pursuing include: symbolic execution for DNN analysis, label-guided clustering to automatically identify input regions that are robust, parallel and compositional approaches to improve formal SMT-based verification, property inference and automated program repair for DNNs, adversarial training and detection, probabilistic reasoning for DNNs. In this talk I will highlight some of the research advances from SafeDNN, that were already published.

Corina Pasareanu↗

Automated Microbial Metabolism Laboratory

The Automated Microbial Metabolism Laboratory (AMML) 1971-1972 program involved the investigation of three separate life detection schemes. The first was a continued further development of the labeled release experiment. The possibility of chamber reuse without inbetween sterilization, to provide comparative biochemical information was tested. Findings show that individual substrates or concentrations of antimetabolites may be sequentially added to a single test chamber. The second detection system which was investigated for possible inclusion in the AMML package of assays, was nitrogen fixation as detected by acetylene reduction. Thirdly, a series of preliminary steps were taken to investigate the feasibility of detecting biopolymers in soil. A strategy for the safe return to Earth of a Mars sample prior to manned landings on Mars is outlined. The program assumes that the probability of indigenous life on Mars is unity and then broadly presents the procedures for acquisition and analysis of the Mars sample in a manner to satisfy the scientific community and the public that adequate safeguards are being taken.

Source record↗

Methodology for physics-informed generation of synthetic neutron time-of-flight measurement data

Accurate neutron cross section data are a vital input to the simulation of nuclear systems for a wide range of applications from energy production to national security. The evaluation of experimental data is a key step in producing accurate cross sections. There is a widely recognized lack of reproducibility in the evaluation process due to its artisanal nature and therefore there is a call for improvement within the nuclear data community. This can be realized by automating/standardizing viable parts of the process, namely, parameter estimation by fitting theoretical models to experimental data. This automation effort could greatly benefit from a synthetic data resource. This work leverages problem-specific physics, Monte Carlo sampling, and a general methodology for data synthesis to generate unlimited, labelled experimental cross-section data that is statistically indistinguishable to the observed data. Heuristic and, where applicable, rigorous statistical comparisons to observed data support this claim. The demonstration is based on/limited to transmission measurements at Rensselaer Polytechnic Institute (RPI) and energy-differential cross sections in the resolved resonance region (RRR). An open-source software is published alongside this article that executes the complete methodology to produce high-utility synthetic datasets. The goal of this work is to provide an approach and corresponding tool that will allow the evaluation community to begin exploring more data-driven, ML-based solutions to long-standing challenges in the field.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Optimizing automatic morphological classification of galaxies with machine learning and deep learning using Dark Energy Survey imaging

There are several supervised machine learning methods used for the application of automated morphological classification of galaxies; however, there has not yet been a clear comparison of these different methods using imaging data, or an investigation for maximizing their effectiveness. We carry out a comparison between several common machine learning methods for galaxy classification [Convolutional Neural Network (CNN), K-nearest neighbour, logistic regression, Support Vector Machine, Random Forest, and Neural Networks] by using Dark Energy Survey (DES) data combined with visual classifications from the Galaxy Zoo 1 project (GZ1). Our goal is to determine the optimal machine learning methods when using imaging data for galaxy classification. We show that CNN is the most successful method of these ten methods in our study. Using a sample of ~2800 galaxies with visual classification from GZ1, we reach an accuracy of ~0.99 for the morphological classification of ellipticals and spirals. The further investigation of the galaxies that have a different ML and visual classification but with high predicted probabilities in our CNN usually reveals the incorrect classification provided by GZ1. We further find the galaxies having a low probability of being either spirals or ellipticals are visually lenticulars (S0), demonstrating that supervised learning is able to rediscover that this class of galaxy is distinct from both ellipticals and spirals. We confirm that ~2.5 percent galaxies are misclassified by GZ1 in our study. After correcting these galaxies’ labels, we improve our CNN performance to an average accuracy of over 0.99 (accuracy of 0.994 is our best result).

79 ASTRONOMY AND ASTROPHYSICS↗

Automating the detection of hydrological barriers and fragmentation in wetlands using deep learning and InSAR

The loss of hydrological connectivity and fragmentation of natural wetlands is a widespread driver of wetland degradation. Understanding where and how natural connectivity is impaired is essential for managing, protecting and remediating these ecosystems. Wetland Interferometric Synthetic Aperture Radar (Wetland InSAR) can provide information on surface flow orientation in wetlands at a high spatial resolution, which can be used for barrier detection. However, the broad application of this approach is constrained by the labour-intensive manual delineation of barriers based on mapped water levels. This study presents the first deep learning-based methodology for the automated detection of hydrological barriers. We trained a deep convolutional network to segment edge features of hydrological barriers in 25 image pairs captured by ALOS PALSAR-1 L-Band InSAR between 2006 and 2011. The training dataset consists of manually labelled and delineated barriers showing abrupt changes in water surface elevation and wrapped interferograms with high coherence. We tested this method across three wetland sites: the Everglades and southern Louisiana wetlands (United States) and the Cienaga de Zapata (Cuba). Across these sites, the convolutional network detected hydrological barriers with up to 84% accuracy. The model performed particularly well for linear hydrological barriers such as roads, dikes, and channels. Notably, some barriers impede flow only seasonally, appearing during low water levels and disappearing when water levels rise. Our automated approach to detecting and assessing wetland hydrologic connectivity can be applied more broadly to support the effective management of fragmented wetland ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Fine-tuning TrailMap: The utility of transfer learning to improve the performance of deep learning in axon segmentation of light-sheet microscopy images

Light-sheet microscopy has made possible the 3D imaging of both fixed and live biological tissue, with samples as large as the entire mouse brain. However, segmentation and quantification of that data remains a time-consuming manual undertaking. Machine learning methods promise the possibility of automating this process. This study seeks to advance the performance of prior models through optimizing transfer learning. We fine-tuned the existing TrailMap model using expert-labeled data from noradrenergic axonal structures in the mouse brain. By changing the cross-entropy weights and using augmentation, we demonstrate a generally improved adjusted F1-score over using the originally trained TrailMap model within our test datasets.

97 MATHEMATICS AND COMPUTING↗

CENC PSN VOID FY20 Report

Maritime trade accounts for approximately 80 percent of international commerce. The high volume of vessels traversing domestic and international ports makes port areas prime targets for terrorism as well as illegal trafficking of drugs and arms (conventional or nuclear). Port security is therefore a worldwide concern affecting global economies, freedom of movement, and national security. However, extensive port monitoring is inherently complex and time consuming — making it truly viable only via an automated framework that can detect potential illicit activity and alert authorities in a timely manner. The development of image processing algorithms for this purpose requires access to large, labeled datasets that cover the breadth of targets of interest as well as the environments that they are observed within. Curated and labeled datasets of this nature are of enormous value to Sandia's Defense Nuclear Nonproliferation and National Security Program portfolios, as well as to Sandia's machine learning/automatic target recognition (ML/ATR) algorithm development and R&D communities. The goal of this project is to create a commercial satellite imagery dataset of labeled maritime vessels in port areas to support the development of ML/ATR algorithms for port security nonproliferation purposes. This dataset — Port Security Nonproliferation Vessel Overhead Imagery Dataset (PSN VOID) — has the potential to support a variety of other ancillary missions, such as maritime domain awareness, domestic and international security, drug interdiction, and weapons trafficking.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Automated Microbial Metabolism Laboratory

The effect of several environmental parameters on previously developed life detection systems is explored. Initial attempts were made to conduct all the experiments in a moist mode (high soil volume to water volume ratio). However, only labeled release and measurement of ATP were found to be feasible under conditions of low moisture. Therefore, these two life detection experiments were used for most of the environmental effects studies. Three soils, Mojave (California desert), Wyaconda (Maryland, sandy loam) and Victoria Valley (Antarctic desert) were generally used throughout. The environmental conditions studied included: incubation temperature 3 C to 80 C, ultraviolet irradiation of soils, variations in soil/liquid ratio, specific atmospheric gases, various antimetabolites, specific substrates, and variation in pH. An experiment designed to monitor nitrogen metabolism was also investigated.

Source record↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

Demonstration of sub-micron UCN position resolution using room-temperature CMOS sensor

High spatial resolution of ultracold neutron (UCN) measurement is of growing interest to UCN experiments such as UCN spectrometers, UCN polarimeters, quantum physics of UCNs, and quantum gravity. Here we utilize physics informed deep learning to enhance the experimental position resolution and to demonstrate sub-micron spatial resolutions for UCN position measurements obtained using a room-temperature CMOS sensor, extending our previous work that demonstrated a position uncertainty of 1.5 microns. We explore the use of the open-source software Allpix Squared to generate experiment-like synthetic hit images with ground-truth position labels. We use physics-informed deep learning by training a fully connected neural network (FCNN) to learn a mapping from input hit images to output hit position. The automated analysis for sub-micron position resolution in UCN detection combined with the fast data rates of current and next generation UCN sources will enable improved precision for future UCN research and applications.

10B nanometer thin film↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Refining fuzzy logic controllers with machine learning

In this paper, we describe the GARIC (Generalized Approximate Reasoning-Based Intelligent Control) architecture, which learns from its past performance and modifies the labels in the fuzzy rules to improve performance. It uses fuzzy reinforcement learning which is a hybrid method of fuzzy logic and reinforcement learning. This technology can simplify and automate the application of fuzzy logic control to a variety of systems. GARIC has been applied in simulation studies of the Space Shuttle rendezvous and docking experiments. It has the potential of being applied in other aerospace systems as well as in consumer products such as appliances, cameras, and cars.

Berenji, Hamid R.↗

Automated phase segmentation and quantification of high-resolution TEM image for alloy design

In the alloy design and development process, a wealth of atomically resolved structural high-resolution transmission electron microscopy (HRTEM) images are produced. Identifying the different nano-precipitate phases and tracking their evolution under various compositions and during manufacturing or post-processing requires hundreds of HRTEM images and thousands of precipitates. The nanoscopic phase information labeling and analysis purely relies on humans are prohibitively costly and time-consuming, sometimes not reliable because of the lack of authoritative knowledge. Here, in this work, we develop a novel unsupervised machine learning approach coupled with adaptive computer vision techniques with features in the Fourier space to automatically determine the number of phases and segment/quantify the phases with nanoscale resolution, allowing for quantitative correlation between nanostructure formation, processing and functional properties. To automate the phase extraction/quantification and ascertain its applicability, we have applied the developed framework to the HRTEM images from several alloy systems, processing conditions, image magnifications, and phase types and morphologies (precipitates, nano-twins, stacking faults, crystalline matrix, and amorphous structures) for verification. This study paves the road for compression, visualization, and translation of raw image structural data into physically relevant information in real-time with minimal human supervision. It shows the promise of enabling high-throughput materials characterization for the acceleration of alloy manufacturing and design.

36 MATERIALS SCIENCE↗

HT-SIP: a semi-automated stable isotope probing pipeline identifies cross-kingdom interactions in the hyphosphere of arbuscular mycorrhizal fungi

Abstract Background Linking the identity of wild microbes with their ecophysiological traits and environmental functions is a key ambition for microbial ecologists. Of many techniques that strive for this goal, Stable-isotope probing—SIP—remains among the most comprehensive for studying whole microbial communities in situ. In DNA-SIP, actively growing microorganisms that take up an isotopically heavy substrate build heavier DNA, which can be partitioned by density into multiple fractions and sequenced. However, SIP is relatively low throughput and requires significant hands-on labor. We designed and tested a semi-automated, high-throughput SIP (HT-SIP) pipeline to support well-replicated, temporally resolved amplicon and metagenomics experiments. We applied this pipeline to a soil microhabitat with significant ecological importance—the hyphosphere zone surrounding arbuscular mycorrhizal fungal (AMF) hyphae. AMF form symbiotic relationships with most plant species and play key roles in terrestrial nutrient and carbon cycling. Results Our HT-SIP pipeline for fractionation, cleanup, and nucleic acid quantification of density gradients requires one-sixth of the hands-on labor compared to manual SIP and allows 16 samples to be processed simultaneously. Automated density fractionation increased the reproducibility of SIP gradients compared to manual fractionation, and we show adding a non-ionic detergent to the gradient buffer improved SIP DNA recovery. We applied HT-SIP to 13 C-AMF hyphosphere DNA from a 13 CO 2 plant labeling study and created metagenome-assembled genomes (MAGs) using high-resolution SIP metagenomics (14 metagenomes per gradient). SIP confirmed the AMF Rhizophagus intraradices and associated MAGs were highly enriched (10–33 atom% 13 C), even though the soils’ overall enrichment was low (1.8 atom% 13 C). We assembled 212 13 C-hyphosphere MAGs; the hyphosphere taxa that assimilated the most AMF-derived 13 C were from the phyla Myxococcota, Fibrobacterota, Verrucomicrobiota, and the ammonia-oxidizing archaeon genus Nitrososphaera . Conclusions Our semi-automated HT-SIP approach decreases operator time and improves reproducibility by targeting the most labor-intensive steps of SIP—fraction collection and cleanup. We illustrate this approach in a unique and understudied soil microhabitat—generating MAGs of actively growing microbes living in the AMF hyphosphere (without plant roots). The MAGs’ phylogenetic composition and gene content suggest predation, decomposition, and ammonia oxidation may be key processes in hyphosphere nutrient cycling.

59 BASIC BIOLOGICAL SCIENCES↗

Introduction to Special Section: Machine Learning for Image-based Geologic Interpretation

Image-based geological interpretation has been a labor-intensive and time-consuming process because it requires well-trained geoscientists to identify geological structures, features, and textures from various types of images. These images include scanning electron microscopic images, optical microscopic images, optical photos, resistivity images, seismic volumes, remote-sensing images, etc. With fast-evolving machine learning (ML) technology and computing power in recent decades, computers can achieve nearhuman-level to super-human-level performance with scalable high efficiency in the computer vision field. These technological revolutions facilitated image-based geological interpretation in petroleum exploration and production. For example, a fault picking method applied to 3-D seismic volume data using deep learning can achieve superior performance in comparison to conventional auto-picking methods. In addition, under the new normal of low oil prices, the petroleum industry seeks cost-effective strategies such as automating traditionally labor-intensive processes. Nevertheless, the potential of applying ML to geological image interpretation is still facing a few key challenges including data scarcity, data distribution, poor data and/or label quality, data leakage, learning algorithms, model architecture, training methodologies, testing and evaluation metrics, hyper-parameters optimization, model drift, production deployment, and the like.

58 GEOSCIENCES↗