Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Image Labeler: Label Earth Science Images for Machine Learning

The application of machine learning for image-based classification of earth science phenomena, such as hurricanes, is relatively new. While extremely useful, the techniques used for image-based phenomena classification require storing and managing an abundant supply of labeled images in order to produce meaningful results. Existing methods for dataset management and labeling include maintaining categorized folders on a local machine, a process that can be cumbersome and not scalable. Image Labeler is a fast and scalable web-based tool that facilitates the rapid development of image-based earth science phenomena datasets, in order to aid deep learning application and automated image classification/detection. Image Labeler is built with modern web technologies to maximize the scalability and availability of the platform. It has a user-friendly interface that allows tagging multiple images relatively quickly. Essentially, Image Labeler improves upon existing techniques by providing researchers with a shareable source of tagged earth science images for all their machine learning needs. Here, we demonstrate Image Labeler’s current image extraction and labeling capabilities including supported data sources, spatiotemporal subsetting capabilities, individual project management and team collaboration for large scale projects.

Acharya, Ashish↗

Auto-Curation of Seismic Event Data for Signal Denoising

Denoising contaminated seismic signals for later processing is a fundamental problem in seismic signals analysis. Neural network approaches have shown success denoising local signals when trained on short-time Fourier transform spectrograms. One challenge of this approach is the onerous process of hand-labeling event signals for training. By leveraging the SCALODEEP seismic event detector, we develop an automated set of techniques for labeling event data. Despite region specific challenges, training the neural network denoiser on machine curated events shows comparable performance to the neural network trained on hand curated events. We showcase our technique with two experiments, one using Utah regional data and one using regional data from the Korean peninsula.

58 GEOSCIENCES↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗

Automated annotation of scientific texts for ML-based keyphrase extraction and validation

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)–based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Implementation of Machine Learning Methods for Crater-Based Navigation

Terrain Relative Navigation methods require surface feature detectors to gain information from images used to improve on-board state estimates. This paper presents the development of a crater detection method based on Machine Learning that can extract data from optical images with different crater shapes and sizes, under varying lighting conditions. This work includes an automated capability for generating labeled training data and iterative testing of the neural network-based crater detector. Preliminary results are included to quantify the detector’s accuracy compared to a known crater catalog, given a set of real lunar images from the Lunar Reconnaissance Orbiter.

Sofia G Catalan↗

Detecting technological maturity from bibliometric patterns

We report the capability to identify emergent technologies based upon easily accessed open-source indicators, such as publications, is important for decision-makers in industry and government. The scientific contribution of this work is the proposition of a machine learning approach to the detection of the maturity of emerging technologies based on publication counts. Time-series of publication counts have universal features that distinguish emerging and growing technologies. We train an artificial neural network classifier, a supervised machine learning algorithm, upon these features to predict the maturity (emergent vs. growth) of an arbitrary technology. With a training set comprised of 22 technologies we obtain a classification accuracy ranging from 58.3% to 100% with an average accuracy of 84.6% for six test technologies. To enhance classifier performance, we augmented the training corpus with synthetic time-series technology life cycle curves, formed by calculating weighted averages of curves in the original training set. Training the classifier on the synthetic data set resulted in improved accuracy, ranging from 83.3% to 100% with an average accuracy of 90.4% for the test technologies. The performance of our classifier exceeds that of competing machine learning approaches in the literature, which report an average classification accuracy of only 85.7% at maximum. Moreover, in contrast to current methods our approach does not require subject matter expertise to generate training labels, and it can be automated and scaled.

97 MATHEMATICS AND COMPUTING↗

SSG-4 - An automated spring small grains proportion estimator

In connection with an implementation of the classification procedures employed in the Large Area Crop Inventory Experiment (LACIE), a human analyst had to provide labeled samples. The present investigation is concerned with an automated proportion estimation procedure which has been derived from the early field-labeling procedures used in LACIE. This procedure was developed for the U.S./Canada Spring Small Grains Pilot Experiment. It is demonstrated that the considered spatial/color-based proportion estimation procedure provides the agricultural remote-sensing community with the basic tools to develop unbiased and highly efficient procedures for obtaining crop area estimates at the end of the season.

Dennis, T. B.↗

Automated detection of photovoltaic cleaning events: A performance comparison of techniques as applied to a broad set of labeled photovoltaic data sets

Extracting accurate soiling loss information from photovoltaic (PV) production data first requires segmenting the time series data per natural or manually occurring cleaning events. Maintenance logs are often incomplete, rain data are often unavailable, and the debate on rain thresholds for cleaning and dew or wind cleanings is still ongoing. The present work aims to overtake these issues by improving automated methods to detect these cleaning events and therefore improve extraction of soiling loss information. Time series power production data from 22 PV inverters were labeled for natural or manually occurring cleaning events. The data sets were carefully selected to include varying degrees of soiling, cleaning events, and noise. Several algorithms, including filtering logic and change point detection, were examined for efficacy at detecting the labeled cleanings. All the methods introduced except for changepoint detection showed significant improvement at detecting the labeled cleaning events per the mean F 1 score. Furthermore, the highest performing cleaning detection algorithm achieved an absolute increase in the mean F 1 score of 43% over the default version of the RdTools stochastic rate and recovery (SRR) algorithm. The highest performing algorithm included irradiance filtering and a cleaning detection threshold, adjusted based on the 40-day centered rolling median of the absolute day-to-day deviations in the daily performance index (PI). Furthermore, these improvements are promising as cleaning detection is an essential step in the automated analysis of PV soiling.

14 SOLAR ENERGY↗

Few-shot Learning for Post-disaster Structure Damage Assessment

Automating post-disaster damage assessment with remote sensing data is critical for faster surveys of structures impacted by natural disasters. One significant obstacle to training state-of-the-art deep neural networks to support this automation is that large quantities of labelled data are often required. However, obtaining those labels is particularly unrealistic to support post-disaster damage assessment in a timely manner. Few-shot learning methods could help to mitigate this by reducing the amount of labelled data required to successfully train a model while achieving satisfactory results. To this end, we explore a feature reweighting method to the YOLOv3 object detection architecture to achieve few-shot learning of damage assessment models on the xBD dataset. Our results show that the feature reweighting approach yield improved mAP over the baseline with significantly fewer labelled samples. In addition, we use t-SNE to analyze the class-specific reweighting vectors generated by the reweighting module in order to evaluate their inter-class and intra-class similarity. We find that the vectors form clusters based on class, and that these clusters overlap with visually similar classes. Those results show the potential to employ this few-shot learning strategy for rapid damage assessment with post-event remote sensing images.

Bowman, Jordan↗

Synopsis of a computer program designed to interface a personal computer with the fast data acquisition system of a time-of-flight mass spectrometer

Briefly described are the essential features of a computer program designed to interface a personal computer with the fast, digital data acquisition system of a time-of-flight mass spectrometer. The instrumentation was developed to provide a time-resolved analysis of individual vapor pulses produced by the incidence of a pulsed laser beam on an ablative material. The high repetition rate spectrometer coupled to a fast transient recorder captures complete mass spectra every 20 to 35 microsecs, thereby providing the time resolution needed for the study of this sort of transient event. The program enables the computer to record the large amount of data generated by the system in short time intervals, and it provides the operator the immediate option of presenting the spectral data in several different formats. Furthermore, the system does this with a high degree of automation, including the tasks of mass labeling the spectra and logging pertinent instrumental parameters.

Bechtel, R. D.↗

Supervised Machine Learning Approach for Classifying Earth Science Publications

The data collections archived and distributed by the GES DISC NASA data center are widely utilized for various Earth Science studies. As these collections are created, many research works are published regarding these collections' algorithms, their validation, and their applications. As NASA data centers collect these publications for public use, it is helpful to categorize them based on how they relate to their associated datasets. Specifically, whether the publication linked to the GES DISC dataset is using it for applicational research, describing the algorithm used for the dataset creation, validating the dataset, or providing a general overview of the data collection. Currently, this process requires simple manual labeling, and as such, it may be possible to solve via automation. To approach this problem, machine learning classifiers were developed to predict a publication's category. Manually labeled publications were used as the training data for the supervised machine learning algorithms, specifically Random Forest and Multinomial Naïve Bayes. After balancing the dataset and implementing the Multinomial Naïve Bayes algorithm, the classification accuracy achieved was substantially higher than the baseline accuracy, thus significantly improving the efficiency of publication labeling.

Rohan Dayal↗

Deep learning classification of lipid droplets in quantitative phase images

We report the application of supervised machine learning to the automated classification of lipid droplets in label-free, quantitative-phase images. By comparing various machine learning methods commonly used in biomedical imaging and remote sensing, we found convolutional neural networks to outperform others, both quantitatively and qualitatively. We describe our imaging approach, all implemented machine learning methods, and their performance with respect to computational efficiency, required training resources, and relative method performance measured across multiple metrics. Overall, our results indicate that quantitative-phase imaging coupled to machine learning enables accurate lipid droplet classification in single living cells. As such, the present paradigm presents an excellent alternative of the more common fluorescent and Raman imaging modalities by enabling label-free, ultra-low phototoxicity, and deeper insight into the thermodynamics of metabolism of single cells.

59 BASIC BIOLOGICAL SCIENCES↗

Category identification of changed land-use polygons in an integrated image processing/geographic information system

A framework is proposed for analyzing ancillary data and developing procedures for incorporating ancillary data to aid interactive identification of land-use categories in land-use updates. The procedures were developed for use within an integrated image processsing/geographic information systems (GIS) that permits simultaneous display of digital image data with the vector land-use data to be updated. With such systems and procedures, automated techniques are integrated with visual-based manual interpretation to exploit the capabilities of both. The procedural framework developed was applied as part of a case study to update a portion of the land-use layer in a regional scale GIS. About 75 percent of the area in the study site that experienced a change in land use was correctly labeled into 19 categories using the combination of automated and visual interpretation procedures developed in the study.

Westmoreland, Sally↗

The value of human data annotation for machine learning based anomaly detection in environmental systems

Anomaly detection is the process of identifying unexpected data samples in datasets. Automated anomaly detection is either performed using supervised machine learning models, which require a labelled dataset for their calibration, or unsupervised models, which do not require labels. While academic research has produced a vast array of tools and machine learning models for automated anomaly detection, the research community focused on environmental systems still lacks a comparative analysis that is simultaneously comprehensive, objective, and systematic. This knowledge gap is addressed for the first time in this study, where 15 different supervised and unsupervised anomaly detection models are evaluated on 5 different environmental datasets from engineered and natural aquatic systems. To this end, anomaly detection performance, labelling efforts, as well as the impact of model and algorithm tuning are taken into account. As a result, our analysis reveals the relative strengths and weaknesses of the different approaches in an objective manner without bias for any particular paradigm in machine learning. Most importantly, our results show that expert-based data annotation is extremely valuable for anomaly detection based on machine learning.

54 ENVIRONMENTAL SCIENCES↗

Automated microbial metabolism laboratory

The design and rationale of an advanced labeled release experiment based on single addition of soil and multiple sequential additions of media into each of four test chambers are outlined. The feasibility for multiple addition tests was established and various details of the methodology were studied. The four chamber battery of tests include: (1) determination of the effect of various atmospheric gases and selection of that gas which produces an optimum response; (2) determination of the effect of incubation temperature and selection of the optimum temperature for performing Martian biochemical tests; (3) sterile soil is dosed with a battery of C-14 labeled substrates and subjected to experimental temperature range; and (4) determination of the possible inhibitory effects of water on Martian organisms is performed initially by dosing with 0.01 ml and 0.5 ml of medium, respectively. A series of specifically labeled substrates are then added to obtain patterns in metabolic 14CO2 (C-14)O2 evolution.

Source record↗

recon3d

SAND2025-00533O recon3d is a software tool that provides automated 3D reconstruction and meshing capabilities. It processes labeled 3D image data from various sources, starting from image stacks, and calculates 3D feature distributions like size, shape, and location. The software also has tools for downscaling rectilinear grid data and creating tetrahedral meshes directly from image data. recon3d can be used by novice users via the command line with a properly formatted configuration file. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Emery, John↗

Improving designer productivity

Designer and design team productivity improves with skill, experience, and the tools available. The design process involves numerous trials and errors, analyses, refinements, and addition of details. Computerized tools have greatly speeded the analysis, and now new theories and methods, emerging under the label Artificial Intelligence (AI), are being used to automate skill and experience. These tools improve designer productivity by capturing experience, emulating recognized skillful designers, and making the essence of complex programs easier to grasp. This paper outlines the aircraft design process in today's technology and business climate, presenting some of the challenges ahead and some of the promising AI methods for meeting these challenges.

Hill, Gary C.↗

Improving designer productivity

Designer and design team productivity improves with skill, experience, and the tools available. The design process involves numerous trials and errors, analyses, refinements, and addition of details. Computerized tools have greatly speeded the analysis, and now new theories and methods, emerging under the label Artificial Intelligence (AI), are being used to automate skill and experience. These tools improve designer productivity by capturing experience, emulating recognized skillful designers, and making the essence of complex programs easier to grasp. This paper outlines the aircraft design process in today's technology and business climate, presenting some of the challenges ahead and some of the promising AI methods for meeting those challenges.

Hill, Gary C.↗