Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Seascape Interface Control Document

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source software, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Structure of the knowledge base for an expert labeling system

One of the principal objectives of the NASA AgRISTARS program is the inventory of global crop resources using remotely sensed data gathered by Land Satellites (LANDSAT). A central problem in any such crop inventory procedure is the interpretation of LANDSAT images and identification of parts of each image which are covered by a particular crop of interest. This task of labeling is largely a manual one done by trained human analysts and consequently presents obstacles to the development of totally automated crop inventory systems. However, development in knowledge engineering as well as widespread availability of inexpensive hardware and software for artificial intelligence work offers possibilities for developing expert systems for labeling of crops. Such a knowledge based approach to labeling is presented.

Rajaram, N. S.↗

An Automated Classification Technique for Detecting Defects in Battery Cells

Battery cell defect classification is primarily done manually by a human conducting a visual inspection to determine if the battery cell is acceptable for a particular use or device. Human visual inspection is a time consuming task when compared to an inspection process conducted by a machine vision system. Human inspection is also subject to human error and fatigue over time. We present a machine vision technique that can be used to automatically identify defective sections of battery cells via a morphological feature-based classifier using an adaptive two-dimensional fast Fourier transformation technique. The initial area of interest is automatically classified as either an anode or cathode cell view as well as classified as an acceptable or a defective battery cell. Each battery cell is labeled and cataloged for comparison and analysis. The result is the implementation of an automated machine vision technique that provides a highly repeatable and reproducible method of identifying and quantifying defects in battery cells.

McDowell, Mark↗

Deep active learning for classifying cancer pathology reports

Abstract Background Automated text classification has many important applications in the clinical setting; however, obtaining labelled data for training machine learning and deep learning models is often difficult and expensive. Active learning techniques may mitigate this challenge by reducing the amount of labelled data required to effectively train a model. In this study, we analyze the effectiveness of 11 active learning algorithms on classifying subsite and histology from cancer pathology reports using a Convolutional Neural Network as the text classification model. Results We compare the performance of each active learning strategy using two differently sized datasets and two different classification tasks. Our results show that on all tasks and dataset sizes, all active learning strategies except diversity-sampling strategies outperformed random sampling, i.e., no active learning. On our large dataset (15K initial labelled samples, adding 15K additional labelled samples each iteration of active learning), there was no clear winner between the different active learning strategies. On our small dataset (1K initial labelled samples, adding 1K additional labelled samples each iteration of active learning), marginal and ratio uncertainty sampling performed better than all other active learning techniques. We found that compared to random sampling, active learning strongly helps performance on rare classes by focusing on underrepresented classes. Conclusions Active learning can save annotation cost by helping human annotators efficiently and intelligently select which samples to label. Our results show that a dataset constructed using effective active learning techniques requires less than half the amount of labelled data to achieve the same performance as a dataset constructed using random sampling.

59 BASIC BIOLOGICAL SCIENCES↗

Empirical Analysis and Automated Classification of Security Bug Reports

With the ever expanding amount of sensitive data being placed into computer systems, the need for effective cybersecurity is of utmost importance. However, there is a shortage of detailed empirical studies of security vulnerabilities from which cybersecurity metrics and best practices could be determined. This thesis has two main research goals: (1) to explore the distribution and characteristics of security vulnerabilities based on the information provided in bug tracking systems and (2) to develop data analytics approaches for automatic classification of bug reports as security or non-security related. This work is based on using three NASA datasets as case studies. The empirical analysis showed that the majority of software vulnerabilities belong only to a small number of types. Addressing these types of vulnerabilities will consequently lead to cost efficient improvement of software security. Since this analysis requires labeling of each bug report in the bug tracking system, we explored using machine learning to automate the classification of each bug report as a security or non-security related (two-class classification), as well as each security related bug report as specific security type (multiclass classification). In addition to using supervised machine learning algorithms, a novel unsupervised machine learning approach is proposed. An ac- curacy of 92%, recall of 96%, precision of 92%, probability of false alarm of 4%, F-Score of 81% and G-Score of 90% were the best results achieved during two-class classification. Furthermore, an accuracy of 80%, recall of 80%, precision of 94%, and F-score of 85% were the best results achieved during multiclass classification.

Cybersecurity↗

Automated pixel screening and selection technique

One approach for estimating crop acreage on the basis of Landsat data involves the sampling of portions of Landsat imagery which represent 5 x 6-n mi areas on the ground, taking into account acreage estimates for each of these areas. Difficulties concerning the estimating procedures are related to labeling. These difficulties arise in particular in connection with mixed pixels. The present investigation is concerned with the employment of a suitable automated technique for overcoming these difficulties. A description of the selected technique is presented, and the test conducted for the evaluation of the technique is discussed. It is found that the employed automated pixel screening and selection procedure provides representative samples of the scene. The amount of time spent on pure pixel selection is significantly reduced by using the automated method.

Dailey, C. L.↗

Proof Rules for Automated Compositional Verification through Learning

Compositional proof systems not only enable the stepwise development of concurrent processes but also provide a basis to alleviate the state explosion problem associated with model checking. An assume-guarantee style of specification and reasoning has long been advocated to achieve compositionality. However, this style of reasoning is often non-trivial, typically requiring human input to determine appropriate assumptions. In this paper, we present novel assume- guarantee rules in the setting of finite labelled transition systems with blocking communication. We show how these rules can be applied in an iterative and fully automated fashion within a framework based on learning.

Barringer, Howard↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, 51%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear hits. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Pusey, Marc L.↗

ClimateNet: an expert-labeled open dataset and deep learning architecture for enabling high-precision analyses of extreme weather

Abstract. Identifying, detecting, and localizing extreme weather events is a crucial first step in understanding how they may vary under different climate change scenarios. Pattern recognition tasks such as classification, object detection, and segmentation (i.e., pixel-level classification) have remained challenging problems in the weather and climate sciences. While there exist many empirical heuristics for detecting extreme events, the disparities between the output of these different methods even for a single event are large and often difficult to reconcile. Given the success of deep learning (DL) in tackling similar problems in computer vision, we advocate a DL-based approach. DL, however, works best in the context of supervised learning – when labeled datasets are readily available. Reliable labeled training data for extreme weather and climate events is scarce. We create “ClimateNet” – an open, community-sourced human-expert-labeled curated dataset that captures tropical cyclones (TCs) and atmospheric rivers (ARs) in high-resolution climate model output from a simulation of a recent historical period. We use the curated ClimateNet dataset to train a state-of-the-art DL model for pixel-level identification – i.e., segmentation – of TCs and ARs. We then apply the trained DL model to historical and climate change scenarios simulated by the Community Atmospheric Model (CAM5.1) and show that the DL model accurately segments the data into TCs, ARs, or “the background” at a pixel level. Further, we show how the segmentation results can be used to conduct spatially and temporally precise analytics by quantifying distributions of extreme precipitation conditioned on event types (TC or AR) at regional scales. The key contribution of this work is that it paves the way for DL-based automated, high-fidelity, and highly precise analytics of climate data using a curated expert-labeled dataset – ClimateNet. ClimateNet and the DL-based segmentation method provide several unique capabilities: (i) they can be used to calculate a variety of TC and AR statistics at a fine-grained level; (ii) they can be applied to different climate scenarios and different datasets without tuning as they do not rely on threshold conditions; and (iii) the proposed DL method is suitable for rapidly analyzing large amounts of climate model output. While our study has been conducted for two important extreme weather patterns (TCs and ARs) in simulation datasets, we believe that this methodology can be applied to a much broader class of patterns and applied to observational and reanalysis data products via transfer learning.

54 ENVIRONMENTAL SCIENCES↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of a macromolecules purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals will show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "bits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Minamitani, Elizabeth Forsythe↗

Preliminary Radiation Testing of a State-of-the-Art Commercial 14nm CMOS Processor - System-on-a-Chip

Hardness assurance test results of Intel state-of-the-art 14nm Broadwell U-series processor System-on-a-Chip (SoC) for total dose are presented, along with first-look exploratory results from trials at a medical proton facility. Test method builds upon previous efforts by utilizing commercial laptop motherboards and software stress applications as opposed to more traditional automated test equipment (ATE).

14nm↗

Preliminary Radiation Testing of a State-of-the-Art Commercial 14nm CMOS Processor/System-on-a-Chip

Hardness assurance test results of Intel state-of-the-art 14nm “Broadwell” U-series processor / System-on-a-Chip (SoC) for total ionizing dose (TID) are presented, along with exploratory results from trials at a medical proton facility. Test method builds upon previous efforts [1] by utilizing commercial laptop motherboards and software stress applications as opposed to more traditional automated test equipment (ATE).

microprocessor↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically can not reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "hits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Pusey, Marc L.↗

Reinforcement learning based automated history matching for improved hydrocarbon production forecast

History matching aims to find a numerical reservoir model that can be used to predict the reservoir performance. An engineer and model calibration (data inversion) method are required to adjust various parameters/properties of the numerical model in order to match the reservoir production history. In this study, we develop deep neural networks within the reinforcement learning framework to achieve automated history matching that will reduce engineers’ efforts, human bias, automatically and intelligently explore the parameter space, and remove the need of large set of labeled training data. To that end, a fast-marching-based reservoir simulator is encapsulated as an environment for the proposed reinforcement learning. The deep neural-network-based learning agent interacts with the reservoir simulator within reinforcement learning framework to achieve the automated history matching. Reinforcement learning techniques, such as discrete Deep Q Network and continuous Deep Deterministic Policy Gradients, are used toth, used to train the learning agents. The continuous actions enable the Deep Deterministic Policy Gradients to explore more states at each iteration in a a learning episode; consequently, a better history matching is achieved using this algorithm as compared to Deep Q Network. For simplified dual-target composite reservoir models, the best history-matching performances of the discrete and continuous learning methods in terms of normalized root mean square errors are 0.0447 and 0.0038, respectively. Furthermore, our study shows that continuous action space achieved by the deep deterministic policy gradient drastically outperforms deep Q network.

42 ENGINEERING↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

CyBERT: Cybersecurity Claim Classification by Fine-Tuning the BERT Language Model

We introduce CyBERT, a cybersecurity feature claims classifier based on bidirectional encoder representations from transformers and a key component in our semi-automated cybersecurity vetting for industrial control systems (ICS). To train CyBERT, we created a corpus of labeled sequences from ICS device documentation collected across a wide range of vendors and devices. This corpus provides the foundation for fine-tuning BERT’s language model, including a prediction-guided relabeling process. We propose an approach to obtain optimal hyperparameters, including the learning rate, the number of dense layers, and their configuration, to increase the accuracy of our classifier. Fine-tuning all hyperparameters of the resulting model led to an increase in classification accuracy from 76% obtained with BertForSequenceClassification’s original architecture to 94.4% obtained with CyBERT. Furthermore, we evaluated CyBERT for the impact of randomness in the initialization, training, and data-sampling phases. CyBERT demonstrated a standard deviation of ±0.6% during validation across 100 random seed values. Finally, we also compared the performance of CyBERT to other well-established language models including GPT2, ULMFiT, and ELMo, as well as neural network models such as CNN, LSTM, and BiLSTM. The results showed that CyBERT outperforms these models on the validation accuracy and the F1 score, validating CyBERT’s robustness and accuracy as a cybersecurity feature claims classifier.

97 MATHEMATICS AND COMPUTING↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1 %, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "hits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. Preliminary experiments show that the presence of the fluorescent probe does not affect the nucleation process or the quality of the X-ray data obtained.

Pusey, Marc L.↗

Automated Identification of Molecular Crystals’ Packing Motifs

Packing motifs—patterns in how molecules orient relative to one another in a crystal structure—are an important concept in many subdisciplines of materials science because of correlations observed between specific packing motifs and properties of interest. That said, packing motif data sets have remained small and noisy due to intensive manual labeling processes and insufficient labeling schemes. The most prominent labeling algorithms calculate relative interplanar angles of nearest neighbor molecules to determine the packing motif of a molecular crystal, but this simple approach can fail when neighbors are naively sampled isotropically around the crystal structure. To remedy this issue, here we propose an optimization algorithm, which rotates the molecular crystal structure to find representative molecules that inform the packing motif. We package this algorithm into an automated framework—Autopack—which both optimally rotates the crystal structure and labels the packing motif based on the appropriate neighboring molecules. In this work, we detail the Autopack framework and its performance, which shows improvements compared to previous state-of-the-art labeling methods, providing the first quantitative point of comparison for packing motif labeling algorithms. Furthermore, using Autopack (available at https://ipo.llnl.gov/technologies/software/autopack), we perform the first large-scale study of potential relationships between chemicals’ compositions and packing motifs, which shows that these relationships are more complex than previously hypothesized from studies that used only tens of polycyclic aromatic hydrocarbon molecules. Autopack’s capabilities help pose next steps for crystal engineering research focusing not only on a molecule’s adoption of a specific packing motif but also on new structure–property relationships.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗