Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

The CanBikeCO Mini Pilot: Preliminary Results and Lessons Learned

In fall 2020, the Colorado Energy Office, as part of the State of Colorado’s “Can Do Colorado” initiative, initiated a project aimed at encouraging energy-efficient transportation during the COVID-19 pandemic. The initial mini-pilot provided e-bikes to 13 low-income households under an individual ownership model. This report assesses the impact of providing this additional mobility option on the travel behavior of participants. It also outlines the lessons learned from deploying a continuous monitoring platform to track the travel behavior. These lessons will influence the evaluation component for the full pilot, which will cover multiple geographic regions, starting in summer 2021, and run for 2 years. The continuous data collection was enabled by a customized version of the open-source e-mission platform, called CanBikeCO, configured with a behavioral gamification feature. The Colorado Energy Office used this system to collect a unique data set consisting of 3 months of partially automated travel diaries, combining sensed and surveyed data and linked with demographic information, from 12 participants. The data collection process worked well overall: users generally liked the app, appreciated the game, and did not complain about battery life. The long tracking period introduced behavioral challenges in user engagement, which we plan to address using repeated patterns and automated status checks for the full pilot. The analysis results, based on the subset of trips with user-reported labels (68%), indicate that the e-bike was the dominant commute mode share (31%), in sharp contrast to the census bicycle commute mode share (<1%). E-bike trips primarily replaced single-occupancy vehicle (SOV) trips (28%), followed closely by walking (24%) and regular bike (20%). The nonmotorized mode replacement corresponds to lower travel time and increased productivity enabled by the program. The emissions impact analysis of the program, computed using trip-level energy intensity factors, indicates savings of 1,367 lbs. of CO 2 . Although the results are strongly positive, the narrow demographic profile of study participants, their limited mobility alternatives, and nonuniform labeling indicate caution in broader interpretation. These preliminary results do suggest that such programs, supported by real-time education and support from program managers, can simultaneously meet equity and sustainability goals. The planned full pilot, addressing the data collection challenges and broadening the geographic scope, will provide additional insights into the generality of this approach.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Challenges in Automated Detection of COVID-19 Misinformation

The COVID-19 pandemic has made the dangers of the spread of misinformation obvious but despite much global effort to curbing its spread, fake information about the pandemic keeps proliferating. In this paper, we address the development of automated methods for verification of claims about COVID-19 and discuss the challenges associated with this task. We focus on labeled data collection, limitations of existing models, and difficulties of applying misinformation detection models in practical applications. Our initial analysis indicates label imbalance may be a particular challenge for developing claim verification models and we discuss options for alleviating this issue.

Herrmannova, Dasha↗

The Clean Energy Mortgage

Explore the source record for details and available documents.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The CanBikeCO Mini Pilot: Procedure and Preliminary Results

In fall 2020, the Colorado Energy Office, as part of the State of Colorado's "Can Do Colorado" initiative, initiated a project aimed at encouraging energy-efficient transportation during the COVID-19 pandemic. The initial mini-pilot provided e-bikes to 13 low-income households under an individual ownership model. This report assesses the impact of providing this additional mobility option on the travel behavior of participants. It also outlines the lessons learned from deploying a continuous monitoring platform to track the travel behavior. These lessons will influence the evaluation component for the full pilot, which will cover multiple geographic regions, start in summer 2021, and run for 2 years. The continuous data collection was enabled by a customized version of the open-source e-mission platform, called CanBikeCO, configured with a behavioral gamification feature. The Colorado Energy Office used this system to collect a unique data set consisting of 3 months of partially automated travel diaries, combining sensed and surveyed data and linked with demographic information, from 12 participants. The data collection process worked well overall: users generally liked the app, appreciated the game, and did not complain about battery life. The long tracking period introduced behavioral challenges in user engagement, which we plan to address using repeated patterns and automated status checks for the full pilot. The analysis results, based on the subset of trips with user-reported labels (68%), indicate that the e-bike was the dominant commute mode share (31%), in sharp contrast to the census bicycle commute mode share (<1%). E-bike trips primarily replaced single-occupancy vehicle (SOV) trips (28%), followed closely by walking (24%) and regular bike (20%). The non-motorized mode replacement corresponds to lower travel time and increased productivity enabled by the program. The emissions impact analysis of the program, computed using trip-level energy intensity factors, indicates savings of 1,367 lbs. of CO2. Although the results are strongly positive, the narrow demographic profile of study participants, their limited mobility alternatives, and nonuniform labeling indicate caution in broader interpretation. These preliminary results do suggest that such programs, supported by real-time education and support from program managers, can simultaneously meet equity and sustainability goals. The planned full pilot, addressing the data collection challenges and broadening the geographic scope, will provide additional insights into the generality of this approach.

ADVANCED PROPULSION SYSTEMS↗

Leveraging generative adversarial networks to create realistic scanning transmission electron microscopy images

Abstract The rise of automation and machine learning (ML) in electron microscopy has the potential to revolutionize materials research through autonomous data collection and processing. A significant challenge lies in developing ML models that rapidly generalize to large data sets under varying experimental conditions. We address this by employing a cycle generative adversarial network (CycleGAN) with a reciprocal space discriminator, which augments simulated data with realistic spatial frequency information. This allows the CycleGAN to generate images nearly indistinguishable from real data and provide labels for ML applications. We showcase our approach by training a fully convolutional network (FCN) to identify single atom defects in a 4.5 million atom data set, collected using automated acquisition in an aberration-corrected scanning transmission electron microscope (STEM). Our method produces adaptable FCNs that can adjust to dynamically changing experimental variables with minimal intervention, marking a crucial step towards fully autonomous harnessing of microscopy big data.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Automatic Segmentation of Building Envelope Point Cloud Data Using Machine Learning

About 50% of buildings in the US were constructed before energy codes were introduced. Modular overclad panel retrofits, in which a new envelope is constructed over the existing building, are a promising solution given that it minimizes occupant disruption and shortens construction time at the jobsite. Current state-of-the-art retrofit panel layout and dimensioning consists of three steps: 1) 3D point cloud data generation of the building envelope using commonly available surveying equipment, 2) manual segmentation of 3D point cloud data by a trained professional to identify and dimension window openings, door openings, and other architectural features, and 3) modular panel layout optimization and dimensioning by an architect or engineer. Among these steps, the second one remains the most difficult and costly because it is very labor-intensive. We propose a methodology to automatically label 3D point cloud data to reduce the time and expense spent in manual segmentation. Machine learning methods were employed to classify the point cloud data into distinct groups, each of which corresponds to different features of the building envelope. After classification, a segmentation algorithm was developed to perform boundary detection and separate the components of the façade. Finally, the algorithm returns the relative positions and dimensions of the features in the building envelope. The measurements obtained with the proposed automated method were compared against the actual dimensions to determine the overall algorithm accuracy. The proposed algorithm can then be used to reduce manual efforts for 3D point cloud labeling before modular panel layout optimization is performed.

Maldonado Puente, Bryan↗

REACTER 2.0: Quantum-Informed Reaction Constraints and Automated Interaction Typing

REACTER is a heuristic method for modeling chemical reactions in classical molecular dynamics simulations, implemented in LAMMPS as fix bond/react. The authors recently extended LAMMPS to support alphanumeric labels for atom types, bond types etc., which enables the pre- and post-reaction templates required by the REACTER protocol to be portable between different simulations and greatly simplifies the task of creating simulation-ready reaction templates. To further increase the generality of reaction templates, support for wildcard characters within atom types has been added, along with the automatic assignment of interaction types for new bonds, angles, etc. based on the involved atom types. In some cases, this feature can express a class of reactions with one pair of reaction templates, where previously dozens may have been required. Advanced reaction constraints have also been added, including an Arrhenius constraint to enforce an effective activation energy, a root-mean-square-deviation option for complex geometrical constraints, as well as a custom constraint that leverages LAMMPS’ powerful built-in variable framework. Other new features include variable support for various inputs (e.g., to allow reaction rates or cutoffs to be dependent on overall conversion), on-the-fly update of molecule IDs, and the ability to create new atoms positioned with respect to the reaction site. The new features are applied to modeling polymeric, thermosetting and composite materials, and advanced applications of the new reaction constraints are demonstrated. For example, REACTER is shown to accurately reproduce mechanically-induced bond breaking, as characterized by third-order DFT-based tight-binding (DFTB3) simulations, via a constraint on the total potential energy of the involved atoms.

polymer simulations↗

AgRISTARS: Foreign commodity production forecasting. Corn/soybean decision logic development and testing

The development and testing of an analysis procedure which was developed to improve the consistency and objectively of crop identification using Landsat data is described. The procedure was developed to identify corn and soybean crops in the U.S. corn belt region. The procedure consists of a series of decision points arranged in a tree-like structure, the branches of which lead an analyst to crop labels. The specific decision logic is designed to maximize the objectively of the identification process and to promote the possibility of future automation. Significant results are summarized.

Dailey, C. L.↗

MultiTaskDeltaNet: change detection-based image segmentation for operando ETEM with application to carbon gasification kinetics

Transforming in situ transmission electron microscopy (TEM) imaging into a tool for spatially-resolved operando characterization of solid-state reactions requires automated, high-precision semantic segmentation of dynamically evolving features. However, traditional deep learning methods for semantic segmentation often face limitations due to the scarcity of labeled data, visually ambiguous features of interest, and scenarios involving small objects. To tackle these challenges, we introduce MultiTaskDeltaNet (MTDN), a novel deep learning architecture that creatively reconceptualizes the segmentation task as a change detection problem. By implementing a unique Siamese network with a U-Net backbone and using paired images to capture feature changes, MTDN effectively leverages minimal data to produce high-quality segmentations. Furthermore, MTDN utilizes a multi-task learning strategy to exploit correlations between physical features of interest. In an evaluation using data from in situ environmental TEM (ETEM) videos of filamentous carbon gasification, MTDN demonstrated a significant advantage over conventional segmentation models, particularly in accurately delineating fine structural features. Notably, MTDN achieved a 10.22% performance improvement over conventional segmentation models in predicting small and visually ambiguous physical features. This work bridges key gaps between deep learning and practical TEM image analysis, advancing automated characterization of nanomaterials in complex experimental settings.

08 HYDROGEN↗

Fluorescent Applications to Crystallization

By covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and tests with model proteins have shown that labeling u to 5 percent of the protein molecules does not affect the X-ray data quality obtained . The presence of the trace fluorescent label gives a number of advantages. Since the label is covalently attached to the protein molecules, it "tracks" the protein s response to the crystallization conditions. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a darker background. Non-protein structures, such as salt crystals, do not show up under fluorescent illumination. Crystals have the highest protein concentration and are readily observed against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. Preliminary tests, using model proteins, indicates that we can use high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that more rapid amorphous precipitation kinetics may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Experiments are now being carried out to test this approach using a wider range, of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗

State Event Models for the Formal Analysis of Human-Machine Interactions

The work described in this paper was motivated by our experience with applying a framework for formal analysis of human-machine interactions (HMI) to a realistic model of an autopilot. The framework is built around a formally defined conformance relation called "fullcontrol" between an actual system and the mental model according to which the system is operated. Systems are well-designed if they can be described by relatively simple, full-control, mental models for their human operators. For this reason, our framework supports automated generation of minimal full-control mental models for HMI systems, where both the system and the mental models are described as labelled transition systems (LTS). The autopilot that we analysed has been developed in the NASA Ames HMI prototyping tool ADEPT. In this paper, we describe how we extended the models that our HMI analysis framework handles to allow adequate representation of ADEPT models. We then provide a property-preserving reduction from these extended models to LTSs, to enable application of our LTS-based formal analysis algorithms. Finally, we briefly discuss the analyses we were able to perform on the autopilot model with our extended framework.

State-event Models↗

Automating Bug Report Classification with Few Shot Learning

Orthogonal defect classification (ODC) is a method used to categorize software defects, providing valuable insights into the development process. This study focuses on automating the classification of software bug reports into different ODC defect types using few shot learning, a machine learning approach that requires minimal labeled data. Previous research has manually classified bug reports or used traditional machine learning algorithms like linear support vector machine, achieving limited success. Our approach uses few shot learning to improve classification accuracy and efficiency. The results show a harmonic mean of recall and precision (i.e., the F1 score) of around 0.6 which is a performance improvement over previous methods. The results highlight the potential benefit of few shot learning techniques and their application in enhancing the safety and reliability of nuclear digital instrumentation and control (DI&C) systems. Future work will explore incorporating advanced techniques to supplement the model's training data and achieve better results.

42 - ENGINEERING↗

Fluorescent Approaches to High Throughput Crystallography

We have shown that by covalently modifying a subpopulation, less than or equal to 1%, of a macromolecule with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification, and the presence of the probe at low concentrations does not affect the X-ray data quality or the crystallization behavior. The presence of the trace fluorescent label gives a number of advantages when used with high throughput crystallizations. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination crystals show up as bright objects against a dark background. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Brightly fluorescent crystals are readily found against less bright precipitated phases, which under white light illumination may obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries as the protein or protein structures is all that shows up. Fluorescence intensity is a faster search parameter, whether visually or by automated methods, than looking for crystalline features. We are now testing the use of high fluorescence intensity regions, in the absence of clear crystalline features or "hits", as a means for determining potential lead conditions. A working hypothesis is that kinetics leading to non-structured phases may overwhelm and trap more slowly formed ordered assemblies, which subsequently show up as regions of brighter fluorescence intensity. Preliminary experiments with test proteins have resulted in the extraction of a number of crystallization conditions from screening outcomes based solely on the presence of bright fluorescent regions. Subsequent experiments will test this approach using a wider range of proteins. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons.

Pusey, Marc L.↗

Leveraging machine learning to enhance aerosol classification using Single-Particle Mass Spectrometry

Advancing automated classification of atmospheric aerosols from Single-Particle Mass Spectrometry (SPMS) data remains challenging due to overlapping ion signatures, compositional diversity, and limited labeled data. This study evaluates supervised and semi-supervised learning frameworks to enhance aerosol identification by jointly leveraging labeled and unlabeled spectra. Four models were compared: a supervised Support Vector Machine (SVM), a self-training SVM, a stacked autoencoder classifier, and a stacked autoencoder trained using a temporal-ensembling Mean Teacher approach. All models achieved high and stable accuracies (90.0 %–91.1 %), surpassing previous results on the same dataset (87 %) and matching the performance of state-of-the-art deep learning methods. Despite small global metric differences (≤ 1 %), semi-supervised variants yielded up to 5 %–10 % improvements for compositionally rare particle types – such as soot (0.77 % of spectra, F1-score: 0.93–0.97) and hazelnut pollen (0.98 % of spectra, F1-score: 0.97–1.00) – equating to roughly ∼ 187 additional correctly classified spectra. These gains are scientifically significant, as such rare particles exert disproportionate influence on radiative absorption and ice nucleation processes; their improved detection reduces modeled uncertainties in aerosol absorption optical depth and mixed-phase cloud ice nucleation rates. The models' residual misclassifications (≈ 9 %) largely arise from true spectral overlap among chemically adjacent species (e.g., Na- vs. K-feldspar, coated vs. uncoated feldspars), reflecting physical compositional continuity rather than algorithmic error. Collectively, these findings demonstrate that leveraging unlabeled data to learn robust spectral representations and refine classification enhances both fidelity and interpretability, bridging data-driven analysis with aerosol–climate process understanding.

54 ENVIRONMENTAL SCIENCES↗

Automated anomaly detection for Orbiter High Temperature Reusable Surface Insulation

The description, analysis, and experimental results of a method for identifying possible defects on High Temperature Reusable Surface Insulation (HRSI) of the Orbiter Thermal Protection System (TPS) is presented. Currently, a visual postflight inspection of Orbiter TPS is conducted to detect and classify defects as part of the Orbiter maintenance flow. The objective of the method is to automate the detection of defects by identifying anomalies between preflight and postflight images of TPS components. The initial version is intended to detect and label gross (greater than 0.1 inches in the smallest dimension) anomalies on HRSI components for subsequent classification by a human inspector. The approach is a modified Golden Template technique where the preflight image of a tile serves as the template against which the postflight image of the tile is compared. Candidate anomalies are selected as a result of the comparison and processed to identify true anomalies. The processing methods are developed and discussed, and the results of testing on actual and simulated tile images are presented. Solutions to the problems of brightness and spatial normalization, timely execution, and minimization of false positives are also discussed.

Cooper, Eric G.↗

Characterizing Interference in Radio Astronomy Observations through Active and Unsupervised Learning

In the process of observing signals from astronomical sources, radio astronomers must mitigate the effects of manmade radio sources such as cell phones, satellites, aircraft, and observatory equipment. Radio frequency interference (RFI) often occurs as short bursts (< 1 ms) across a broad range of frequencies, and can be confused with signals from sources of interest such as pulsars. With ever-increasing volumes of data being produced by observatories, automated strategies are required to detect, classify, and characterize these short "transient" RFI events. We investigate an active learning approach in which an astronomer labels events that are most confusing to a classifier, minimizing the human effort required for classification. We also explore the use of unsupervised clustering techniques, which automatically group events into classes without user input. We apply these techniques to data from the Parkes Multibeam Pulsar Survey to characterize several million detected RFI events from over a thousand hours of observation.

Doran, G.↗

Image Labeler: A Web Interface to Catalog Earth Science Events

Advances in machine learning (ML) have made it possible to automatically detect Earth science phenomena from satellite imagery. While useful, ML algorithms typically require an extensive dataset containing labeled images for training. Systematic labeling and management of such datasets is quite cumbersome. With this in mind, we present the Image Labeler. Image Labeler is a fast and scalable cloud-based tool that facilitates the rapid development of Earth science event databases, in order to aid automated ML-based image classification.

Case Study↗

Characterizing Interference in Radio Astronomy Observations through Active and Unsupervised Learning

In the process of observing signals from astronomical sources, radio astronomers must mitigate the effects of man-made radio sources such as cell phones, satellites, aircraft, and observatory equipment. Radio frequency interference (RFI) often occurs as short bursts (< 1 ms) across a broad range of frequencies, and can be confused with signals from sources of interest such as pulsars. With ever-increasing volumes of data being produced by observatories, automated strategies are required to detect, classify, and characterize these short “transient” RFI events. We investigate an active learning approach in which an astronomer labels events that are most confusing to a classifier, minimizing the human effort required for classification. We also explore the use of unsupervised clustering techniques, which automatically group events into classes without user input. We apply these techniques to data from the Parkes Multibeam Pulsar Survey to characterize several million detected RFI events from over a thousand hours of observation

Doran, G.↗