Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Automated labeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Automated Label‐Free Assay for Viral Detection and Inhibitor Screening via Biomembrane‐Functionalized Microelectrode Arrays

Most virus infection assays have indirect readout such as virus number following entry (e.g., PCR, cell lysis). While effective, these technologies are labor‐intensive, require specialized environments (e.g., sterile or RNA‐free), and detect later‐stage viral events like lysis or cell death, lacking sensitivity to early fusion events. To address these limitations, we present biologically relevant 2D membrane materials, host‐cell‐derived supported lipid bilayers (hcd‐SLBs), integrated with organic microelectrode arrays (OMEAs) for detection of severe acute respiratory syndrome coronavirus 2 (SARS‐CoV‐2) fusion. By overexpressing angiotensin‐converting enzyme 2 (ACE2) receptors on the native membranes, the platform functions as a viral sensor capable of detecting virus pseudo particles (VPPs) through the late pathway. Additionally, hcd‐SLBs extracted from human lung epithelium expressing native ACE2 detect fusion events through the early pathway. The platform's utility as a drug‐screening tool is demonstrated by testing antibodies targeting either the ACE2 on the host membrane or the viral spike (S) proteins. To enhance the throughput, microfluidics are integrated for automation and OMEAs are incorporated within each channel, miniaturizing the testing units. This system supports high‐throughput data generation, automation, and scalability, providing an efficient platform for viral fusion detection that advances the study of pathogen‐host interactions and accelerates antiviral drug discovery.

Biology↗

Classifying handedness in chiral nanomaterials using label error robust deep learning

Abstract High-throughput scanning electron microscopy (SEM) coupled with classification using neural networks is an ideal method to determine the morphological handedness of large populations of chiral nanoparticles. Automated labeling removes the time-consuming manual labeling of training data, but introduces label error, and subsequently classification error in the trained neural network. Here, we evaluate methods to minimize classification error when training from automated labels of SEM datasets of chiral Tellurium nanoparticles. Using the mirror relationship between images of opposite handed particles, we artificially create populations of varying label error. We analyze the impact of label error rate and training method on the classification error of neural networks on an ideal dataset and on a practical dataset. Of the three training methods considered, we find that a pretraining approach yields the most accurate results across label error rates on ideal datasets, where size and other morphological variables are held constant, but that a co-teaching approach performs the best in practical application.

36 MATERIALS SCIENCE↗

Automated instant labeling chemistry workflow for real-time monitoring of monoclonal antibody N -glycosylation

With the transition toward continuous bioprocessing, process analytical technology (PAT) is becoming necessary for rapid and reliable in-process monitoring during biotherapeutics manufacturing. Bioprocess 4.0 is looking to build end-to-end bioprocesses that include PAT-enabled real-time process control. This is especially important for drug product quality attributes that can change during bioprocessing, such as protein N-glycosylation, a critical quality attribute for most monoclonal antibody (mAb) therapeutics. Glycosylation of mAbs is known to influence their efficacy as therapeutics and is regulated for a majority of mAb products on the market today. Currently, there is no method to truly measure N-glycosylation using on-line PAT, hence making it impractical to design upstream process control strategies. We recently described the N-GLYcanyzer workflow: an integrated PAT unit that measures mAb N-glycosylation within 3 hours of automated sampling from a bioreactor. Here, we integrated Agilent's Instant Procainamide (InstantPC) based chemistry workflow into the N-GLYcanyzer PAT unit to allow for nearly 10× faster near real-time analysis of mAb glycoforms. Furthermore, our methodology is explained in detail to allow for replication of the PAT workflow as well as present a case study demonstrating the use of this PAT to autonomously monitor a mammalian cell perfusion process at the bench scale to gain increased knowledge of mAb glycosylation dynamics during continuous biologics manufacturing using Chinese hamster ovary (CHO) cells.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AI-Enabled Operations at Fermi Complex: Multivariate Time Series Prediction for Outage Prediction and Diagnosis

The Main Control Room of the Fermilab accelerator complex continuously gathers extensive time-series data from thousands of sensors monitoring the beam. However, unplanned events such as trips or voltage fluctuations often result in beam outages, causing operational downtime. This downtime not only consumes operator effort in diagnosing and addressing the issue but also leads to unnecessary energy consumption by idle machines awaiting beam restoration. The current threshold-based alarm system is reactive and faces challenges including frequent false alarms and inconsistent outage-cause labeling. To address these limitations, we propose an AI-enabled framework that leverages predictive analytics and automated labeling. Using data from $2,703$ Linac devices and $80$ operator-labeled outages, we evaluate state-of-the-art deep learning architectures, including recurrent, attention-based, and linear models, for beam outage prediction. Additionally, we assess a Random Forest-based labeling system for providing consistent, confidence-scored outage annotations. Our findings highlight the strengths and weaknesses of these architectures for beam outage prediction and identify critical gaps that must be addressed to fully harness AI for transitioning downtime handling from reactive to predictive, ultimately reducing downtime and improving decision-making in accelerator management.

Jain, Milan [PNL, Richland] (ORCID:000000021676111↗

Automated Credibility Assessments of User Features in Scientific Software

Scientific software (SciSoft) is complex, often containing a mixture of production capabilities co-mingled with features under active research and development. Furthermore, SciSoft is often developed over decades by non-computer scientists who may not have a strong background in or prioritize software architecture design, testing, and quality (e.g., test coverage). These conditions lead to difficulty in understanding which software components or functions implement what user-facing features and therefore those features’ software quality pedigree. This lack of understanding poses challenges in assessing readiness and credibility of user features, and often relies on a SciSoft subject matter expert’s (SME) laborious investigation and assertion. This final report of a one-year Computing and Information Sciences Lab Directed Research and Development project presents a general framework for modeling SciSoft architecture as a direct relationship between user features and the software components/functions that implement them. Our approach leverages automated labeling of the SciSoft’s regression test suite and employs machine learning algorithms to construct the architecture model. We demonstrate this framework on the Solid Mechanics component of the SIERRA multi-physics engineering analysis suite developed at Sandia National Laboratories.

97 MATHEMATICS AND COMPUTING↗

Active Learning Meets Foundation Models: Fast Remote Sensing Data Annotation for Object Detection

Object detection in remote sensing demands extensive, high-quality annotations—a process that is both labor-intensive and time-consuming. In this work, we introduce a real-time active learning and semi-automated labeling framework that leverages foundation models to streamline dataset annotation for object detection in remote sensing imagery. For example, by integrating a Segment Anything Model (SAM), our approach generates mask-based bounding boxes that serve as the basis for dual sampling: (a) uncertainty estimation to pinpoint challenging samples, and (b) diversity assessment to ensure broad data coverage. Furthermore, our Dynamic Box Switching Module (DBS) addresses the well-known cold start problem for object detection models by replacing its suboptimal initial predictions with SAM-derived masks, thereby enhancing early-stage localization accuracy. Extensive evaluations on multiple remote sensing datasets plus a real-world user study, demonstrate that our framework not only reduces annotation effort, but also significantly boosts detection performance compared to traditional active learning sampling methods. The code for training and the user interface will be made available.

Burges, Marvin [ORNL] (ORCID:0000000312690769)↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

The L-CAPE Project at FNAL

The controls system at FNAL records data asynchronously from several thousand Linac devices at their respective cadences, ranging from 15Hz down to once per minute. In case of downtimes, current operations are mostly reactive, investigating the cause of an outage and labeling it after the fact. However, as one of the most upstream systems at the FNAL accelerator complex, the Linac’s foreknowledge of an impending downtime as well as its duration could prompt downstream systems to go into standby, potentially leading to energy savings. The goals of the Linac Condition Anomaly Prediction of Emergence (L-CAPE) project that started in late 2020 are (1) to apply data-analytic methods to improve the information that is available to operators in the control room, and (2) to use machine learning to automate the labeling of outage types as they occur and discover patterns in the data that could lead to the prediction of outages. We present an overview of the challenges in dealing with time-series data from 2000+ devices, our approach to developing an ML-based automated outage labeling system, and the status of augmenting operations by identifying the most likely devices predicting an outage.

43 PARTICLE ACCELERATORS↗

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation↗

Creating a Training Dataset for Semantic Segmentation of Canal Networks for Irrigation Modernization

Canal infrastructure has provided critical irrigation water to the western United States for over a century. To continue providing vital water resources to the semi-arid West, irrigation systems must undergo maintenance and modernization. Many canal companies are resource-constrained, and because funding opportunities often require detailed knowledge of existing infrastructure, they can struggle to secure financial capital. We address this problem by creating training data for a semantic segmentation deep learning model to map canal networks throughout the western United States. To create a diverse and robust training dataset, we labelled 1-m NAIP imagery with the locations of no canals, wet canals, and dry/vegetated canals. Since creating these datasets is time consuming, we first developed a preprocessing methodology to identify canals within our four study areas. We used NAIP imagery and provided canal centerline data to buffer, standardize, and cluster the imagery, automating the labeling process as much as possible. However, this still required manual cleaning and manual classification of canal type. Challenges arose when canals were interrupted (e.g., road culverts or piped sections) or when nearby features shared similar characteristics (e.g., irrigated fields, trees, and shadows). Combining automated preprocessing with manual refinement produced four detailed canal masks to be used in the semantic segmentation model developed by Richard Tapia.

13 - HYDRO ENERGY↗

The Geothermal Artificial Intelligence for geothermal exploration

Exploration of geothermal resources involves analysis and management of a large number of uncertainties, which makes investment and operations decisions challenging. Remote Sensing (RS), Machine Learning (ML) and Artificial Intelligence (AI) have potential in managing the challenges of geothermal exploration. In this paper, we present a methodology that integrates RS, ML and AI to create an initial assessment of geothermal potential, by resorting to known indicators of geothermal areas namely mineral markers, surface temperature, faults and deformation. We demonstrated the implementation of the method in two sites (Brady and Desert Peak geothermal sites) that are close to each other but have different characteristics (Brady having clear surface manifestations and Desert Peak being a blind site). Here, we processed various satellite images and geospatial data for mineral markers, temperature, faults and deformation and then implemented ML methods to obtain pattern of surface manifestation of geothermal sites. We developed an AI that uses patterns from surface manifestations to predict geothermal potential of each pixel. We tested the Geothermal AI using independent data sets obtaining accuracy of 92-95%; also tested the Geothermal AI trained on one site by executing it for the other site to predict the geothermal / non-geothermal delineation, the Geothermal AI performed quite well in prediction with 72-76% accuracy.

15 GEOTHERMAL ENERGY↗

A hybrid machine-learning approach for analysis of methane hydrate formation dynamics in porous media with synchrotron CT imaging

Fast multi-phase processes in methane hydrate bearing samples pose a challenge for quantitative micro-computed tomography study and experiment steering due to complex tomographic data analysis involving time-consuming segmentation procedures. This is because of the sample's multi-scale structure, which changes over time, low contrast between solid and fluid materials, and the large amount of data acquired during dynamic processes. Here, a hybrid approach is proposed for the automatic segmentation of tomographic data from time-resolved imaging of methane gas-hydrate formation in sandy granular media, which includes a deep-learning 3D U-Net model. To prepare a training dataset for the 3D U-Net, a technique to automate data labeling based on sample-specific information about the mineral matrix immobility and occasional fluid movement in pores is proposed. Automatic segmentation allowed for studying properties of the hydrate growth in pores, as well as dynamic processes such as incremental flow and redistribution of pore brine. Results of the quantitative analysis showed that for typical gas-hydrate stability parameters (100 bar methane pressure, 7°C temperature) the rate of formation is slow (less than 1% per hour), after which the surface area of contact between brine and gas increases, resulting in faster formation (2.5% per hour). Hydrate growth reaches the saturation point after 11 h of the experiment. Finally, the efficacy of the proposed segmentation scheme in on-the-fly automatic data analysis and experiment steering with zooming to regions of interest is demonstrated.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

End-to-end deep learning pipeline for real-time Bragg peak segmentation: from training to large-scale deployment

X-ray crystallography reconstruction, which transforms discrete X-ray diffraction patterns into three-dimensional molecular structures, relies critically on accurate Bragg peak finding for structure determination. As X-ray free electron laser (XFEL) facilities advance toward MHz data rates (1 million images per second), traditional peak finding algorithms that require manual parameter tuning or exhaustive grid searches across multiple experiments become increasingly impractical. While deep learning approaches offer promising solutions, their deployment in high-throughput environments presents significant challenges in automated dataset labeling, model scalability, edge deployment efficiency, and distributed inference capabilities. We present an end-to-end deep learning pipeline with three key components: (1) a data engine that combines traditional algorithms with our peak matching algorithm to generate high-quality training data at scale, (2) a modular architecture that scales from a few million to hundreds of million parameters, enabling us to train large expert-level models offline while deploying smaller, distilled models at the edge, and (3) a decoupled producer-consumer architecture that separates specialized data source layer from model inference, enabling flexible deployment across diverse computing environments. Using this integrated approach, our pipeline achieves accuracy comparable to traditional methods tuned by human experts while eliminating the need for experiment-specific parameter tuning. Although current throughput requires optimization for MHz facilities, our system's scalable architecture and demonstrated model compression capabilities provide a foundation for future high-throughput XFEL deployments.

Wang, Cong↗

Prime agricultural land monitoring and assessment component of the California Integrated Remote Sensing System

The use of digital LANDSAT techniques for monitoring agricultural land use conversions was studied. Two study areas were investigated: one in Ventura County and the other in Fresno County (California). Ventura test site investigations included the use of three dates of LANDSAT data to improve classification performance beyond that previously obtained using single data techniques. The 9% improvement is considered highly significant. Also developed and demonstrated using Ventura County data is an automated cluster labeling procedure, considered a useful example of vertical data integration. Fresno County results for a single data LANDSAT classification paralleled those found in Ventura, demonstrating that the urban/rural fringe zone of most interest is a difficult environment to classify using LANDSAT data. A general raster to vector conversion program was developed to allow LANDSAT classification products to be transferred to an operational county level geographic information system in Fresno.

Estes, J. E.↗

A Hybrid Approach to Labeling Datasets in Earth Science Publications

NASA Data Centers provide the public with thousands of datasets that result in published papers, reports, and conference proceedings. Collecting accurate metrics on usage of these datasets is key to connecting different areas of knowledge and evaluating the datasets’ impact. While most of the datasets have Digital Object Identifiers (DOIs) assigned, most publications do not cite them hampering the automated search of these publications. Instead, articles mention attributes like organization, instrument, mission, variable, or a publication describing the dataset. Often only domain experts can deduce the dataset that was used in the publication text. The lack of a citation slows the spread of information and reduces the research’s impact. With thousands of papers produced each year, an automated means of labeling datasets is critical. This paper explores a hybrid approach of heuristics and a Natural Language Processing (NLP) Named Entity Recognition (NER) model to find and label the datasets used within Earth Science papers. Heuristics are used to produce the labelled sentences and any potential dataset candidates that can be derived from a sentence. The heuristic labels the sentences with the names of mission, instrument, re-analysis models, and science keywords taken from the Global Change Master Directory (GCMD) ontology. Additionally, it uses those labels to generate the dataset citation candidates. If the mission, instrument, and variable are sufficient to create the citation for the dataset the citation and the label the domain expert reviews the output without going through the NLP model. If the extracted label is not sufficient to label the dataset on its own, the sentence and its associated dataset labels will be inputted into the NER model. The model outputs the labeled sentence and the potential dataset candidates with their associated probabilities. The domain expert then reviews the NER model’s output and the correct labels are determined. The newly labelled papers can then be used as additional training data. This creates an iterative process for the approach to continuously improve. Because all the possible mentions are gathered by the model, the domain expert can quickly and easily label the papers resulting in large time savings.

Jacob Atkins↗

An automated liquid jet for fluorescence dosimetry and microsecond radiolytic labeling of proteins

X-ray radiolytic labeling uses broadband X-rays for in situ hydroxyl radical labeling to map protein interactions and conformation. High flux density beams are essential to overcome radical scavengers. However, conventional sample delivery environments, such as capillary flow, limit the use of a fully unattenuated focused broadband beam. An alternative is to use a liquid jet, and we have previously demonstrated that use of this form of sample delivery can increase labeling by tenfold at an unfocused X-ray source. Here we report the first use of a liquid jet for automated inline quantitative fluorescence dosage characterization and sample exposure at a high flux density microfocused synchrotron beamline. Our approach enables exposure times in single-digit microseconds while retaining a high level of side-chain labeling. This development significantly boosts the method’s overall effectiveness and efficiency, generates high-quality data, and opens up the arena for high throughput and ultrafast time-resolved in situ hydroxyl radical labeling.

59 BASIC BIOLOGICAL SCIENCES↗

Automated 3D cytoplasm segmentation in soft X-ray tomography

Cells’ structure is key to understanding cellular function, diagnostics, and therapy development. Soft X-ray tomography (SXT) is a unique tool to image cellular structure without fixation or labeling at high spatial resolution and throughput. Fast acquisition times increase demand for accelerated image analysis, like segmentation. Currently, segmenting cellular structures is done manually and is a major bottleneck in the SXT data analysis. This paper introduces ACSeg, an automated 3D cytoplasm segmentation model. ACSeg is generated using semi-automated labels and 3D U-Net and is trained on 43 SXT tomograms of immune T cells, rapidly converging to high-accuracy segmentation, therefore reducing time and labor. Furthermore, adding only 6 SXT tomograms of other cell types diversifies the model, showing potential for optimal experimental design. ACSeg successfully segmented unseen tomograms and is published on Biomedisa, enabling high-throughput analysis of cell volume and structure of cytoplasm in diverse cell types.

59 BASIC BIOLOGICAL SCIENCES↗

Automated microbial metabolism laboratory

The labeled release concept was advanced to accommodate a post- Viking mission designed to extend the search, to confirm the presence of, and to characterize any Martian life found, and to obtain preliminary information on control of the life detected. The advanced labeled release concept utilizes four test chambers, each of which contains either an active or heat sterilized sample of the Martian soil. A variety of C-14 labeled organic substrates can be added sequentially to each soil sample and the resulting evolved radioactive gas monitored. The concept can also test effects of various inhibitors and environmental parameters on the experimental response. The current Viking '75 labeled release hardware is readily adaptable to the advanced labeled release concept.

Source record↗