Engineering PapersSearch

SEARCH · Engineering Papers

Results for “labeled data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Using partially labeled data for normal mixture identification with application to class definition

The problem of estimating the parameters of a normal mixture density when, in addition to the unlabeled samples, sets of partially labeled samples are available is addressed. The density of the multidimensional feature space is modeled with a normal mixture. It is assumed that the set of components of the mixture can be partitioned into several classes and that training samples are available from each class. Since for any training sample the class of origin is known but the exact component of origin within the corresponding class is unknown, the training samples as considered to be partially labeled. The EM iterative equations are derived for estimating the parameters of the normal mixture in the presence of partially labeled samples. These equations can be used to combine the supervised and nonsupervised learning processes.

Shahshahani, Behzad M.

Attention-based functional-group coarse-graining: a deep learning framework for molecular prediction and design

Machine learning (ML) offers considerable promise for the design of new molecules and materials. In real-world applications, the design problem is often domain-specific, and suffers from insufficient data, particularly labeled data, for ML training. In this study, we report a data-efficient, deep-learning framework for molecular discovery that integrates a coarse-grained functional-group representation with a self-attention mechanism to capture intricate chemical interactions. Our approach exploits group-contribution concepts to create a graph-based intermediate representation of molecules, serving as a low-dimensional embedding that substantially reduces the data demands typically required for training. Using a self-attention mechanism to learn the subtle but highly relevant chemical context of functional groups, the method proposed here consistently outperforms existing approaches for predictions of multiple thermophysical properties. In a case study focused on adhesive polymer monomers, we train on a limited dataset comprising only 6,000 unlabeled and 600 labeled monomers. The resulting chemistry prediction model achieves over 92% accuracy in forecasting properties directly from SMILES strings, exceeding the performance of current state-of-the-art techniques. Furthermore, the latent molecular embedding is invertible, enabling the design pipeline to automatically generate new monomers from the learned chemical subspace. We illustrate this functionality by targeting several properties, including high and low glass transition temperatures (Tg), and demonstrate that our model can identify new candidates with values that surpass those in the training set. The ease with which the proposed framework navigates both chemical diversity and data scarcity offers a promising route to accelerate and broaden the search for functional materials.

Han, Ming [Univ. of Chicago, IL (United States)]

The Fifth Calibration/Data Product Validation Panel Meeting

The minutes and associated documents prepared from presentations and meetings at the Fifth Calibration/Data Product Validation Panel meeting in Boulder, Colorado, April 8 - 10, 1992, are presented. Key issues include (1) statistical characterization of data sets: finding statistics that characterize key attributes of the data sets, and defining ways to characterize the comparisons among data sets; (2) selection of specific intercomparison exercises: selecting characteristic spatial and temporal regions for intercomparisons, and impact of validation exercises on the logistics of current and planned field campaigns and model runs; and (3) preparation of data sets for intercomparisons: characterization of assumptions, transportable data formats, labeling data files, content of data sets, and data storage and distribution (EOSDIS interface).

Source record

Cell Kinetic and Histomorphometric Analysis of Microgravitational Osteopenia: PARE.03B

Previous methods of identifying cells undergoing DNA synthesis (S-phase) utilized H-3 thymidine (3HT) autoradiography. 5-Bromo-2'-deoxyuridine (BrdU) immunohistochemistry is a nonradioactive alternative method. This experiment compared the two methods using the nuclear volume model for osteoblast histogenesis in two different embedding media. Twenty Sprague-Dawley rats were used, with half receiving 3HT (1 micro Ci/g) and the other half BrdU (50 microgram/g). Condyies were embedded (one side in paraffin, the other in plastic) and S-phase nuclei were identified using either autoradiography or immunohistochemistry. The fractional distribution of preosteoblast cell types and the percentage of labeled cells (within each cell fraction and label index) were calculated and expressed as mean q standard error. Chi-Square analysis showed only a minor difference in the fractional distribution of cell types. However, there were significant differences (p less than 0.05) by ANOVA, in the nuclear labeling of specific cell types. With the exception of the less-differentiated A+A'cells, more BrdU label was consistently detected in paraffin than in plastic-embedded sections. In general, more nuclei were labeled with 3H-thymidine than with BrdU in both types of embedding media. Labeling index data (labeled cells/total cells sampled x 100) indicated that BrdU in paraffin, but not plastic gave the same results as 3HT in either embedding method. Thus, we conclude that the two labeling methods do not yield the same results for the nuclear volume model and that embedding media is an important factor whenusing BrdU. As a result of this work, 3HT was chosen for used in the PARE.03 flight experiments.

Roberts, W. Eugene

Cell Kinetic and Histomorphometric Analysis of Microgravitational Osteopenia: PARE.03B

Previous methods of identifying cells undergoing DNA synthesis (S-phase) utilized 3H-thymidine (3HT) autoradiography. 5-Bromo-2'-deoxyuridine (BrdU) immunohistochemistry is a nonradioactive alternative method. This experiment compared the two methods using the nuclear volume model for osteoblast histogenesis in two different embedding media. Twenty Sprague-Dawley rats were used, with half receiving 3HT (1 micro-Ci/g) and the other half BrdU (50 micro-g/g). Condyles were embedded (one side in paraffin, the other in plastic) and S-phase nuclei were identified using either autoradiography or immunohistochemistry. The fractional distribution of preosteoblast cell types and the percentage of labeled cells (within each cell fraction and label index) were calculated and expressed as mean +/- standard error. Chi-Square analysis showed only a minor difference in the fractional distribution of cell types. However, there were,significant differences (p less than 0.05) by ANOVA, in the nuclear labeling of specific cell types. With the exception of the less-differentiated A+A' cells, more BrdU label was consistently detected in paraffin than in plastic-embedded sections. In general, more nuclei were labeled with 3H-thymidine than with BrdU in both types of embedding media (Fig 2.). Labeling index data (labeled cells/total cells sampled x 100) indicated that BrdU in paraffin, but not plastic gave the same results as 3HT in either embedding method. Thus, we conclude that the two labeling methods do not yield the same results.

Roberts, W. Eugene

Positive and Negative Ion CIMS Measurements, SGP, May-June 2024

The negative ion measurements were taken using two LTOF-CIMS with two different nitrate inlets. The data labeled Aerodyne_Inlet used an Aerodyne nitrate inlet. Data labeled PCC used a custom transverse inlet. The positive ion measurements were taken using one LTOF-CIMS with hydronium as an ionization reagent using a custom transverse inlet with an inlet voltage difference of 700 V. See https://doi.org/10.1021/acs.jpca.2c01672 for more information. Peaks were autofit using tofware. Positive ion data is published in the form of peak signal/reagent signal. Data files of peak signal/reagent signal for the two mass spectrometers for the negative ion measurements are separate. Concentrations of H2SO4 are combined into one file. When measurements were overlapping, measurements from the Aerodyne nitrate inlet were used.

54 ENVIRONMENTAL SCIENCES

Value of Information App (Value of Information App for Binary Geothermal Decisions and Binary Geothermal Possibilities) (Negative/Positive) [SWR-25-15]

Code base to run Streamlit Value of Information App for binary decision with geothermal techno economics. An open-source VOI app that models binary decisions (e.g. do something (drill) or walk away (do nothing)) and binary geothermal scenarios (positive or negative) has been developed. Users can input their anticipated economic values (profits or losses) directly into the value matrix to represent all four combinations of these actions and geothermal possibilities. VOI in general requires probabilities to be assigned for “probability of success”, or probability of experiencing a positive geothermal scenario versus negative. The users of the App can toggle this probability of success both in the demo problem and in the Value of Imperfect Information problem. The VOI App allows users to upload their own labeled data to evaluate how well it allows them to distinguish between positive versus negative sites. We have been using IGNENIOUS data to test and demonstrate; industry members have prepared their own labeled data, and have present their examples from diverse use cases at a conference workshop. The VOI App is open to the public at: https://voigeothermalrising.streamlit.app

Trainor-Guitton, Whitney [National Renewable Energ

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat

Development of biological and nonbiological explanations for the Viking label release data

The plausibility that hydrogen peroxide, widely distributed within the Mars surface material, was responsible for the evocative response obtained by the Viking Labeled Release (LR) experiment on Mars was investigated. Although a mixture of gamma Fe2O3 and silica sand stimulated the LR nutrient reaction with hydrogen peroxide and reduced the rate of hydrogen decomposition under various storage conditions, the Mars analog soil prepared by the Viking Inorganic Analysis Team to match the Mars analytical data does not cause such effects. Nor is adequate resistance to UV irradiation shown. On the basis of the results and consideration presented while the hydrogen peroxide theory remains the most, if not only, attractive chemical explanation of the LR data, it remains unconvincing on critical points. Until problems concerning the formation and stabilization of hydrogen peroxide on the surface of Mars can be overcome, adhere to the scientific evidence requires serious consideration of the biological theory.

Source record

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science

Leveraging 13C-Labeling to Assign Molecular Formulas to Unknown Yeast Metabolites

Mass spectrometry analyses have identified tens of thousands of unknown small molecule-associated peaks in different biological specimens. Notably, even the simplest and best studied organisms like Escherichia coli and Saccharomyces cerevisiae yield thousands of unknown peaks. A key question is how many of these reflect actual novel endogenous metabolites. To explore this, Mahieu and Patti used complete 13 C -labeling in E. coli to credential peaks as biological. This reduced the number of unknowns by more than 90%. Here, we carry out similar uniform 13 C-labeling in the Baker’s yeast S. cerevisiae and two less-studied bioenergy-relevant yeasts Rhodotorula toruloides (lipid producer) and Issatchenkia orientalis (organic acid producer). Identification of unknown metabolite peaks and their molecular formulas is facilitated through software tailored for 13 C labeling data and resulting knowledge of carbon atom count. A classification model evaluates the plausibility of each candidate formula, with peaks lacking plausible candidate formulas unlikely to reflect metabolite molecular ions. This approach prioritizes about one hundred candidate abundant unknown metabolites with logical molecular formulas. Most of these are species-specific rather than conserved across yeasts, and more are found in the nonmodel yeasts than S. cerevisiae. Thus, 13 C-labeling data on unknown metabolites highlights the potential for discovering new metabolites and pathways in nonmodel yeasts.

Carbon

Advanced Semi-Supervised Learning with Uncertainty Estimation for Phase Identification in Distribution Systems

The integration of advanced metering infrastructure (AMI) into power distribution networks generates valuable data for tasks such as phase identification; however, the limited and unreliable availability of labeled data in the form of customer phase connectivity presents challenges. To address this issue, we propose a semi-supervised learning (SSL) framework that effectively leverages labeled and unlabeled data. Our approach incorporates self-training, label spreading, and Bayesian neural networks (BNNs) to enhance phase identification with AMI data. Our method uses an ensemble of multilayer perceptron classifiers in a self-training setup, iteratively adding high-confidence pseudo-labels to improve robustness. We also apply label spread to propagate labels based on data similarity, which enhances generalization across diverse distributions. In addition, we employ a BNNs with uncertainty estimation, boosting confidence in predictions and reducing phase identification errors. In our case study, we achieved approximately 98% +/- 0.08 accuracy with uncertainty using minimal and unreliable labeled data from a real U.S. utility, Duquesne Light Company. Our SSL approach, combined with uncertainty estimation, provides an efficient solution for phase identification in AMI data, ultimately improving the reliability of smart grid applications.

24 POWER TRANSMISSION AND DISTRIBUTION

Advanced Semi-Supervised Learning With Uncertainty Estimation for Phase Identification in Distribution Systems

The integration of advanced metering infrastructure (AMI) into power distribution networks generates valuable data for tasks such as phase identification; however, the limited and unreliable availability of labeled data in the form of customer phase connectivity presents challenges. To address this issue, we propose a semi-supervised learning (SSL) framework that effectively leverages labeled and unlabeled data. Our approach incorporates self-training, label spreading, and Bayesian neural networks (BNNs) to enhance phase identification with AMI data. Our method uses an ensemble of multilayer perceptron classifiers in a self-training setup, iteratively adding high-confidence pseudo-labels to improve robustness. We also apply label spread to propagate labels based on data similarity, which enhances generalization across diverse distributions. In addition, we employ a BNNs with uncertainty estimation, boosting confidence in predictions and reducing phase identification errors. In our case study, we achieved approximately 98% +/- 0.08 accuracy with uncertainty using minimal and unreliable labeled data from a real U.S. utility, Duquesne Light Company. Our SSL approach, combined with uncertainty estimation, provides an efficient solution for phase identification in AMI data, ultimately improving the reliability of smart grid applications.

24 POWER TRANSMISSION AND DISTRIBUTION

Science information systems: Archive, access, and retrieval

The objective of this research is to develop technology for the automated characterization and interactive retrieval and visualization of very large, complex scientific data sets. Technologies will be developed for the following specific areas: (1) rapidly archiving data sets; (2) automatically characterizing and labeling data in near real-time; (3) providing users with the ability to browse contents of databases efficiently and effectively; (4) providing users with the ability to access and retrieve system independent data sets electronically; and (5) automatically alerting scientists to anomalies detected in data.

Campbell, William J.

Curation and Dissemination of Complex Multi-Modal Datasets for Radiation Detection, Localization, and Tracking

The PANDAWN sensor network in Chicago, IL, is a state-of-the-art testbed for networked, multi-modal sensing. It integrates AI/data science methods into its operation, from data acquisition to automated data labeling and curation workflows. The curation and dissemination of diverse multi-modal datasets will enable the development of new radiological/nuclear (R/N) detection, localization, and tracking algorithms and methods relevant across the nonproliferation mission space. This article first introduces the PANDAWN sensor network and the features that make it stand out from previous multi-modal data acquisition efforts. We then review the various data streams acquired on the PANDAWN nodes and present the implementation of an automated data curation pipeline that includes the labeling of radiation and contextual data streams. Here, we finally provide a short overview of different studies that leveraged the curated datasets.

Data curation

AI Applications to Physics Experiments at Jefferson Lab

We survey how AI/ML is being deployed across Jefferson Lab's experimental and accelerator programs. In EPSCI, Hydra applies computer vision to automate real-time data-quality monitoring across all four experimental halls, replacing manual inspection of hundreds to thousands of histograms per shift. AIEC (AI Experiment Controls) uses ML to stabilize drift chamber gains and is now part of standard CEBAF production running, while AI Optimized Polarization (AIOP) targets autonomous control of polarized targets and photon beam angular alignment. In CASA, cavity fault classification models identify faulted cavities and trip types from waveform data with ~85% and ~78% agreement to labeled data, respectively, and are deployed in production; a separate effort applies LLMs and hybrid search to make the CEBAF operations logbook AI-ready. QCD-focused work includes transformer- and GAN-based generative models for particle-level event simulation, with distributed GAN training scaling studies on Polaris. Additional efforts span ML-on-FPGA for the EIC and a new Data Science Department coordinating anomaly detection, uncertainty quantification, and HPC-scalable ML lab-wide. Collectively, these projects illustrate AI's growing role in improving efficiency across JLab's nuclear physics mission.

Mei, Xinxin [Thomas Jefferson National Accelerator