Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES↗

Investigating the Future of Scientific Data Search [Slides]

Searching for usable, actionable, data in a trustworthy manner is a challenge across scientific communities. Artificial Intelligence (AI) and Machine Learning (ML) techniques may be leveraged to increase the utility of scientific data by: Demystify unstructured data to aid curation & sharing Surfacing hard to find datasets. User Experience (UX) Research can help uncover scientists needs & challenges finding data and using AI/ML enabled tools.

97 MATHEMATICS AND COMPUTING↗

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES↗

Automating data analysis for hydrogen/deuterium exchange mass spectrometry using data-independent acquisition methodology

We present a hydrogen/deuterium exchange workflow coupled to tandem mass spectrometry (HX-MS 2 ) that supports the acquisition of peptide fragment ions alongside their peptide precursors. The approach enables true auto-curation of HX data by mining a rich set of deuterated fragments, generated by collisional-induced dissociation (CID), to simultaneously confirm the peptide ID and authenticate MS 1 -based deuteration calculations. The high redundancy provided by the fragments supports a confidence assessment of deuterium calculations using a combinatorial strategy. The approach requires data-independent acquisition (DIA) methods that are available on most MS platforms, making the switch to HX-MS 2 straightforward. Importantly, we find that HX-DIA enables a proteomics-grade approach and wide-spread applications. Considerable time is saved through auto-curation and complex samples can now be characterized and at higher throughput. We illustrate these advantages in a drug binding analysis of the ultra-large protein kinase DNA-PKcs, isolated directly from mammalian cells.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-Species Complex and Standard Metabolomic Samples with Verified Truth Annotations Dataset

This dataset contains 4523251 (~6.35 GB) metabolite-spectra matches following identification with CoreMS. Data were manually curated as true positives, true negatives, or unknowns. Calculations for spectral similarity scores were carried out with two methods for a total of ~12.7 GB (6.35 * 2) of data. They are all .tsv files, though can easily be changed to .txt. The file types are: * human cerebrospinal fluid (CSF), human blood plasma human urine: already published here https://www.nature.com/articles/s41597-021-00894-y, • purchased FAMES standards • fungi species (A. niger, A. nidulans, T. reesei) • soil crust

59 BASIC BIOLOGICAL SCIENCES↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Global root traits (GRooT) database

Motivation: Trait data are fundamental to the quantitative description of plant form and function. Although root traits capture key dimensions related to plant responses to changing environmental conditions and effects on ecosystem processes, they have rarely been included in large-scale comparative studies and global models. For instance, root traits remain absent from nearly all studies that define the global spectrum of plant form and function. Thus, to overcome conceptual and methodological roadblocks preventing a widespread integration of root trait data into large-scale analyses we created the Global Root Trait (GRooT) Database. GRooT provides readyto- use data by combining the expertise of root ecologists with data mobilization and curation. Specifically, we (a) determined a set of core root traits relevant to the description of plant form and function based on an assessment by experts, (b) maximized species coverage through data standardization within and among traits, and (c) implemented data quality checks. Main types of variables contained: GRooT contains 114,222 trait records on 38 continuous root traits. Spatial location and grain: Global coverage with data from arid, continental, polar, temperate and tropical biomes. Data on root traits were derived from experimental studies and field studies. Time period and grain: Data were recorded between 1911 and 2019. Major taxa and level of measurement: GRooT includes root trait data for which taxonomic information is available. Trait records vary in their taxonomic resolution, with subspecies or varieties being the highest and genera the lowest taxonomic resolution available. It contains information for 184 subspecies or varieties, 6,214 species, 1,967 genera and 254 families. Owing to variation in data sources, trait records in the database include both individual observations and mean values. Software format: GRooT includes two csv files. A GitHub repository contains the csv files and a script in R to query the database.

59 BASIC BIOLOGICAL SCIENCES↗

INGENIOUS Thermal Conductivity Measurement Source Categorization

Thermal conductivity (TC) data taken for different wells at a specified drill depth. This is an abridged version of the complete SMU heat flow database, downloaded from the SMU node of the NGDS at the beginning of INGENIOUS (approximately April 2021), and filtered to the INGENIOUS study area. This National Geothermal Data System (NGDS) project aggregates geothermal data collected and curated by the SMU Geothermal Laboratory and its partner organizations. All columns in this database are the same as the SMU database, except for 2 additions associated with this project. Repeated columns are for data correlation purposes. Column descriptions and data types are the same as previous iterations of the SMU database. The new values that are the addition are two new columns developed as part of the INGENIOUS project: INGENIOUS TC Value | INGENIOUS notes INGENIOUS notes are individual notes that were written for specific data points during the analysis process. There are not always notes associated with each input value. INGENIOUS TC Value includes 4 values: 1. Assumed Measured These are values that are assumed to be measured thermal conductivity values, either within a specific well or within the same study region. Many of these have either a published reference, a reported standard deviation, or a unique thermal conductivity value. 2. Data release - assumed measured These are values in the SMU database that are from proprietary data that were added to the SMU database and are labeled as data release for their reference. These values were searched for in person at the SMU Geothermal Laboratory as well as virtual examination of data available on the NGDS. For many of these, there are reported thermal conductivity values associated with the heat flow data in the database, but no specific table or reference to measurements in the original data release files. 3. Known measured These are values that have a reported measurement, either as an original file in the SMU data files on the NGDS or a reported table in a publication. In the rare circumstances, Maria Richards or David Blackwell confirmed measurement. Confirmation of measurement would be written in the INGENIOUS notes column. 4. Unmeasured Unmeasured values are those that are known to be unmeasured, either estimated from another report or no information given. In the SMU database, there are wells that have a heat flow but no thermal conductivity. These are categorized as unmeasured. There are also heat flow values that are stated to have estimated or generalized average thermal conductivity values for the region and rock type. Because these are known to be unmeasured, they are categorized as such. 5. Blank Blank values are either A quality or X quality. These quality values are stated in the INGENIOUS notes. These values were not going to change associated with the heat flow analysis, so these were not examined.

15 GEOTHERMAL ENERGY↗

Polymer informatics: Current status and critical next steps

Artificial intelligence (AI) based approaches are beginning to impact several domains of human life, science and technology. Polymer informatics is one such domain where AI and machine learning (ML) tools are being used in the efficient development, design and discovery of polymers. Surrogate models are trained on available polymer data for instant property prediction, allowing screening of promising polymer candidates with specific target property requirements. Questions regarding synthesizability, and potential (retro)synthesis steps to create a target polymer, are being explored using statistical means. Data-driven strategies to tackle unique challenges resulting from the extraordinary chemical and physical diversity of polymers at small and large scales are being explored. Other major hurdles for polymer informatics are the lack of widespread availability of curated and organized data, and approaches to create machine-readable representations that capture not just the structure of complex polymeric situations but also synthesis and processing conditions. Methods to solve inverse problems, wherein polymer recommendations are made using advanced AI algorithms that meet application targets, are being investigated. As various parts of the burgeoning polymer informatics ecosystem mature and become integrated, efficiency improvements, accelerated discoveries and increased productivity can result. Here in this paper, we review emergent components of this polymer informatics ecosystem and discuss imminent challenges and opportunities.

36 MATERIALS SCIENCE↗

Comparison of Machine Learning and Deep Learning for View Identification from Cardiac Magnetic Resonance Images

Background: Artificial intelligence is increasingly utilized to aid in the interpretation of cardiac magnetic resonance (CMR) studies. One of the first steps is the identification of the imaging plane depicted, which can be achieved by both deep learning (DL) and classical machine learning (ML) techniques without user input. We aimed to compare the accuracy of ML and DL for CMR view classification and to identify potential pitfalls during training and testing of the algorithms. Methods: To train our DL and ML algorithms, we first established datasets by retrospectively selecting 200 CMR cases. The models were trained using two different cohorts (passively and actively curated) and applied data augmentation to enhance training. Once trained, the models were validated on an external dataset, consisting of 20 cases acquired at another center. We then compared accuracy metrics and applied class activation mapping (CAM) to visualize DL model performance. Results: The DL and ML models trained with the passively-curated CMR cohort were 99.1% and 99.3% accurate on the validation set, respectively. However, when tested on the CMR cases with complex anatomy, both models performed poorly. After training and testing our models again on all 200 cases (active cohort), validation on the external dataset resulted in 95% and 90% accuracy, respectively. The CAM analysis depicted heat maps that demonstrated the importance of carefully curating the datasets to be used for training. Conclusions: Both DL and ML models can accurately classify CMR images, but DL outperformed ML when classifying images with complex heart anatomy.

artificial intelligence↗

Combustion machine learning: Principles, progress and prospects

Progress in combustion science and engineering has led to the generation of large amounts of data from large-scale simulations, high-resolution experiments, and sensors. This corpus of data offers enormous opportunities for extracting new knowledge and insights—if harnessed effectively. Machine learning (ML) techniques have demonstrated remarkable success in data analytics, thus offering a new paradigm for data-intense analyses and scientific investigations through combustion machine learning (CombML). While data-driven methods are utilized in various combustion areas, recent advances in algorithmic developments, the accessibility of open-source software libraries, the availability of computational resources, and the abundance of data have together rendered ML techniques ubiquitous in scientific analysis and engineering. This article examines ML techniques for applications in combustion science and engineering. Starting with a review of sources of data, data-driven techniques, and concepts, we examine supervised, unsupervised, and semi-supervised ML methods. Various combustion examples are considered to illustrate and to evaluate these methods. Next, we review past and recent applications of ML approaches to problems in combustion, spanning fundamental combustion investigations, propulsion and energy-conversion systems, and fire and explosion hazards. Challenges unique to CombML are discussed and further opportunities are identified, focusing on interpretability, uncertainty quantification, robustness, consistency, creation and curation of benchmark data, and the augmentation of ML methods with prior combustion-domain knowledge.

33 ADVANCED PROPULSION SYSTEMS↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗

Planet Microbe: a platform for marine microbiology to discover and analyze interconnected ‘omics and environmental data

In recent years, large-scale oceanic sequencing efforts have provided a deeper understanding of marine microbial communities and their dynamics. These research endeavors require the acquisition of complex and varied datasets through large, interdisciplinary and collaborative efforts. However, no unifying framework currently exists for the marine science community to integrate sequencing data with physical, geological, and geochemical datasets. Planet Microbe is a web-based platform that enables data discovery from curated historical and on-going oceanographic sequencing efforts. In Planet Microbe, each ‘omics sample is linked with other biological and physiochemical measurements collected for the same water samples or during the same sample collection event, to provide a broader environmental context. This work highlights the need for curated aggregation efforts that can enable new insights into high-quality metagenomic datasets.

59 BASIC BIOLOGICAL SCIENCES↗

Energy Community Atlas

The Energy Community Atlas provides efficient access to authoritative, curated, and relevant data that is vital to supporting energy planning, development, and economic growth across the U.S. In this effort, researchers at the National Energy Technology Laboratory (NETL) are utilizing advanced data visualization and transformation capabilities to develop an integrated, data atlas and resource focused on supporting energy community transitions to new manufacturing opportunities. Specifically, this project is working to find, acquire, integrate, and virtually host in a user-friendly, public and private solution from available resources, relevant to understanding and characterizing fossil energy communities themselves and inform energy planning, development, and economic growth opportunities, including opportunities for co-development to support manufacturing, critical materials, and more. This Atlas when complete is to offer a one-stop-shop for stakeholders to derive new insights to accelerate energy investments and strategic decision support needs. These are following datasets that are available as part of this ongoing project • Energy Community Atlas Map Package - This is ArcPro Map package and it contains all of the symbolized layers along with ArcPro map and geodatabase • Energy Community Atlas ArcGIS REST service - https://www.arcgis.com/apps/mapviewer/index.html?panel=gallery&suggestField=true&layers=537ced69bd88440380a62c2ec8aca30c • README Energy Community Atlas - Read me word document that has details about feature classes in Map package, ArcPro map and ArcGIS Rest Service

Bipartisan Infrastructure Law↗

Smart Methane Emission Detection System Development (Final Report)

Working with the Department of Energy's National Energy Technology Laboratory, Southwest Research Institute® (SwRI®) developed a system to identify methane leaks reliably, accurately, and autonomously at critical midstream sections of the natural gas distribution network in real-time for the purpose of mitigating methane emissions using Optical Gas Imaging (OGI) cameras. SwRI's Smart Leak Detection – Methane (SLED/M) adds a high degree of automation to the process of methane leak detection to minimize sources of human error, minimize response time to a leak event, and maximize midstream visibility. Furthermore, SwRI has been working towards integrating Quantitative OGI (QOGI) capabilities into this existing technology. By leveraging Deep Learning, SwRI now has the capability to estimate fugitive emission leak rates quickly and reliably, which allows operators to detect emissions, quantify leak rate, prioritize repairs, and validate the repairs in a single instrument. The next generation QOGI technology leverages the same cameras used in Leak Detection and Repair (LDAR) programs, with improvements in safety and speed for traditional quantification-based repairs, ultimately leading to less overhead cost for the operators. The goals for this research were to develop two types of models with the following goals: Run in real-time on the edge (≥ 12 Hz), Classification: Achieve less than 5% false positive detection, Classification: Achieve ≥ 95% methane plume detection rate, Regression: achieve ≤ 10 standard cubic feet per hour (scfh) prediction > 70% of the time. In order to achieve these results, multiple infrared (IR) and other sensors were investigated in tandem with the midwave IR (MWIR) OGI to provide additional information to train the underlying models. Information on atmospheric conditions including humidity, temperature, pressure, and solar radiation was provided by a weather station. Several machine learning and deep learning architectures and methods, including looking at quantized classification networks and regressions networks, were explored. As further data was collected, curated, and labeled, it allowed for more refined regressive networks to be adequately trained, leading to better insight into the true flow rates being observed. An important valuable deliverable of this research effort was the development of an advanced network which underwent multiple iterations capable of giving a continuous output. The current network has a predicted mean average percentage error (MAPE) of 12.3% just outside our target goal of 10.00%, but an accuracy of 97.78% at ±50 scfh, well within the overall goal for the Department of Energy (DOE) program. Upon closer inspection, it was observed that more than 10% of datapoints contributing to the MAPE predictions were the result of low flow rate predictions and are beyond the sensitivity of instrument measurement as a result of normal operational variation and noise.

03 NATURAL GAS↗

IDB Database Tables

The International Database of Reference Gamma-Ray Spectra of Various Nuclear Matter is designed to hold curated gamma spectral data is hosted by the International Atomic Energy Agency on its public facing web site. The database used to hold the spectral data was designed by Sandia National Labs under the auspices of the State Department’s Support Program. This document describes the tables and entity relationships that make up the database.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Summary of Responses to the Request for Information (RFI) on Partnerships for Transformational Artificial Intelligence Models

The Department of Energy (DOE) issued a Request for Information (RFI) in December 2025 inviting public comments regarding partnerships for transformational Artificial Intelligence (AI) models for the Genesis Mission Consortium, a public-private partnership platform. This RFI solicited feedback from industry, nonprofit organizations, universities, independent research organizations and other stakeholders. Specifically, the RFI asked three questions on (1) mobilizing DOE National Laboratories to curate the scientific data in a responsible and privacy-preserving manner, (2) the extent to which existing general-purpose AI models can be leveraged and which scientific disciplines are priorities for such model development, and (3) mechanisms by which these AI models can be provided to scientific communities. This document summarizes the input from 194 unique nonproprietary responses from businesses, universities, nonprofit organizations, research institutes and laboratories as well as a variety of other contributors, including individual contributions.

97 MATHEMATICS AND COMPUTING↗

Developing predictive models for µ opioid receptor binding using machine learning and deep learning techniques

Opioids exert their analgesic effect by binding to the µ opioid receptor (MOR), which initiates a downstream signaling pathway, eventually inhibiting pain transmission in the spinal cord. However, current opioids are addictive, often leading to overdose contributing to the opioid crisis in the United States. Therefore, understanding the structure-activity relationship between MOR and its ligands is essential for predicting MOR binding of chemicals, which could assist in the development of non-addictive or less-addictive opioid analgesics. This study aimed to develop machine learning and deep learning models for predicting MOR binding activity of chemicals. Chemicals with MOR binding activity data were first curated from public databases and the literature. Molecular descriptors of the curated chemicals were calculated using software Mold2. The chemicals were then split into training and external validation datasets. Random forest, k-nearest neighbors, support vector machine, multi-layer perceptron, and long short-term memory models were developed and evaluated using 5-fold cross-validations and external validations, resulting in Matthews correlation coefficients of 0.528–0.654 and 0.408, respectively. Furthermore, prediction confidence and applicability domain analyses highlighted their importance to the models’ applicability. Our results suggest that the developed models could be useful for identifying MOR binders, potentially aiding in the development of non-addictive or less-addictive drugs targeting MOR.

Research & Experimental Medicine↗