Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Evaluating the factors influencing accuracy, interpretability, and reproducibility in the use of machine learning classifiers in biology to enable standardization

The complexity and variability of biological data has promoted the increased use of machine learning methods to understand processes and predict outcomes. These same features complicate reliable, reproducible, interpretable, and responsible use of such methods, resulting in questionable relevance of the derived. outcomes. Here we systematically explore challenges associated with applying machine learning to predict and understand biological processes using a well- characterized in vitro experimental system. We evaluated factors that vary while applying machine learning classifers: (1) type of biochemical signature (transcripts vs. proteins), (2) data curation methods (pre- and post-processing), and (3) choice of machine learning classifier. Using accuracy, generalizability, interpretability, and reproducibility as metrics, we found that the above factors significantly mod- ulate outcomes even within a simple model system. Our results caution against the unregulated use of machine learning methods in the biological sciences, and strongly advocate the need for data standards and validation tool-kits for such studies.

59 BASIC BIOLOGICAL SCIENCES↗

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database↗

Towards Diverse and Representative Global Pretraining Datasets for Remote Sensing Foundation Models

The design of a pretraining dataset is emerging as a critical component for the generality of foundation models. In the remote sensing realm, large volumes of imagery and benchmark datasets exist that can be leveraged to pretrain foundation models, however using this imagery in absence of a well-crafted sampling strategy is inefficient and has the potential to create biased and less generalizable models. Here, we provide a discussion and vision for the curation and assessment of pretraining datasets for remote sensing geospatial foundation models. We highlight the importance of geographic, temporal, and image acquisition diversity and review possible strategies to enable such diversity at global scale. In addition to these characteristics, support for various spatial-temporal pretext tasks within the dataset is also critical. Ultimately, our primary objective is to place emphasis on and draw attention to the data curation stage of the foundation model development pipeline. By doing so, we think it is possible to reduce biases of geospatial foundation models, as well as enable broader generalization to downstream remote sensing tasks and applications.

Arndt, Jacob↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

High-throughput predictions of metal–organic framework electronic properties: theoretical challenges, graph neural networks, and data exploration

Abstract With the goal of accelerating the design and discovery of metal–organic frameworks (MOFs) for electronic, optoelectronic, and energy storage applications, we present a dataset of predicted electronic structure properties for thousands of MOFs carried out using multiple density functional approximations. Compared to more accurate hybrid functionals, we find that the widely used PBE generalized gradient approximation (GGA) functional severely underpredicts MOF band gaps in a largely systematic manner for semi-conductors and insulators without magnetic character. However, an even larger and less predictable disparity in the band gap prediction is present for MOFs with open-shell 3 d transition metal cations. With regards to partial atomic charges, we find that different density functional approximations predict similar charges overall, although hybrid functionals tend to shift electron density away from the metal centers and onto the ligand environments compared to the GGA point of reference. Much more significant differences in partial atomic charges are observed when comparing different charge partitioning schemes. We conclude by using the dataset of computed MOF properties to train machine-learning models that can rapidly predict MOF band gaps for all four density functional approximations considered in this work, paving the way for future high-throughput screening studies. To encourage exploration and reuse of the theoretical calculations presented in this work, the curated data is made publicly available via an interactive and user-friendly web application on the Materials Project.

36 MATERIALS SCIENCE↗

Agnostic capture of pathogens for the detection and diagnostics of emerging threats

The continued emergence of pathogens, whether novel, re-emerging, or engineered, poses a persistent global biosecurity and public health challenge. Recent outbreaks, including COVID-19, Lassa fever, Marburg virus, mpox, and avian influenza, underscore the urgent need for robust systems that enable rapid surveillance, early diagnosis, and timely countermeasures before widespread human transmission occurs. In this article, we focus on early detection technologies and systematically evaluate current diagnostic and sensing modalities. We highlight sequencing and spectroscopy as two complementary approaches capable of providing broad, agnostic detection and rich biological insight. Our analysis emphasizes that scientific innovation alone is insufficient: effective preparedness also requires improved data curation, integration, and sharing to build AI-ready resources that accelerate future responses. We argue for coordinated advances in both technological capabilities and supporting infrastructure to enable the rapid identification and characterization of emerging pathogens and to fully leverage modern science against evolving infectious threats.

Environmental health↗

Incorporating Diurnal and Meter-Scale Variations of Ambient CO 2 Concentrations in Development of Direct Air Capture Technologies

To be implemented on climate-relevant scales, direct air capture of CO 2 (DAC) will require large capital-intensive facilities and careful attention to cost minimization. In making decisions among potential sites for DAC facilities, all of the factors that will impact process cost and efficiency should be considered. In this paper we focus on a factor that has previously received little attention in the DAC community, namely variations in atmospheric conditions on hourly time scales and length scales of meters. We present data curated from extensive previous studies of biosphere-atmosphere fluxes with observations of CO 2 concentration, temperature, and relative humidity (RH) with hourly resolution from many sites in North America. These include locations where typical diurnal variations in CO 2 concentration during summer months exceeds 150 ppm. These variations are larger than the seasonal variations that exist between averaged CO 2 concentrations in winter and summer, and they are highly correlated with diurnal variations in temperature and RH. Diurnal variations are dependent on the height above ground at which CO 2 concentrations are measured, with smaller variations existing at heights of 10 m or more than at ground level. We illustrate the potential implications of these short-term variations for the operation and optimization of a DAC process with process-level calculations for a specific adsorption-based process using amine-rich adsorbents.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advancing the Prediction of MS/MS Spectra Using Machine Learning

Tandem mass spectrometry (MS/MS) is an important tool for the identification of small molecules and metabolites where resultant spectra are most commonly identified by matching them with spectra in MS/MS reference libraries. While popular, this strategy is limited by the contents of existing reference libraries. In response to this limitation, various methods are being developed for the in silico generation of spectra to augment existing libraries. Recently, machine learning and deep learning techniques have been applied to predict spectra with greater speed and accuracy. Here, in this work, we investigate the challenges these algorithms face in achieving fast and accurate predictions on a wide range of small molecules. The challenges are often amplified by the use of generic machine learning benchmarking tactics, which lead to misleading accuracy scores. Curating data sets, only predicting spectra for sufficiently high collision energies, and working more closely with experimental mass spectrometrists are recommended strategies to improve overall prediction accuracy in this nuanced field.

47 OTHER INSTRUMENTATION↗

A Climatology and Life‐Cycle Characteristics of Atmospheric Fronts and Their Associated Precipitation

Abstract Atmospheric fronts are one of the main sources of mid‐latitude variability. We employ a novel method for identifying and tracking fronts and frontal precipitation. Thermal and dynamical variables are used to identify fronts as areal objects in space, which are tracked in time using the open‐source TempestExtremes software package. Precipitation objects are co‐located to identify frontal precipitation. The method is subjected to validation and sensitivity tests using manually curated data from the National Weather Service. Climatologies of fronts and frontal precipitation are computed from reanalysis and observations; fronts are present upwards of 14% of the time in the storm tracks, and represent the majority (up to 90%) of total and extreme precipitation. Novel aspects of the method are showcased through the lifetime characteristics of fronts across North America. Three sets of warm and cold fronts were discovered, and their duration, distance‐traveled, and translation velocity are examined. Plain Language Summary Mid‐latitude low‐pressure systems and weather fronts are important for our day‐to‐day experience of weather events, particularly in the mid‐latitudes. This work makes use of standardized atmospheric data and creates a method of automatically tracking these important atmospheric features and their precipitation to quantify their relative role in global precipitation. Weather fronts are persistent in the mid‐latitudes and are associated with the majority of precipitation–particularly the most intense precipitation. Trajectories of fronts over North America are categorized to create a set of archetypal fronts that occur in that region. The differences between these types of fronts are characterized. Key Points An automated, efficient, and skillful frontal detection algorithm is developed and validated Fronts contribute a larger fraction of extreme precipitation than all precipitation in mid‐latitude storm tracks Fronts across North America have substantial variation in characteristics depending on their origin location

extratropical cyclone↗

Document-Based Nuclear Archaeology

Deeper reductions in the nuclear arsenals will require better understanding of historic fissile material management and production. The concept of “nuclear archaeology” has been considered since the 1990s to provide the tools and methods to develop independent production estimates, primarily based on nuclear forensic techniques. Here, we propose to add a framework for reconstructing the history of a nuclear program that complements traditional nuclear archaeology techniques by examining the role of operating records to support such an effort. As a test case, we use the JEEP II reactor, a 2 MW civilian research reactor at Norway’s Institute for Energy Technology (IFE), in operation for more than fifty years, however, recently shut down permanently. We have collected, analyzed, and started to preserve the reactor’s operating records, which exist on both analog and digital media, and to simulate parts of its history using OpenMC/ONIX neutronics calculations. Here, a particular focus of this project has been on digital data curation and preservation to confirm and maintain the integrity, authenticity, and provenance of these records. In developing guidelines for best practices that conform to existing standards for long-term digital preservation and curation, we hope this project can help lay the basis for future nuclear archaeology efforts to support nuclear arms control and disarmament.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Why is the winner the best?

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multi- center study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), im- age preprocessing (97%), data curation (79%), and post- processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work.

Eisenmann, Matthias↗

A Comprehensive Calibration Framework for the Northwest River Forecast Center

We present a comprehensive framework developed by the Northwest River Forecast Center for calibrating hydrologically diverse basins. The framework includes models for snow, soil moisture, routing, channel loss, and consumptive use. Data inputs include a wide range of open-access datasets for meteorology, land use, topography, and land cover. The framework uses conceptual hydrologic models to handle basins with various hydrologic regimes including rain-driven and snowmelt-dominated basins. We also develop a flexible automatic calibration system that can handle numerous unobservable model parameters in a computationally efficient manner. A single-basin automatic calibration run can typically be completed on a modern laptop in under 10 min. We found that model performance metrics for this new approach match the quality of the NWRFC's previous labor-intensive manual calibrations. The model performance also rivals that of a state-of-the-art deep learning model at a fraction of the computational cost. This framework presents a new standard for the quality of calibrations possible with lumped conceptual hydrologic models, combining careful data curation, an objective calibration framework, and expert local knowledge. In addition, we have made software packages available for the entire suite of National Weather Service River Forecast System models, including SAC-SMA, SNOW-17, and Lag-K. These modern interfaces are intended to increase accessibility and facilitate future research.

Forecasting↗

Li-Ion Battery Electrode Contact Resistance Estimation by Mechanical Peel Test

Li-ion battery electrode electronic properties, including bulk conductivity and contact resistance, are critical parameters affecting cell performance and fast-charge capability. Contact resistance between the coating and current collector is often the largest electronic resistance in an electrode and is affected by chemical, microstructural, and interfacial variations. Direct measurements of contact resistance and bulk conductivity have proven to be challenging. In their absence, a mechanical electrode peel test is often used to compare adhesion and electrical contact resistance. However, using a micro-flexible-surface probe, contact resistance can be directly determined. Here, this work compares contact resistance and mechanical peel strength of multiple commercial-grade HE5050 and NCM523 cathodes and graphite and silicon anodes. It was found that peel strength correlates well with contact resistance in a carefully curated data set (p < 0.05) and in some situations may be a good metric to estimate electrical properties. However, there were distinct outliers in the data set, indicating that peel strength may not accurately reflect electrical properties when there is significant variation in electrode composition. These results illustrate the value of the micro-flexible-surface probe in quantifying contact resistance and bulk conductivity to better understand how battery composition and processing steps affect microstructure and resulting cell performance.

25 ENERGY STORAGE↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

America Resilient Climate Conference

On April 14, 2021, scientists, policymakers, and other interested parties from research institutes, academia, and other organizations gathered together virtually at the America Resilient Climate Conference to discuss one of the most pressing challenges of the 21st century: building resilience to climate change. Climate change affects the security and health of all Americans. Coastal areas are enduring more frequent and severe flooding due to sea level rise and storm surge; western states and Alaska have experienced increasingly devastating wildfires, driven in part by hotter, drier, and longer fire seasons; and communities across the nation have suffered through extreme precipitation events and heat waves. Even if emissions are reduced aggressively in the near future, the world—and the United States—will continue to feel the impacts of climate change for decades to come, due to the continued accumulation of greenhouse gasses in the atmosphere. Consequently, it is essential to act now to protect natural and human assets from the gradual—as well as extreme—impacts of a changing climate. To build resilient communities, leaders and community members need science-based information about the potential impacts climate change will have decades into the future and for specific regions. Therefore, it is essential to develop high-resolution climate models that can project both various climate impacts and the interactions between earth system variables and humans down to regional and local scales. Collecting and curating data for such models and their computational requirements poses large challenges. In the future, artificial intelligence will be needed to increase their accuracy and reduce associated uncertainties. The investments in Earth system science and artificial intelligence made by the U.S. Department of Energy and other federal entities will be essential in addressing these challenges.

54 ENVIRONMENTAL SCIENCES↗

CTSA MHRI Datasets

Oak Ridge National Laboratory (ORNL) has collaborated with MedStar Health Research Institute (MHRI) to develop, test, and validate health outcomes using electronic health records (EHRs) from hospitals associated with participating Clinical and Translational Science Awards (CTSA). MHRI, the research organization of MedStar Health (MSH), has a history of initiating projects, both in the laboratory and in the field, that serve the needs of medically underserved and disenfranchised groups. In this document, the ORNL team is providing data curation documentation for the following datasets, which are publicly available: 1. Area Deprivation Index (ADI) 2015 block group; 2. Child Opportunity Index 2015 tract; 3. Low food access 2017 block group; 4. Neighborhood deprivation index 2017 tract; 5. Social Capital Index 2014 county; and, 6. Social Vulnerability Index 2014 tract.

54 ENVIRONMENTAL SCIENCES↗

Validation of LOCA2 and STAR-ESDM Statistically Downscaled Products

The National Climate Assessment (NCA) is the preeminent national report examining current and future risks posed by climate change. Countless agencies, policymakers, stakeholders and other end-users rely upon guidance from the NCA to plan for an uncertain future. These groups all depend on modern curated data, provided alongside the NCA, to quantify the impact of climate change on metrics of relevance for their decision processes. In its fifth iteration (NCA5), two statistically downscaled ensemble products, each providing data at grid spacing of approximately 5km over the contiguous United States, were selected to accompany the report. These include LOCalized Analogs version 2 (LOCA2) and Seasonal Trends and Analysis of Residuals Empirical-Statistical Downscaling Model (STAR-ESDM). Both data products are produced through a process known as statistical downscaling, where relatively coarse Global Climate Model (GCM) data is refined to locally relevant scales through the application of scientifically-supported empirical and algorithmic relationships. In support of the NCA effort, this report provides an independent validation of these two products against historical observations, with a focus on precipitation and near-surface temperature variables. Based on the results of this validation, several recommendations are provided related to the use of these data products. The structure of this report is as follows: In section 2, we review three gridded observational products that are used as part of our intercomparison. In section 3, we describe the two statistical downscaling techniques and their corresponding datasets that are the focus of this study. In section 4, the methodology we employ for validation is described. Section 5 provides results of the validation, which in turn motivate our recommendations on the use of these data products. A brief summary is provided in section 6.

54 ENVIRONMENTAL SCIENCES↗

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 1

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024). Three distinct rounds of FSP experiments were performed by the experimental team, producing replicate samples utilizing across different nominal processing conditions (Condition IDs) listed in Table 1. The starting material on which FSP was applied was commercially available unprocessed stainless-steel type 316L material. Chosen processing conditions were very diverse, and some were intentionally chosen to produce defects. Several samples experienced tool breakage during experimentation, so a full set of three replicates was not produced for every nominal processing condition.

36 MATERIALS SCIENCE↗