Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Document-Based Nuclear Archaeology

Deeper reductions in the nuclear arsenals will require better understanding of historic fissile material management and production. The concept of “nuclear archaeology” has been considered since the 1990s to provide the tools and methods to develop independent production estimates, primarily based on nuclear forensic techniques. Here, we propose to add a framework for reconstructing the history of a nuclear program that complements traditional nuclear archaeology techniques by examining the role of operating records to support such an effort. As a test case, we use the JEEP II reactor, a 2 MW civilian research reactor at Norway’s Institute for Energy Technology (IFE), in operation for more than fifty years, however, recently shut down permanently. We have collected, analyzed, and started to preserve the reactor’s operating records, which exist on both analog and digital media, and to simulate parts of its history using OpenMC/ONIX neutronics calculations. Here, a particular focus of this project has been on digital data curation and preservation to confirm and maintain the integrity, authenticity, and provenance of these records. In developing guidelines for best practices that conform to existing standards for long-term digital preservation and curation, we hope this project can help lay the basis for future nuclear archaeology efforts to support nuclear arms control and disarmament.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Why is the winner the best?

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multi- center study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), im- age preprocessing (97%), data curation (79%), and post- processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work.

Eisenmann, Matthias↗

A Comprehensive Calibration Framework for the Northwest River Forecast Center

We present a comprehensive framework developed by the Northwest River Forecast Center for calibrating hydrologically diverse basins. The framework includes models for snow, soil moisture, routing, channel loss, and consumptive use. Data inputs include a wide range of open-access datasets for meteorology, land use, topography, and land cover. The framework uses conceptual hydrologic models to handle basins with various hydrologic regimes including rain-driven and snowmelt-dominated basins. We also develop a flexible automatic calibration system that can handle numerous unobservable model parameters in a computationally efficient manner. A single-basin automatic calibration run can typically be completed on a modern laptop in under 10 min. We found that model performance metrics for this new approach match the quality of the NWRFC's previous labor-intensive manual calibrations. The model performance also rivals that of a state-of-the-art deep learning model at a fraction of the computational cost. This framework presents a new standard for the quality of calibrations possible with lumped conceptual hydrologic models, combining careful data curation, an objective calibration framework, and expert local knowledge. In addition, we have made software packages available for the entire suite of National Weather Service River Forecast System models, including SAC-SMA, SNOW-17, and Lag-K. These modern interfaces are intended to increase accessibility and facilitate future research.

Forecasting↗

Li-Ion Battery Electrode Contact Resistance Estimation by Mechanical Peel Test

Li-ion battery electrode electronic properties, including bulk conductivity and contact resistance, are critical parameters affecting cell performance and fast-charge capability. Contact resistance between the coating and current collector is often the largest electronic resistance in an electrode and is affected by chemical, microstructural, and interfacial variations. Direct measurements of contact resistance and bulk conductivity have proven to be challenging. In their absence, a mechanical electrode peel test is often used to compare adhesion and electrical contact resistance. However, using a micro-flexible-surface probe, contact resistance can be directly determined. Here, this work compares contact resistance and mechanical peel strength of multiple commercial-grade HE5050 and NCM523 cathodes and graphite and silicon anodes. It was found that peel strength correlates well with contact resistance in a carefully curated data set (p < 0.05) and in some situations may be a good metric to estimate electrical properties. However, there were distinct outliers in the data set, indicating that peel strength may not accurately reflect electrical properties when there is significant variation in electrode composition. These results illustrate the value of the micro-flexible-surface probe in quantifying contact resistance and bulk conductivity to better understand how battery composition and processing steps affect microstructure and resulting cell performance.

25 ENERGY STORAGE↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

America Resilient Climate Conference

On April 14, 2021, scientists, policymakers, and other interested parties from research institutes, academia, and other organizations gathered together virtually at the America Resilient Climate Conference to discuss one of the most pressing challenges of the 21st century: building resilience to climate change. Climate change affects the security and health of all Americans. Coastal areas are enduring more frequent and severe flooding due to sea level rise and storm surge; western states and Alaska have experienced increasingly devastating wildfires, driven in part by hotter, drier, and longer fire seasons; and communities across the nation have suffered through extreme precipitation events and heat waves. Even if emissions are reduced aggressively in the near future, the world—and the United States—will continue to feel the impacts of climate change for decades to come, due to the continued accumulation of greenhouse gasses in the atmosphere. Consequently, it is essential to act now to protect natural and human assets from the gradual—as well as extreme—impacts of a changing climate. To build resilient communities, leaders and community members need science-based information about the potential impacts climate change will have decades into the future and for specific regions. Therefore, it is essential to develop high-resolution climate models that can project both various climate impacts and the interactions between earth system variables and humans down to regional and local scales. Collecting and curating data for such models and their computational requirements poses large challenges. In the future, artificial intelligence will be needed to increase their accuracy and reduce associated uncertainties. The investments in Earth system science and artificial intelligence made by the U.S. Department of Energy and other federal entities will be essential in addressing these challenges.

54 ENVIRONMENTAL SCIENCES↗

CTSA MHRI Datasets

Oak Ridge National Laboratory (ORNL) has collaborated with MedStar Health Research Institute (MHRI) to develop, test, and validate health outcomes using electronic health records (EHRs) from hospitals associated with participating Clinical and Translational Science Awards (CTSA). MHRI, the research organization of MedStar Health (MSH), has a history of initiating projects, both in the laboratory and in the field, that serve the needs of medically underserved and disenfranchised groups. In this document, the ORNL team is providing data curation documentation for the following datasets, which are publicly available: 1. Area Deprivation Index (ADI) 2015 block group; 2. Child Opportunity Index 2015 tract; 3. Low food access 2017 block group; 4. Neighborhood deprivation index 2017 tract; 5. Social Capital Index 2014 county; and, 6. Social Vulnerability Index 2014 tract.

54 ENVIRONMENTAL SCIENCES↗

Validation of LOCA2 and STAR-ESDM Statistically Downscaled Products

The National Climate Assessment (NCA) is the preeminent national report examining current and future risks posed by climate change. Countless agencies, policymakers, stakeholders and other end-users rely upon guidance from the NCA to plan for an uncertain future. These groups all depend on modern curated data, provided alongside the NCA, to quantify the impact of climate change on metrics of relevance for their decision processes. In its fifth iteration (NCA5), two statistically downscaled ensemble products, each providing data at grid spacing of approximately 5km over the contiguous United States, were selected to accompany the report. These include LOCalized Analogs version 2 (LOCA2) and Seasonal Trends and Analysis of Residuals Empirical-Statistical Downscaling Model (STAR-ESDM). Both data products are produced through a process known as statistical downscaling, where relatively coarse Global Climate Model (GCM) data is refined to locally relevant scales through the application of scientifically-supported empirical and algorithmic relationships. In support of the NCA effort, this report provides an independent validation of these two products against historical observations, with a focus on precipitation and near-surface temperature variables. Based on the results of this validation, several recommendations are provided related to the use of these data products. The structure of this report is as follows: In section 2, we review three gridded observational products that are used as part of our intercomparison. In section 3, we describe the two statistical downscaling techniques and their corresponding datasets that are the focus of this study. In section 4, the methodology we employ for validation is described. Section 5 provides results of the validation, which in turn motivate our recommendations on the use of these data products. A brief summary is provided in section 6.

54 ENVIRONMENTAL SCIENCES↗

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 1

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024). Three distinct rounds of FSP experiments were performed by the experimental team, producing replicate samples utilizing across different nominal processing conditions (Condition IDs) listed in Table 1. The starting material on which FSP was applied was commercially available unprocessed stainless-steel type 316L material. Chosen processing conditions were very diverse, and some were intentionally chosen to produce defects. Several samples experienced tool breakage during experimentation, so a full set of three replicates was not produced for every nominal processing condition.

36 MATERIALS SCIENCE↗

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 2

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024).

36 MATERIALS SCIENCE↗

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 3

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024).

36 MATERIALS SCIENCE↗

Materials Characterization, Prediction, and Control Project: Summary Report on Material Characterization, Part 4

The Pacific Northwest National Laboratory (PNNL) undertook the Materials Characterization, Prediction, and Control (MCPC) Laboratory Directed Research and Development Project to advance understanding of nuclear material processing and enable multifold acceleration in the development and qualification of new material systems in national security and advanced energy applications (Smith 2021). The MCPC Project executed research across three scientific vertices—material characterization, predictive modeling, and data analytics—with extensive support by a data curation and management team. The central technical objective in the MCPC Project was to improve the prediction and characterization of the process-structure-property relationships within the microstructurally refined region of stainless-steel samples prepared utilizing friction stir processing (FSP). Application of the FSP technique is well established at PNNL within the Solid Phase Processing capability through many years of investment across a range of materials and applications (PNNL 2024).

36 MATERIALS SCIENCE↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

Automating methods for estimating metabolite volatility

The volatility of metabolites can influence their biological roles and inform optimal methods for their detection. Yet, volatility information is not readily available for the large number of described metabolites, limiting the exploration of volatility as a fundamental trait of metabolites. Here, we adapted methods to estimate vapor pressure from the functional group composition of individual molecules (SIMPOL.1) to predict the gas-phase partitioning of compounds in different environments. We implemented these methods in a new open pipeline called volcalc that uses chemoinformatic tools to automate these volatility estimates for all metabolites in an extensive and continuously updated pathway database: the Kyoto Encyclopedia of Genes and Genomes (KEGG) that connects metabolites, organisms, and reactions. We first benchmark the automated pipeline against a manually curated data set and show that the same category of volatility (e.g., nonvolatile, low, moderate, high) is predicted for 93% of compounds. We then demonstrate how volcalc might be used to generate and test hypotheses about the role of volatility in biological systems and organisms. Specifically, we estimate that 3.4 and 26.6% of compounds in KEGG have high volatility depending on the environment (soil vs. clean atmosphere, respectively) and that a core set of volatiles is shared among all domains of life (30%) with the largest proportion of kingdom-specific volatiles identified in bacteria. With volcalc , we lay a foundation for uncovering the role of the volatilome using an approach that is easily integrated with other bioinformatic pipelines and can be continually refined to consider additional dimensions to volatility. The volcalc package is an accessible tool to help design and test hypotheses on volatile metabolites and their unique roles in biological systems.

59 BASIC BIOLOGICAL SCIENCES↗

A multi-scale time-series dataset of anthropogenic heat from buildings in Los Angeles County

The dataset contains hourly Anthropogenic heat (AH) from buildings in Los Angeles County, based on weather data from 2018. The hourly AH is aggregated at three spatial resolutions: 450m x 450m grid, 12km x 12km grid, and census tract. The AH is broken down into three components: building envelope surface convection, heating, ventilation, and air conditioning (HVAC) system heat release, and zone exfiltration and exhaust air heat loss. The dataset is created with the physics-based EnergyPlus building energy models to calculate individual buildings' AH considering WRF-UCM simulated microclimate conditions. Please refer to the paper "A multi-scale time-series dataset of anthropogenic heat from buildings in Los Angeles County" for more information about the data generation workflow and the data validation procedure. The data set contains two folders: the "output_data" folder holds the simulation results (EP_output and EP_output_csv), building metadata (building_metadata.geojson and building_metadata.csv), aggregated heat emission and energy consumption time-series data (hourly_heat_energy), and geographical data (geo_data) associated with the GEOID referenced in heat and energy consumption data. The "input_data" folder contains the raw data used to generate files in the "output_data" folder as well as data sets used in the validation. The code repository (https://github.com/IMMM-SFA/xu_etal_2022_sdata) holds the processing scripts for data curation, validation, and visualization.

Energy↗

Lunar Glovebox Balance with Wireless Technology

The most important equipment required for processing lunar samples is a high-quality mass balance for maintaining accurate weight inventory, security, and scientific study. After careful review, a Curation Office memo by Michael Duke in 1978 chose the Mettler PL200 to be used for sample weight measurements inside the gloveboxes (Fig. 3). These commercial off-the-shelf (COTS) balances did not meet the strict accepted material requirements in the Lunar lab. As a result, each balance housing, weighing pan, and wiring was custom retrofitted to meet Lunar Operating Procedure (LOP) 54 requirements [for material construction restrictions]. The original design drawings for the custom housings, readout support stands, and wiring were done by the JSC engineering directorate. The 1977- 1978 schematics, drawings, and files are now housed in the curation Data Center. Per the design specifications, the housing was fabricated from aluminum grade 6061 T6, seamless welds, and anodized per MIL-A-8625 type I, class I. The balance feet were TFE Teflon and any required joints were sealed with Viton A gaskets. The readout display and support stands outside the glovebox were fabricated from 300 series stainless steel with #4 finish and mounted to the glovebox with welded bolts. Wire harnesses that linked the balance with the outside display and power were encapsulated with TFE Teflon and transported through custom Deutsch wire bulk head pass-through systems from inside to outside the glovebox. These Deutsch connectors were custom fabricated with 316L stainless steel bodies, Viton A O-rings, aluminum 6061 with electroless nickel plating, Teflon (replacing the silicone), and gold crimp connectors (no soldering). Many of the Deutsch connectors may have been used in the Apollo program high vacuum complex in building 37 and date to about 1968 to 1970.

Zeigler, Ryan A.↗

Extravehicular Activity Mission System Software (EMSS) - Enabling Human Planetary Exploration Data Within The Broader Planetary Data Ecosystem

The planetary science community is once again on the verge of generating, capturing and analyzing human planetary exploration data, this time via the Artemis program. Artemis missions will involve robotic missions in addition to human extravehicular activity (EVA) where crew will be generating scientific data [1]. Present-day robotic mission data expectations for data archiving involves ingesting data into the Planetary Data System (PDS), but how might PDS be leveraged/adapted/ready (or not) for human spaceflight mission data, particularly EVA data that includes non-scientific data that provides important context to the scientific data gathered on the lunar surface? This question has broader implications than what this abstract can answer, but we wanted to pose the question to 1) get conversations started and 2) highlight how operations software data handling could play a role in overall data curation.

M J Miller↗

NASA Life Sciences Portal (NLSP): Supporting Scientific Transparency and Reproducibility

NASA’s Life Sciences Ports (NLSP) serves the scientific community by providing curated data from space life science experiment. The Human Research Program (HRP) with the help of NLSP is currently transforming their life sciences data archive systems and processes to improve compliance with the FAIR principles [1]. Some of these improvements will at the same time support the twin pillars of Open Science [2]: transparency of methods and reproducibility of results. Scientific transparency is marked by the easily intelligible communication of what has been investigated: what were the procedures for collecting sample and the characteristics of samples collected? what kinds of measurements were made, what were the environmental conditions of the measurements? What were the analysis techniques of the collected data? Reproducibility of the results and findings from the investigation requires a high level of transparency for all but the simplest investigations; the slightest deviation in communicating and replicating complex experimental procedures or data analyses can often yield quite different data and even findings, thwarting their validation. One of the ways the NLSP is aiming to improve the communication of scientific information is through the use of ontology-driven metadata. Ontologies are powerful, graph-based knowledge representation structures, which can be leveraged to increase data interoperability, the area of the FAIR principles in which many data systems most lack compliance. Over the past decade, there has been a concerted effort in the biomedical community to develop modular and narrowly focused domain and application-specific ontologies in a common, open-source framework, the Open Biological and Biomedical Ontology (OBO) Foundry [3]. The open sharing and modular nature of this effort promises huge increases in harmonized data sharing for systems that leverage these models. Which is in line with the FAIR Data Principles of Findability, Accessibility, Interoperability, and Reuse for scientific data management and stewardship. 1. Wilkinson, M.D., et al., The FAIR Guiding Principles for scientific data management and stewardship. Sci Data, 2016. 3: p. 160018. 2. National Academies of Sciences, E. and Medicine, Open Science by Design: Realizing a Vision for 21st Century Research. 2018, Washington, DC: The National Academies Press. 232. 3. Smith, B., et al., The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol, 2007. 25(11): p. 1251-5.

Life Sciences data↗