Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram↗

Seascape Interface Control Document (V.1)

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source codes, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Seascape Interface Control Document (V. 2)

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source codes, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Seascape Interface Control Document

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source software, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Workshop Summary: Bridging the Gap Between Atmospheric Science and Grid Integration

The need for dedicated, accurate, expertly curated weather data is increasingly important as the share of variable renewable energy increases on the power system. Projections for futures with very high (50+% annual energy) shares of variable generation require ongoing assessment of data requirements from industry stakeholders in their power system operation and planning contexts. In March 2024, NREL organized a workshop entitled "Bridging the Gap Between Atmospheric Science and Grid Integration Workshop", which brought atmospheric scientists and power system experts together to refine the requirements of atmospheric datasets for grid integration, and to describe a holistic approach to creating new and regularly updated national scale wind datasets for power system planning and operations. The results of this workshop are being used to inform the near-term development and a longer-term strategy for DOE to produce relevant wind resource datasets and inform wider use of wind/solar/load data sets in power system planning. This presentation provides an overview of a preworkshop survey, an assessment of current state of the art of national-scale datasets for wind resource assessment and grid integration, insights on appropriate uses of the WTK-LED, power system perspectives on data needs, as well as recommended next steps as discussed in the workshop and how these steps support longer-term strategies.

17 WIND ENERGY↗

The Colorado East River Community Observatory Data Collection

Abstract The U.S. Department of Energy's (DOE) Colorado East River Community Observatory (ER) in the Upper Colorado River Basin was established in 2015 as a representative mountainous, snow‐dominated watershed to study hydrobiogeochemical responses to hydrological perturbations in headwater systems. The ER is characterized by steep elevation, geologic, hydrologic and vegetation gradients along floodplain, montane, subalpine, and alpine life zones, which makes it an ideal location for researchers to understand how different mountain subsystems contribute to overall watershed behaviour. The ER has both long‐term and spatially‐extensive observations and experimental campaigns carried out by the Watershed Function Scientific Focus Area (SFA), led by Lawrence Berkeley National Laboratory, and researchers from over 30 organizations who conduct cross‐disciplinary process‐based investigations and modelling of watershed behaviour. The heterogeneous data generated at the ER include hydrological, genomic, biogeochemical, climate, vegetation, geological, and remote sensing data, which combined with model inputs and outputs comprise a collection of datasets and value‐added products within a mountainous watershed that span multiple spatiotemporal scales, compartments, and life zones. Within 5 years of collection, these datasets have revealed insights into numerous aspects of watershed function such as factors influencing snow accumulation and melt timing, water balance partitioning, and impacts of floodplain biogeochemistry and hillslope ecohydrology on riverine geochemical exports. Data generated by the SFA are managed and curated through its Data Management Framework. The SFA has an open data policy, and over 70 ER datasets are publicly available through relevant data repositories. A public interactive map of data collection sites run by the SFA is available to inform the broader community about SFA field activities. Here, we describe the ER and the SFA measurement network, present the public data collection generated by the SFA and partner institutions, and highlight the value of collecting multidisciplinary multiscale measurements in representative catchment observatories.

54 ENVIRONMENTAL SCIENCES↗

A Proposed Geospatial Data Preservation Strategy for DOE's Office of Legacy Management

Identifying and planning preservation and curation activities associated with geospatial data will improve the ability of the U.S. Department of Energy Office of Legacy Management (LM) to support their core mission of protecting human health and the environment. This report documents the development LM's strategy for preserving and curating geospatial data within the context of LM's data-lifecycle-management framework. The strategy consists of preservation and curation elements, specific activities, and key enabling factors that ensure LM's geospatial data is maintained. Preservation elements enable the effective preservation of LM's geospatial data and recognizes that strategies need to be flexible to adapt to ongoing changes in scale, technology, and standards. Key enabling factors are intended to highlight critical data management responsibilities that must be addressed by LM to meet its preservation and curation objectives. A summary of best practices for geospatial data preservation is provided as part of the strategy.

97 MATHEMATICS AND COMPUTING↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

A Reproducible Validation of Algorithms for Estimating Array Tilt and Azimuth from Photovoltaic Power Time Series

In this research, we assess the viability of four different, publicly available algorithms for estimating the azimuth and tilt parameters of solar photovoltaic systems using only the associated AC power time series data and site latitude-longitude coordinates. In this work, we curated a benchmarking data set of 44 fixed-tilt systems, comprising 275 measured AC power inverter data streams, with known azimuth and tilt parameters. Additionally, we isolated test cases in the data set with real- world issues, including shading and clipping, to determine how algorithm performance varies based on the presence of these phenomena. Using this data set for benchmarking, we evaluated the estimated vs. actual system characteristics for each algorithm, as well as the associated algorithm execution time using a standardized benchmarking process. The two highest performing algorithms were the Solar Data Tools and the PVWatts 5- based methods, which both achieved a median absolute error of approximately 5 and 1 degrees for azimuth and tilt, respectively. During run time analysis, the SDT method was approximately 5 times faster than the PVWatts 5-based method, with the median execution time for a stream varying between 6 and 8 seconds vs. a median run time of 31 seconds for the PVWatts 5-based method.

azimuth↗

A Reproducible Validation of Algorithms for Estimating Array Tilt and Azimuth from Photovoltaic Power Time Series

In this research, we assess the viability of four different, publicly available algorithms for estimating the azimuth and tilt parameters of solar photovoltaic systems using only the associated AC power time series data and site latitude-longitude coordinates. In this work, we curated a benchmarking data set of 44 fixed-tilt systems, comprising 275 measured AC power inverter data streams, with known azimuth and tilt parameters. Additionally, we isolated test cases in the data set with real-world issues, including shading and clipping, to determine how algorithm performance varies based on the presence of these phenomena. Using this data set for benchmarking, we evaluated the estimated vs. actual system characteristics for each algorithm, as well as the associated algorithm execution time using a standardized benchmarking process. The two highest performing algorithms were the Solar Data Tools and the PVWatts 5-based methods, which both achieved a median absolute error of approximately 5 and 1 degrees for azimuth and tilt, respectively. During run time analysis, the SDT method was approximately 5 times faster than the PVWatts 5-based method, with the median execution time for a stream varying between 6 and 8 seconds vs. a median run time of 31 seconds for the PVWatts 5-based method.

algorithm validation↗

IDB Data Loader

The International Database of Reference Gamma-Ray Spectra of Various Nuclear Matter is designed to hold curated gamma spectral data and will be hosted by the International Atomic Energy Agency on its public facing web site. Currently, the database to be hosted is given to the International Atomic Energy Agency by Sandia. This document describes the application used by Sandia to load spectral data into a database.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology↗

Science Validation for Dark Energy Research with Optical Imaging Surveys

The Universe has been expanding at an accelerating rate over the past several billion years, as though an unknown form of dark energy permeates all of space. The ultimate scientific goal of the proposed research is to distinguish between different physical mechanisms that could account for this observed accelerated expansion, for example, the zero-point energy of the vacuum, a dynamical form of energy that varies in time and/or space, or a modification to our theory of gravity. Wide- field optical imaging surveys of the night sky can test these competing models by measuring both the cosmic expansion history and the growth of large-scale structure. To this end, the Dark Energy Survey (DES) has cataloged several hundred million galaxies and thousands of supernovae. The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), will enlarge the census to billions of galaxies and hundreds of thousands of supernovae. A critical question for these dark energy experiments is whether systematic uncertainties can continue to be controlled at a level to keep pace with the statistical precision offered by such enormous datasets. The immediate research objectives of this project were (1) to prepare and validate input datasets that are the foundation of cosmological analyses with DES, and (2) to prepare for value-added characterization of Rubin Observatory commissioning data to inform early operations and accelerate the realization of dark energy science from LSST data products. For DES, we assembled and curated cosmology-ready data releases that include value-added components such as enhanced photometric and astrometric calibrations, alternative source extraction algorithms, maps of the survey coverage and survey conditions, object classifications, object quality selections, galaxy shapes, and photometric redshifts. We used the galaxy clustering technique to validate the photometric redshift distributions of various galaxy samples to be used as lenses in combined studies of galaxy clustering and weak gravitational lensing. The galaxy clustering redshift analysis was enhanced by use of a larger sample of reference galaxies from the eBOSS spectroscopic survey that extends to higher redshifts. For LSST, we prepared for science validation studies of commissioning data aimed at dark energy science capability that extend beyond the normative system-level tests to be done by the Rubin Observatory Construction Project. We identified a set of proposed survey strategies and candidate target fields that could be observed during the commissioning period to enhance science validation activities related to studies of dark energy.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology

PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArchive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide.

59 BASIC BIOLOGICAL SCIENCES↗

CoRE MOF DB: A curated experimental metal-organic framework database with machine-learned properties for integrated material-process screening

Here, we present an updated version of the Computation-Ready, Experimental (CoRE) Metal-Organic Framework (MOF) database, which includes a curated set of computation-ready MOF crystal structures designed for high-throughput computational materials discovery. Data collection and curation procedures were improved from the previous version to enable more frequent updates in the future. Machine-learning-predicted properties, such as stability metrics and heat capacities, are included in the dataset to streamline screening activities. An updated version of MOFid was developed to provide detailed information on metal nodes, organic linkers, and topologies of an MOF structure. DDEC6 partial atomic charges of MOFs were assigned based on a machine-learning model. Gibbs ensemble Monte Carlo simulations were used to classify the hydrophobicity of MOFs. The finalized dataset was subsequently used to perform integrated material-process screening for various carbon-capture conditions using high-fidelity temperature-swing adsorption (TSA) simulations. Our workflow identified multiple MOF candidates that are predicted to outperform CALF-20 for these applications.

CoRE MOF database↗