Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “research data management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Geospatial Data Platform for All

Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.

data platform↗

A Lakehouse Architecture for the Management and Analysis of Heterogeneous Data for Biomedical Research and Mega-biobanks

Data Lakehouse is a new paradigm in data architectures that embodies and integrates already established concepts for the systematic management of disparate, large-scale data – a data lake for heterogeneous data management, use of open standards for high-performance querying, and systematic maintenance of the data "freshness". In addition to being a new concept, the data lakehouse is also still a conceptual construct. Many projects that use the lakehouse require maturing, empirical studies, and specific implementations. In this paper, we present our implementation of the data lakehouse concept in a biomedical research and health data analytics domain, and we discuss the implementation of some unique and novel features such as support for specialized access controls in support of HIPAA regulation and IRB protocols, and support for the FAIR standard.

Begoli, Edmon↗

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)↗

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Open Source Scalable Data Services and Data Fusion for Biological and Environmental Sciences (SBIR Phase I Final Scientific/ Technical Report)

The overarching goal of the project is to develop an integrated open-source scientific data management system (Apache V2 license), ResonantEco, that meets the need of biological and environmental researchers and developers for data management, curation, and data processing for analyses with a wide range of scale and complexity. ResonantEco will provide web enabled data services with features such as unified data interfaces and federated views of data and metadata for heterogeneous data sources with an interactive web client for data exploration. Our use of the term fusion is taken from geospatial (GIS) domain where data fusion is often synonymous with data integration. In particular, data integration in ResonantEco involves combining data residing in different sources and providing users with a unified view of them.

99 GENERAL AND MISCELLANEOUS↗

Seismology in the Cloud: A New Streaming Workflow

Data-intensive research in seismology is experiencing a recent boom, driven in part by large volumes of available data and advances in the growing field of data science. However, there are significant barriers to processing large data volumes, such as long retrieval times from data repositories, complex data management, and limited computational resources. New tools and platforms have reduced the barriers to entry for scientific cluster computing, including the maturation of the commercial cloud as an accessible instrument for research. Here, we build a customized research cluster in the cloud to test a new workflow for large-scale seismic analysis, in which data are processed as a stream (retrieved on-the-fly and acted upon without storing), with data from the Incorporated Research Institutions for Seismology Data Management Center. We use this workflow to deploy a spectral peak detection algorithm over 5.6 TB of compressed continuous seismic data from 2074 stations of the USArray Transportable Array EarthScope network. Using a 50-node cluster in the cloud, we completed the noise survey in 80 hr, with an average data throughput of 1.7 GB per minute. By varying cluster sizes, we find the scaling of our analysis to be sublinear, due to a combination of algorithmic limitations and data center response times. The cloud-based streaming workflow represents an order-of-magnitude increase in acquisition and processing speed compared to a traditional download-store-process workflow, and offers the additional benefits of employing a flexible, accessible, and widely used computing architecture. It is limited, however, due to its reliance on Internet transfer speeds and data center service capacity, and may not work well for repeated analyses or those for which even higher data throughputs are needed. These research applications will require a new class of cloud-native approaches in which both data and analysis are in the cloud.

58 GEOSCIENCES↗

Sandia National Laboratories Ecosystem for Open Science: Metadata Schema v0.2 Description.

The Ecosystem for Open Science (eOS) initiative was established in 2019. Its objective is improving openness and sharing of data and information across Defense Nuclear Nonproliferation (DNN) Research and Development (R&D) activities. To support this initiative, the eOS team at Sandia National Laboratories (SNL) developed metadata and data standards and proposed a machine-readable metadata schema. The nuclear explosion monitoring field was selected as a focus area due to its the wide range of pertinent phenomenologies.We developed the DCAT-eOS-AP metadata schema extending the Data Catalog Vocabulary version 2 (DCATv2) standard using an application profile (AP), to fit the needs of multi-disciplinary NA-22 projects. The DCAT-eOS-AP metadata schema describes data at different levels of granularity ranging from general descriptions to more domain-specific granular metadata. Its implementation and serialization is flexible with the ability to include new file or data types. Thus, it will scale with the ever-increasing data management needs of government research. Due to the multitude of phenomenologies represented in the DCAT-eOS-AP schema, we anticipate that it will be easily extensible to various projects across many DOE mission areas. This document describes data management challenges faced within the DNN R&D portfolio and provides insight on how metadata and data standards/guidelines combined with a comprehensive metadata schema can add value to programs throughout the Department of Energy (DOE). It reviews the importance of metadata standards, FAIR (Findability, Accessibility, Interoperability, and Reusability) data principles, and metadata schemas. Additionally, it summarizes input from subject matter experts (SME) at SNL and other National Laboratories that resulted in metadata and data standards/guidelines encompassing domains relevant to NA-22 projects. Finally, we discuss the DCAT-eOS-AP metadata development. Implementation recommendations and future development directions are included for those keen on adopting the DCAT-eOS-AP metadata schema.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Curifactory: A research experiment manager

Curifactory is a command line tool and framework for organizing Python experiment code, configuration parameters, and results. It is an opinionated and lightweight approach to workflow management infrastructure and is primarily intended to support researchers conducting experiments on one machine. This software was developed to support the reproducibility of results for several data science projects in the Nuclear Nonproliferation Division at Oak Ridge National Laboratory. Curifactory is intended to be a general framework and is not specific to machine learning or data science. It can aid in any field in which experiments are primarily computation-based studies and can be implemented in Python (e.g., high-energy physics, astronomy, computational chemistry). Here, the design emphasizes the automated caching of intermediate data analysis artifacts to speed up development involving computationally intensive tasks. It also allows for data provenance and experiment reproduction. Individual experiment runs are tracked through logs and their output reports, and entire copies of a run with all cached data and metadata can be exported for others to run using Curifactory on another machine. Curifactory experiments can either be integrated into a project from the beginning or can be written on top of an existing codebase without needing significant modification. A few important views of the Curifactory library can be seen in Figure 1.

97 MATHEMATICS AND COMPUTING↗

AmeriFlux FLUXNET-1F US-HWB USDA ARS Pasture Sytems and Watershed Management Research Unit- Hawbecker Site

This is the AmeriFlux Management Project (AMP) created FLUXNET-1F version of the carbon flux data for the site US-HWB USDA ARS Pasture Sytems and Watershed Management Research Unit- Hawbecker Site. This is the FLUXNET version of the carbon flux data for the site US-HWB USDA ARS Pasture Sytems and Watershed Management Research Unit- Hawbecker Site produced by applying the standard ONEFlux (1F) software. Site Description - Hawbecker farm is owned by Penn State University. The farming that took place was performed by their Farm Operations Division. The ground is rolling terrain, next to wooded areas, the Beef and Sheep Reasearch Farm, The University Airport, and other large fields maintained by Farm Operations. At the time of this collection period, the site housed another Meteorological Site, a Phenocam, and the GraceNet plots. Crop Rotation during period was 2 years Alfalfa, 1 year Corn for grain and 1 year Wheat.

Goslee, Sarah↗

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache↗

Energy Material Network Data Hubs

In early 2015 the United States Department of Energy conceived of a consortium of collaborative bodies based on shared expertise, data, and resources that could be targeted towards the more difficult problems in energy materials research. The concept of virtual laboratories had been envisioned and discussed earlier in the decade in response to the advent of the Materials Genome Initiative and similar scientific thrusts. To be effective, any virtual laboratory needed a robust method for data management, communication, security, data sharing, dissemination, and demonstration to work efficiently and effectively for groups of remote researchers. With the accessibility of new, easily deployed cloud technology and software frameworks, such individual elements could be integrated, and the required collaboration architecture is now possible. The developers have leveraged open-source software frameworks, customized them, and merged them into a platform to enable collaborative energy materials science, regardless of the geographic dispersal of the people and resources. After five years in operations, the systems are demonstratively an effective platform for enabling research within the Energy Material Networks (EMN). This paper will show the design and development of a secured scientific data sharing platform, the ability to customize the system to support diverse workflows, and examples of the enabled research and results connected with some of the Energy Material Networks.

97 MATHEMATICS AND COMPUTING↗

AmeriFlux US-HWB USDA ARS Pasture Sytems and Watershed Management Research Unit- Hawbecker Site

This is the AmeriFlux version of the carbon flux data for the site US-HWB USDA ARS Pasture Sytems and Watershed Management Research Unit- Hawbecker Site. Site Description - Hawbecker farm is owned by Penn State University. The farming that took place was performed by their Farm Operations Division. The ground is rolling terrain, next to wooded areas, the Beef and Sheep Reasearch Farm, The University Airport, and other large fields maintained by Farm Operations. At the time of this collection period, the site housed another Meteorological Site, a Phenocam, and the GraceNet plots. Crop Rotation during period was 2 years Alfalfa, 1 year Corn for grain and 1 year Wheat.

Goslee, Sarah↗

NDMAS System and Process Description

The U. S. Department of Energy (DOE) has made a significant investment in research to develop the next generation of reactor technologies as well as to improve the performance and lengthen the life cycle of existing nuclear reactors. Data collected to demonstrate new concepts may also be used in the future to support licensing of these technologies. Provenance of these data must be preserved. The Nuclear Data Management and Analysis System (NDMAS) was established to manage and preserve data collected by fuels and materials research conducted by the high-temperature, gas-cooled reactor program. The scope of NDMAS is expanding to include other nuclear research programs that have the shared need to preserve the provenance of research data. Nuclear research funded by DOE is conducted by Idaho National Laboratory (INL), universities, other national laboratories, foreign research partners, and private companies. This research will generate a large amount of data from a variety of sources over a period of many years. Managing the data generated by the research and development projects presents a significant challenge for retaining data integrity and availability.

Transforming Drainage Research Data (USDA-NIFA Award No. 2015-68007-23193)

This dataset contains research data compiled by the “Managing Water for Increased Resiliency of Drained Agricultural Landscapes” project a.k.a. Transforming Drainage. This project was funded from 2015-2021 by the United States Department of Agriculture, National Institute of Food and Agriculture (USDA-NIFA, Award No. 2015-68007-23193). Data are also available from a separate web-accessible application (drainagedata.org). At drainagedata.org, users can visualize the data with customized tools, query based on specific sites and measurements of interest, and access site photographs, maps, summaries, and publications. Additional data or edits made following the publication of this data here at USDA NAL Ag Data Commons will be posted under the Versions tab on drainagedata.org. These data began in 1996 and include plot- and field-level measurements for 39 experiments across the Midwest and North Carolina. Practices studied include controlled drainage, drainage water recycling, and saturated buffers. In total, 219 variables are reported and span 207 site-years for tile drainage, 154 for nitrate-N load, 181 for water quality, 92 for water table, and 201 for crop yield.

Modeling↗

Data for publication: "A fresh take: Seasonal changes in terrestrial freshwater inputs impact salt marsh hydrology and vegetation dynamics"

This data repository contains data associated with the manuscript "A fresh take: Seasonal changes in terrestrial freshwater inputs impact salt marsh hydrology and vegetation dynamics". This study was conducted at the Elkhorn Slough National Estuarine Research Reserve in Watsonville, California from October 2019 - June 2022. We sought to understand the role of shallow freshwater inputs from adjacent uplands on salt marsh hydrologic behavior and vegetation productivity. This dataset contains CSV files of the following: daily salt marsh subsurface water level and pore water conductivity, monthly vegetation survey measurements, soil core data. Estuary surface water level, conductivity, and local precipitation were downloaded from the National Estuarine Research Reserve System (Centralized Data Management Office) at https://cdmo.baruch.sc.edu/.

54 ENVIRONMENTAL SCIENCES↗