Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “system metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Generic, Extensible, Configurable Push-Pull Framework for Large-Scale Science Missions

The push-pull framework was developed in hopes that an infrastructure would be created that could literally connect to any given remote site, and (given a set of restrictions) download files from that remote site based on those restrictions. The Cataloging and Archiving Service (CAS) has recently been re-architected and re-factored in its canonical services, including file management, workflow management, and resource management. Additionally, a generic CAS Crawling Framework was built based on motivation from Apache s open-source search engine project called Nutch. Nutch is an Apache effort to provide search engine services (akin to Google), including crawling, parsing, content analysis, and indexing. It has produced several stable software releases, and is currently used in production services at companies such as Yahoo, and at NASA's Planetary Data System. The CAS Crawling Framework supports many of the Nutch Crawler's generic services, including metadata extraction, crawling, and ingestion. However, one service that was not ported over from Nutch is a generic protocol layer service that allows the Nutch crawler to obtain content using protocol plug-ins that download content using implementations of remote protocols, such as HTTP, FTP, WinNT file system, HTTPS, etc. Such a generic protocol layer would greatly aid in the CAS Crawling Framework, as the layer would allow the framework to generically obtain content (i.e., data products) from remote sites using protocols such as FTP and others. Augmented with this capability, the Orbiting Carbon Observatory (OCO) and NPP (NPOESS Preparatory Project) Sounder PEATE (Product Evaluation and Analysis Tools Elements) would be provided with an infrastructure to support generic FTP-based pull access to remote data products, obviating the need for any specialized software outside of the context of their existing process control systems. This extensible configurable framework was created in Java, and allows the use of different underlying communication middleware (at present, both XMLRPC, and RMI). In addition, the framework is entirely suitable in a multi-mission environment and is supporting both NPP Sounder PEATE and the OCO Mission. Both systems involve tasks such as high-throughput job processing, terabyte-scale data management, and science computing facilities. NPP Sounder PEATE is already using the push-pull framework to accept hundreds of gigabytes of IASI (infrared atmospheric sounding interferometer) data, and is in preparation to accept CRIMS (Cross-track Infrared Microwave Sounding Suite) data. OCO will leverage the framework to download MODIS, CloudSat, and other ancillary data products for use in the high-performance Level 2 Science Algorithm. The National Cancer Institute is also evaluating the framework for use in sharing and disseminating cancer research data through its Early Detection Research Network (EDRN).

Foster, Brian M.↗

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the "FAIRness" of NASA's GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.↗

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the “FAIRness” of NASA’s GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.↗

Building a Virtual Solar Observatory: I Look Around and There's a Petabyte Following Me

The 2001 July NASA Senior Review of Sun-Earth Connections missions and data centers directed the Solar Data Analysis Center (SDAC) to proceed in studying and implementing a Virtual Solar Observatory (VSO) to ease the identification of and access to distributed archives of solar data. Any such design (cf. the National Virtual Observatory and NASA's Planetary Data System) consists of three elements: the distributed archives, a "broker" facility that translates metadata from all partner archives into a single standard for searches, and a user interface to allow searching, browsing, and download of data. Three groups are now engaged in a six-month study that will produce a candidate design and implementation roadmap for the VSO. We hope to proceed with the construction of a prototype VSO in US fiscal year 2003, with fuller deployment dependent on community reaction to and use of the capability. We therefore invite as broad as possible public comment and involvement, and invite interested parties to a "birds of a feather" session at this meeting. VSO is partnered with the European Grid of Solar Observations (EGSO), and if successful, we hope to be able to offer the VSO as the basis for the solar component of a Living With a Star data system.

Gurman, J. B.↗

infrastore [SWR-26-077]

Infrastore is time-series storage for energy-systems simulations, backed by HDF5 + SQLite, with Rust, Python, Julia, gRPC, and CLI bindings. It is a Rust library for managing time-series data in power-systems and energy simulations. Numerical arrays are persisted in HDF5, and the metadata associating each array with its owning component lives in SQLite. Identical arrays are stored once and shared through content addressing. It ships native Rust, Python (PyO3), and Julia (C ABI) interfaces, the infrastore command-line tool, and a read-only gRPC server with a Rust client. Documentation: https://natlabrockies.github.io/infrastore/latest/ — start with the Quick Start or the Architecture.

Thom, Daniel [National Laboratory of the Rockies (↗

Hawaii Wave Surge Energy Converter (HAWSEC) OSU O.H. Hinsdale Basin

The following information and metadata applies to both the Phase I (Hydrodynamics) and Phase II (Full System Power Take-Off) zip folders which contain testing data from the OSU (Oregon State University) O.H. Hinsdale Wave Research Laboratory, from both OSU and the University of Hawaii at Manoa (UH). See zip folders provided further below in the downloads section. For experimental data of the full system, including PTO, see Phase II dataset. There are two main directories in each Phases's zip folder: "OSU_data" and "UH_data". The "OSU_data" directory contains data collected from their DAQ (data acquisition system), which includes all wave gauge observations, as well as body motions derived from their Qualisys motion tracking system. The organization of the directory follows OSU's convention. Detailed information on the instrument setup can be found under "OSU_data/docs/setup/instm_locations". The experiments conducted are documented in the "OSU_data/docs/daq_logs", which provides the trial number to the corresponding data located under "OSU_data/data" in several formats (e.g., ".mat" and ".txt"). Inside the trial directory, data is provided for each of the instruments defined in "OSU_data/docs/setup/instm_locations". The "UH_data" directory contains data collected from their DAQ. The data is stored in a ".tdms" file format. There are free plug-ins for Microsoft Excel and MathWorks MATLAB to read the ".tdms" format. Below are a few links providing methods to read in the data, but a Google search should identify alternatives sources if these no longer exist (valid as of January 2024): Excel: http://www.ni.com/example/27944/en/ MATLAB: https://www.mathworks.com/matlabcentral/fileexchange/30023-tdms-reader The Excel plugin is recommend to get a quick overview of the data. The UH data is organized by directory name, in which the sub-directories for each experiment contains a directory whose name defines the wave height and period for the experimental data within. For example, a directory name "H02_T0275" corresponds to an experiment with wave height 0.1m and a period of 2.75s. For random wave data, the gamma value is also included in the directory name. For example, a directory name "H02_T0225_G18" corresponds to an experiment with a significant wave height of 0.2m, a peak period of 2.25s, and a gamma value of 1.8, with each spectra being a TMA spectrum. For the free decay experiments, the directory name is defined by the initial angular displacement. For example, a directory name "ang05_run01" corresponds to an experiment with an initial angular displacement of 5 degrees. There is a dataset in the UH data for each corresponding experiment defined in the OSU DAQ logs. The ".tdms" data is output from the DAQ at fixed intervals. Therefore, if multiple files are contained within the folder, the data will need to be stitched together. Within the UH dataset, there are two input channels from the OSU DAQ providing a random square wave signal for time synchronization ("ENV-WHT-0010") and a high/low signal ("ENV-WHT-0012") to identify when the wave maker is active (+5V). The UH data is logged as a collection of channel outputs. Channels not in use for the OSU testing (either Phase I or Phase II) are marked "nan" below. If the sensor is disconnected, it will record noise throughout the experiment. Below are the channel definitions in terms of what they measure: GPS Time = time CYL-POS-0001 = position between flap and fixed reference CYL-LCA-0001 = force between flap and hydraulic cylinder REC-LPT-0001 = nan REC-HPT-0001 = nan REC-HPT-0002 = nan REC-HPT-0003 = nan HHT-HPT-0001 = pressure at exhaust ("head" only) REC-FQC-0001 = nan REC-FQC-0002 = nan HHT-FQC-0001 = flow at exhaust ("head" only) ENV-WHT-0001 = nan ENV-WHT-0002 = nan ENV-WHT-0003 = nan ENV-WHT-0010 = random signal from OSU DAQ ENV-WHT-0012 = high/low signal from OSU DAQ Also included is a calibration curve to convert the string pot data to flap pi...

16 TIDAL AND WAVE POWER↗

Applying Waveform Correlation and Waveform Template Metadata to Mining Blasts to Reduce Analyst Workload

Organizations that monitor for underground nuclear explosive tests are interested in techniques that automatically characterize mining blasts to reduce the human analyst effort required to produce high - quality event bulletins. Waveform correlation is effective in finding similar waveforms from repeating seismic events, including mining blasts. In this study we use waveform template event metadata to seek corroborating detections from multiple stations in the International Monitoring System of the Preparatory Commission for the Comprehensive Nuclear-Test-Ban Treaty Organization. We build upon events detected in a prior waveform correlation study of mining blasts in two geographic regions, Wyoming and Scandinavia. Using a set of expert analyst-reviewed waveform correlation events that were declared to be true positive detections, we explore criteria for choosing the waveform correlation detections that are most likely to lead to bulletin-worthy events and reduction of analyst effort.

47 OTHER INSTRUMENTATION↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Using semantic data modeling techniques to organize an object-oriented database for extending the mass storage model

A methodology for optimizing organization of data obtained by NASA earth and space missions is discussed. The methodology uses a concept based on semantic data modeling techniques implemented in a hierarchical storage model. The modeling is used to organize objects in mass storage devices, relational database systems, and object-oriented databases. The semantic data modeling at the metadata record level is examined, including the simulation of a knowledge base and semantic metadata storage issues. The semantic data model hierarchy and its application for efficient data storage is addressed, as is the mapping of the application structure to the mass storage.

Campbell, William J.↗

Data catalog for JPL Physical Oceanography Distributed Active Archive Center (PO.DAAC)

The Physical Oceanography Distributed Active Archive Center (PO.DAAC) archive at the Jet Propulsion Laboratory contains satellite data sets and ancillary in-situ data for the ocean sciences and global-change research to facilitate multidisciplinary use of satellite ocean data. Geophysical parameters available from the archive include sea-surface height, surface-wind vector, surface-wind speed, surface-wind stress vector, sea-surface temperature, atmospheric liquid water, integrated water vapor, phytoplankton pigment concentration, heat flux, and in-situ data. PO.DAAC is an element of the Earth Observing System Data and Information System and is the United States distribution site for TOPEX/POSEIDON data and metadata.

Digby, Susan↗

Ocean Data from MODIS at the NASA Goddard DAAC

Terra satellite carrying the Moderate Resolution Imaging Spectroradiometer (MODIS) was successfully launched on December 18, 1999. Some of the 36 different wavelengths that MODIS samples have never before been measured from space. New ocean data products, which have not been derived on a global scale before, are made available for research to the scientific community. For example, MODIS uses a new split window in the four-micron region for the better measurement of Sea Surface Temperature (SST), and provides the unprecedented ability (683 nm band) to measure chlorophyll fluorescence. At full ocean production, more than a thousand different ocean products in three major categories (ocean color, sea surface temperature, and ocean primary production) are archived at the NASA Goddard Earth Sciences (GES) Distributed Active Archive Center (DAAC) at the rate of approx. 230GB/day. The challenge is to distribute such large volumes of data to the ocean community. It is achieved through a combination of public and restricted EOS Data Gateways, the GES DAAC Search and Order WWW interface, and an FTP site that contains samples of MODIS data. A new Search and Order WWW interface at http://acdisx.gsfc.nasa.gov/data/ developed at the GES DAAC is based on a hierarchical organization of data, will always return non-zero results. It has a very convenient geographical representation of five-minute data granule coverage for each day MODIS Data Support Team (MDST) continues the tradition of quality support at the GES DAAC for the ocean color data from the Coastal Zone Color Scanner (CZCS) and the Sea Viewing Wide Field-of-View Sensor (SeaWiFS) by providing expert assistance to users in accessing data products, information on visualization tools, documentation for data products and formats (Hierarchical Data Format-Earth Observing System (HDF-EOS)), information on the scientific content of products and metadata. Visit the MDST website at http://daac.gsfc.nasa.gov/CAMPAIGN DOCS/MODIS/index.html

Leptoukh, Gregory G.↗

Using Open and Interoperable Ways to Publish and Access LANCE AIRS Near-Real Time Data

The Atmospheric Infrared Sounder (AIRS) Near-Real Time (NRT) data from the Land Atmosphere Near real-time Capability for EOS (LANCE) element at the Goddard Earth Sciences Data and Information Services Center (GES DISC) provides information on the global and regional atmospheric state, with very low temporal latency, to support climate research and improve weather forecasting. An open and interoperable platform is useful to facilitate access to, and integration of, LANCE AIRS NRT data. As Web services technology has matured in recent years, a new scalable Service-Oriented Architecture (SOA) is emerging as the basic platform for distributed computing and large networks of interoperable applications. Following the provide-register-discover-consume SOA paradigm, this presentation discusses how to use open-source geospatial software components to build Web services for publishing and accessing AIRS NRT data, explore the metadata relevant to registering and discovering data and services in the catalogue systems, and implement a Web portal to facilitate users' consumption of the data and services.

Zhao, Peisheng↗

Bias control for a memory device

Methods, systems, and devices for bias control for a memory device are described. A memory system may store indication of whether data is coherent. In some examples, the indication may be stored as metadata, where a first value indicates that the data is not coherent and a second value or a third value indicate that the data is coherent. When a processing unit or other component of the memory system processes a command to access data, the memory system may operate according to a device bias mode when the indication is the first value, and according to a host bias mode when the indication is the second value or the third value.

97 MATHEMATICS AND COMPUTING↗

Bias control for a memory device

Methods, systems, and devices for bias control for a memory device are described. A memory system may store indication of whether data is coherent. In some examples, the indication may be stored as metadata, where a first value indicates that the data is not coherent and a second value or a third value indicate that the data is coherent. When a processing unit or other component of the memory system processes a command to access data, the memory system may operate according to a device bias mode when the indication is the first value, and according to a host bias mode when the indication is the second value or the third value.

Walker, Dean↗

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗