Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

NSSDC provides network access to key data via NDADS

The National Space Science Data Center (NSSDC) is making a growing fraction of its most customer-desirable data electronically accessible via both the local and wide area networks. NSSDC is witnessing a great increase in its data dissemination owing to this network accessibility. To provide its customers the best data accessibility, the NSSDC makes data available from a nearline, mass storage system, the NSSDC Data Archive and Dissemination Service (NDADS). The NDADS, the initial version was made available in January 1992, is a customized system of hardware and software that provides users access to the nearline data via ANONYMOUS FTP, an e-mail interface (ARMS), and a C-based software library. In January 1992, the NDADS registered 416 requests for 1,957 files. By December of 1994, NDADS had been populated with 800 gigabytes of electronically accessible data and had registered 1458 requests for 20,887 files. In this report we describe the NDADS system, both hardware and software. Later in the report, we discuss some of the lessons that were learned as a result of operating NDADS, particularly in the area of ingest and dissemination.

Behnke, Jeanne↗

Advanced communications technologies for image processing

It is essential for image analysts to have the capability to link to remote facilities as a means of accessing both data bases and high-speed processors. This can increase productivity through enhanced data access and minimization of delays. New technology is emerging to provide the high communication data rates needed in image processing. These developments include multi-user sharing of high bandwidth (60 megabits per second) Time Division Multiple Access (TDMA) satellite links, low-cost satellite ground stations, and high speed adaptive quadrature modems that allow 9600 bit per second communications over voice-grade telephone lines.

Likens, W. C.↗

Online Visualization and Value Added Services of MERRA-2 Data at GES DISC

NASA climate reanalysis datasets from MERRA-2, distributed at the Goddard Earth Sciences Data and Information Services Center (GES DISC), have been used in broad research areas, such as climate variations, extreme weather, agriculture, renewable energy, and air quality, etc. The datasets contain numerous variables for atmosphere, land, and ocean, grouped into 95 products. The total archived volume is approximately 337 TB ( approximately 562K files) at the end of October 2017. Due to the large number of products and files, and large data volumes, it may be a challenge for a user to find and download the data of interest. The support team at GES DISC, working closely with the MERRA-2 science team, has created and is continuing to work on value added data services to best meet the needs of a broad user community. This presentation, using aerosol over Asia Monsoon as an example, provides an overview of the MERRA-2 data services at GES DISC, including: How to find the data? How many data access methods are provided? What are the best data access methods for me? How do download the subsetted (parameter, spatial, temporal) data and save in preferred spatial resolution and data format? How to visualize and explore the data online? In addition, we introduce a future online analytic tool designed for supporting application research, focusing on long-term hourly time-series data access and analysis.

Shen, Suhung↗

PERSIANN-Unet: A Global Deep Learning Framework for Near-Real-Time Precipitation Estimation Using Infrared Data

Access to high-quality, high-resolution, near-real-time precipitation data is essential for hydrological and meteorological research and disaster mitigation. Traditional tools such as rain gauges and radar networks, though effective, have limitations, including sparse coverage in remote areas and high operational costs. Satellite data, with its global coverage and high spatial and temporal resolutions, mitigates limitations in coverage. Satellite precipitation products like Hydro Estimator (HE), Integrated Multi-satellitE Retrievals for Global Precipitation Measurement (IMERG), and Precipitation Estimation from Remotely Sensed Information using Artificial Neural Networks (PERSIANN) utilize both geosynchronous thermal infrared (IR) and passive microwave (PMW) data in their operation. PMW sensors offer detailed atmospheric profiles but suffer from higher latency, whereas IR sensors provide lower latency but only capture cloud-top information. Despite this constraint, IR data remains attractive for low-latency precipitation estimation. Recent advances in deep learning, particularly convolutional neural networks (CNNs), have further improved satellite precipitation retrievals. This study introduces PERSIANN-Unet (PUnet or PERSIANN V3), a quasi-global algorithm covering 60°N–60°S that combines IR data, monthly climatology, and the UNet architecture to produce half-hourly precipitation estimates at 0.04° resolution. The product is evaluated against HE, IMERG, and PDIR-Now for 2022–2023. Results show that PUnet closely matches its training target, IMERG V07 Final, at the global scale, and performance is further evaluated against Stage IV as a reference over CONUS. Training PUnet on IMERG (2016–2021) leverages a high-quality, integrated PMW IR-gauge precipitation product while developing an IR-based framework not reliant on PMW availability. By operating on a single global image, PUnet avoids tile partitioning and blending steps, reducing edge discontinuities, and produces more spatially consistent precipitation fields across hemispheres.

Phu Nguyen↗

Survey of Federal, National, and International standards applicable to the NASA applications data services

An applications data service (ADS) was developed to meet the challenges in the data access and integration. The ADS provides a common service to locate and access applications data electronically and integrate the cross correlative data sets required by multiple users. Its catalog and network services increase data visibility as well as provide the data in a more rapid manner and a usable form.

Kuch, T.↗

Observational Requirements and Measurement Systems Concepts

Six philosophical elements that a measurement strategy for dynamic qualities should contain are: systemic global observations; nested coverage; continuous upgrading of observational capability and testing of promising instrument concepts; orbit considerations; quality control: data validity and management; and data interpretation. The criteria for determining an observational program should include the following: significance of the primary scientific questions, technical readiness; potential of undeveloped techniques, engineering and economic feasibility; and data access and data systems.

Source record↗

WetNet operations

WetNet is an interdisciplinary Earth science data analysis and research project with an emphasis on the study of the global hydrological cycle. The project goals are to facilitate scientific discussion, collaboration, and interaction among a selected group of investigators by providing data access and data analysis software on a personal computer. The WetNet system fulfills some of the functionality of a prototype Product Generation System (PGS), Data Archive and Distribution System (DADS), and Information Management System for the Distributed Active Archive Center. The PGS functionality is satisfied in WetNet by processing the Special Sensor Microwave/Imager (SSM/I) data into a standard format (McIDAS) data sets and generating geophysical parameter Level II browse data sets. The DADS functionality is fulfilled when the data sets are archived on magneto optical cartridges and distributed to the WetNet investigators. The WetNet data sets on the magneto optical cartridges contain the complete WetNet processing, catalogue, and menu software in addition to SSM/I orbit data for the respective two week time period.

Goodman, H. Michael↗

Strawman Philosophical Guide for Developing International Network of GPM GV Sites

The creation of an international network of ground validation (GV) sites that will support the Global Precipitation Measurement (GPM) Mission's international science programme will require detailed planning of mechanisms for exchanging technical information, GV data products, and scientific results. An important component of the planning will be the philosophical guide under which the network will grow and emerge as a successful element of the GPM Mission. This philosophical guide should be able to serve the mission in developing scientific pathways for ground validation research which will ensure the highest possible quality measurement record of global precipitation products. The philosophical issues, in this regard, partly stem from the financial architecture under which the GV network will be developed, i.e., each participating country will provide its own financial support through committed institutions -- regardless of whether a national or international space agency is involved.At the 1st International GPM Ground Validation Workshop held in Abingdon, UK in November-2003, most of the basic tenants behind the development of the international GV network were identified and discussed. Therefore, with this progress in mind, this presentation is intended to put forth a strawman philosophical guide supporting the development of the international network of GPM GV sites, noting that the initial progress has been reported in the Proceedings of the 1st International GPM GV Workshop -- available online. The central philosophical issues themselves, all flow from the fact that each participating institution can only bring to the table, GV facilities and scientific personnel that are affordable to the sanctioning (funding) national agency (be that a research, research-support, or operational agency). This situation imposes on the network, heterogeneity in the measuring sensors, data collection periods, data collection procedures, data latencies, and data reporting capabilities. Therefore, in order for the network to be effective in supporting the central scientific goals of the GPM mission, there must be a basic agreed upon doctrine under which the network participants function vis-a-vis: (1) an overriding set of general scientific requirements, (2) a minimal set of policies governing the free flow of GV data between the scientific participants, (3) a few basic definitions concerning the prioritization of measurements and their respective value to the mission, (4) a few basic procedures concerning data formats, data reporting procedures, data access, and data archiving, and (5) a simple means to differentiate GV sites according to their level of effort and ability to perform near real-time data acquisition - data reporting tasks. Most important, in case they choose to operate as a near real-time data collection-data distribution site, they would be expected to operate under a fairly narrowly defined protocol needed to ensure smooth GV support operations. This presentation will suggest measures responsive to items (1) - (5) from which to proceed,. In addition, this presentation will seek to stimulate discussion and debate concerning how much heterogeneity is tolerable within the eventual GV site network, given that the any individual GV site can only be considered scientifically useful if it supports the achievement of the central GPM Mission goals. Only ground validation research that has a direct connection to the space mission should be considered justifiable given the overarching scientific goals of the mission. Therefore each site will have to seek some level of accommodation to what the GPM Mission requires in the way of retrieval error characterization, retrieval error detection and reporting, and generation of GV data products that support assessment and improvement of the mission's standard precipitation retrieval algorithms. These are all important scientific issues that will be best resolved in open scientific debate.

Smith, Eric A.↗

Remote Data Exploration with the Interactive Data Language (IDL)

A difficulty for many NASA researchers is that often the data to analyze is located remotely from the scientist and the data is too large to transfer for local analysis. Researchers have developed the Data Access Protocol (DAP) for accessing remote data. Presently one can use DAP from within IDL, but the IDL-DAP interface is both limited and cumbersome. A more powerful and user-friendly interface to DAP for IDL has been developed. Users are able to browse remote data sets graphically, select partial data to retrieve, import that data and make customized plots, and have an interactive IDL command line session simultaneous with the remote visualization. All of these IDL-DAP tools are usable easily and seamlessly for any IDL user. IDL and DAP are both widely used in science, but were not easily used together. The IDL DAP bindings were incomplete and had numerous bugs that prevented their serious use. For example, the existing bindings did not read DAP Grid data, which is the organization of nearly all NASA datasets currently served via DAP. This project uniquely provides a fully featured, user-friendly interface to DAP from IDL, both from the command line and a GUI application. The DAP Explorer GUI application makes browsing a dataset more user-friendly, while also providing the capability to run user-defined functions on specified data. Methods for running remote functions on the DAP server were investigated, and a technique for accomplishing this task was decided upon.

Galloy, Michael↗

GeoNEX: A Cloud Gateway for Near Real-time Processing of Geostationary Satellite Products

The emergence of a new generation of geostationary satellite sensors provides land andatmosphere monitoring capabilities similar to MODIS and VIIRS with far greater temporal resolution (5-15 minutes). However, processing such large volume, highly dynamic datasets requires computing capabilities that (1) better support data access and knowledge discovery for scientists; (2) provide resources to enable real-time processing for emergency response (wildfire, smoke, dust, etc.); and (3) provide reliable and scalable services for the broader user community. This paper presents an implementation of GeoNEX (Geostationary NASA-NOAA Earth Exchange) services that integrate scientific algorithms with Amazon Web Services (AWS) to provide near realtime monitoring (~5 minute latency) capability in a hybrid cloud-computing environment. It offers a user-friendly, manageable and extendable interface and benefits from the scalability provided by Amazon Web Services. Four use cases are presented to illustrate how to (1) search and access geostationary data; (2) configure computing infrastructure to enable near real-time processing; (3) disseminate and utilize research results, visualizations, and animations to concurrent users; and (4) use a Jupyter Notebook-like interface for data exploration and rapid prototyping. As an example of (3), the Wildfire Automated Biomass Burning Algorithm (WF_ABBA) was implemented on GOES-16 and -17 data to produce an active fire map every 5 minutes over the conterminous US. Details of the implementation strategies, architectures, and challenges of the use cases are discussed.

GeoNEX↗

TPSAS-NF1676L-16833-DND

Semantic Infrastructure is central to realizing the first goal of the ASDC's Strategic Plan: expanding the ASDC's customer base by improving access to ASDC data. ASDC data comprises a widely heterogeneous set of complex products which presents two significant challenges in data access: Helping customers discover, among many available options, the most suitable data products for their purpose; and Guiding customers to easily and appropriately use products. Data products differ significantly in terms of how the data was collected and processed, even with similar subject matter. Understanding differences is critical to using data effectively. To reach a broader customer range, the ASDC must provide prospective users with enough information to quickly and meaningfully compare and evaluate data products. Data formats and structures also differ among products. Applications displaying and analyzing data need access to federated and semantically disambiguated data. Semantic technologies offer functionality for addressing this issue. Ontologies can provide robust, stable domain models serving as common schema for discovering, evaluating, comparing, and integrating data from disparate products. Reasoning engines and triple stores can leverage ontologies to support intelligent search applications allowing users to discover, query, retrieve, and easily reformat data from a broad spectrum of sources.

Beth Huffer↗

Why We Do What We Do: Data Reuse, Open Access and Privacy in Data Management at the Life Sciences Data Archive

As custodians of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and the archive’s evolving data management practices; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of data analysis and aggregation tools.

data management↗

The Virtual Space Physics Observatory: Quick Access to Data and Tools

The Virtual Space Physics Observatory (VSPO; see http://vspo.gsfc.nasa.gov) has grown to provide a way to find and access about 375 data products and services from over 100 spacecraft/observatories in space and solar physics. The datasets are mainly chosen to be the most requested, and include most of the publicly available data products from operating NASA Heliophysics spacecraft as well as from solar observatories measuring across the frequency spectrum. Service links include a "quick orbits" page that uses SSCWeb Web Services to provide a rapid answer to questions such as "What spacecraft were in orbit in July 1992?" and "Where were Geotail, Cluster, and Polar on 2 June 2001?" These queries are linked back to the data search page. The VSPO interface provides many ways of looking for data based on terms used in a registry of resources using the SPASE Data Model that will be the standard for Heliophysics Virtual Observatories. VSPO itself is accessible via an API that allows other applications to use it as a Web Service; this has been implemented in one instance using the ViSBARD visualization program. The VSPO will become part of the Space Physics Data Facility, and will continue to expand its access to data. A challenge for all VOs will be to provide uniform access to data at the variable level, and we will be addressing this question in a number of ways.

Cornwell, Carl↗

Why We Do What We Do: Data Reuse, Open Access, and Privacy in Data Management at the Life Sciences Data Archive

As custodian of the unique and irreplaceable collections of human subject research data generated by the Human Research Program and its predecessors throughout the agency’s history, the Life Sciences Data Archive (LSDA) is charged with protecting participants’ privacy and implementing their consent decisions as it provides retrospective data for use in new studies. This active, stewardship-focused approach to data management and preservation shapes the products that LSDA provides to researchers and the responsibilities of researchers in using the data and publishing their results. This presentation reviews how federal and agency mandates shape LSDA’s data management procedures and expectations for researchers. Topics covered will include LSDA’s movement towards implementation of the FAIR (Findable, Accessible, Interoperable, Reusable) principles and how the archive’s evolving data management practices support FAIR-ness; collaboration between LSDA and the Lifetime Surveillance of Astronaut Health (LSAH) project (the repository of astronaut medical data); LSDA’s response to the challenges of performing its stewardship role and maintaining trust given the public profiles of the subjects whose data it preserves; and the ever-increasing challenges to expectations of subject privacy stemming from the growing power and ubiquity of of data analysis and aggregation tools.

Data↗

Project Integration Architecture (PIA) and Computational Analysis Programming Interface (CAPRI) for Accessing Geometry Data from CAD Files

Integration of a supersonic inlet simulation with a computer aided design (CAD) system is demonstrated. The integration is performed using the Project Integration Architecture (PIA). PIA provides a common environment for wrapping many types of applications. Accessing geometry data from CAD files is accomplished by incorporating appropriate function calls from the Computational Analysis Programming Interface (CAPRI). CAPRI is a CAD vendor neutral programming interface that aids in acquiring geometry data directly from CAD files. The benefits of wrapping a supersonic inlet simulation into PIA using CAPRI are; direct access of geometry data, accurate capture of geometry data, automatic conversion of data units, CAD vendor neutral operation, and on-line interactive history capture. This paper describes the PIA and the CAPRI wrapper and details the supersonic inlet simulation demonstration.

Benyo, Theresa L.↗

Governing Data Findability, Accessibility, Interoperability and Reusability (FAIR) Compliance

The most recent data strategy documents at both the federal and NASA levels stipulate that systems should strive for the data they manage to be Findable, Accessible, Interoperable, and Reusable (FAIR). The NASA Life Sciences Portal (NLSP) has already begun leading efforts in this area for HRP, initiating efforts to comply with the FAIR principles. The broad interpretation of the FAIR principles has led to a plethora of tools that use a splay of metrics specifically but variably developed to judge how compliant data and systems are with the principles. A recent review [3] identified and studied 1,180 metrics across 20 publicly available tools for checking FAIR compliance of data and systems. Because of their very recent development, many organizations and data systems managers and developers have not yet had adequate time or resources to understand these FAIR compliance tools and metrics, their variations in design, accuracy or ease of application to their specific data sets and systems. Thus, it would be best for larger organizations like NASA to approach formulating a strategy for governance of FAIR compliance that can be flexibly applied and is adaptable to an evolving awareness knowledge of FAIR compliance methods and tools. In September 2024, the NASA Science Mission Directorate(SMD) organized a workshop on NASA science data repositories, including the topics of implementing FAIR and governing FAIR compliance across SMD. The initial part of these FAIR discussions focused on developing consensus around required science metadata fields. This is challenging given the diverse nature of NASA’s scientific data portfolio, the variety of metadata models and vocabularies used, and variable level of resources available to curate these data. Later discussion focused on three possible approaches to governing FAIR compliance: distributed, in which various programs, projects or systems define their own methods for assessing FAIR compliance, reporting results up appropriate management lines; centralized, in which higher-level organization(s) specify compliance tools or methods for the various data systems; and multi-level, in which a group comprised of individuals with expertise from multiple levels with organizations is formed to provide guidance and/or specifications for governing FAIR compliance. We report on the recommendations this session yielded, and how these might be shaped specifically to help implement and govern the compliance with FAIR of Human Research Program data and systems.

governance↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Near 40 Years MERRA-2 Data at NASA GES DISC -Opportunity and Challenge to Support Extremes Study

To the end of 2019, 40 years NASA climate reanalysis data sets from the Modern Era Retrospective-analysis for Research and Applications, Version 2 (MERRA-2) will be available at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). MERRA-2 consists atmosphere, land, and ocean data, which may be used for the studies ranging from the short scale weather events to the large scale decadal vulnerabilities. The hourly products, such as precipitation, soil moisture, temperature, and aerosols etc., have been used widely to study extreme events.In supporting users from broad communities, GES DISC have developed various data access services, including subsetter for downloading only data of interest with preferred format; OPeNDAP - for machine-to-machine data access; and Giovanni- for online visualization and analysis, etc. A big challenge for extreme study is to downloading and processing long-term hourly or daily data. The data downloading performance is not very satisfied by many users with current services and the native archived data structure. Late June 2019, many people in Europe had experienced extreme heat waves. The temperatures in several countries exceeded 40°C (104°F). For example, MERRA-2 shows that the near surface daily maximum temperature of June 28 2019 over Marseille, a city in southern France, reached 41.1 °C (106°F), which is the record breaking temperature in the last 40 years. GES DISC is working together with domain science experts to improve the performance of long time series access, making analysis ready data sets in supporting application researches, such as extreme study. In this presentation, using Europe heat wave as an example, we will show prototype of the in developing service for finding extremes from near 40 years MERRA-2 data at a given location. MERRA-2 data can be accessed from NASA GES DISC(https://disc.gsfc.nasa.gov/ ) by search keyword "MERRA-2".

Shen, Suhung↗