Impactful Implementation of Next-Generation Ocean Data Infrastructure through Scalable Community and Stakeholder Engagement Programs
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
There are many software reliability models which try to predict future performance of software based on data generated by the debugging process. Unfortunately, the models appear to be unable to account for the random nature of the data. If the same code is debugged multiple times and one of the models is used to make predictions, intolerable variance is observed in the resulting reliability predictions. It is believed that data replication can remove this variance in lab type situations and that it is less than scientific to talk about validating a software reliability model without considering replication. It is also believed that data replication may prove to be cost effective in the real world, thus the research centered on verification of the need for replication and on methodologies for generating replicated data in a cost effective manner. The context of the debugging graph was pursued by simulation and experimentation. Simulation was done for the Basic model and the Log-Poisson model. Reasonable values of the parameters were assigned and used to generate simulated data which is then processed by the models in order to determine limitations on their accuracy. These experiments exploit the existing software and program specimens which are in AIR-LAB to measure the performance of reliability models.
An ongoing NASA-funded project has removed a longstanding barrier to accessing NASA data (i.e., accessing archived time-step array data as point-time series) for selected variables of the North American and Global Land Data Assimilation Systems (NLDAS and GLDAS, respectively) and other EOSDIS (Earth Observing System Data Information System) data sets (e.g., precipitation, soil moisture). These time series (data rods) are pre-generated. Data rods Web services are accessible through the CUAHSI Hydrologic Information System (HIS) and the Goddard Earth Sciences Data and Information Services Center (GES DISC) but are not easily discoverable by users of other non-NASA data systems. The Global Earth Observation System of Systems (GEOSS) is a logical mechanism for providing access to the data rods. An ongoing GEOSS Water Services project aims to develop a distributed, global registry of water data, map, and modeling services cataloged using the standards and procedures of the Open Geospatial Consortium and the World Meteorological Organization. The ongoing data rods project has demonstrated the feasibility of leveraging the GEOSS infrastructure to help provide access to time series of model grid information or grids of information over a geographical domain for a particular time interval. A recently-begun, related NASA-funded ACCESS-GEOSS project expands on these prior efforts. Current work is focused on both improving the performance of the generation of on-the-fly (OTF) data rods and the Web interfaces from which users can easily discover, search, and access NASA data.
For over 30 years, NASA has invested in life sciences research both in space and on the ground. Data accessibility is an important tool for researchers, and NASA has committed to preserving this vital resource for ongoing use. The Life Sciences Data Archive’s multi-center collaboration between NASA’s Johnson Space Center, Ames Research Center, and Kennedy Space Center is geared toward preserving unique and high-value data from a wide variety of disciplines, data collection methods, and species within NASA’s Life Sciences Portal (NLSP). The data generated by the Human Research Program (HRP) require a systematic approach to data preservation that accounts for diverse data sources, formats, physical storage requirements, and security and privacy protections. This poster presentation will provide a guide to the repositories where the various types of human, non-human animal, plant, and microbial data NASA generates are archived and tips for navigating these data collections. Topics will include where different types of data, metadata, and biospecimens are archived or preserved, how the federated repositories work together as a data preservation ecosystem, and how researchers can access each repository’s collections.
For over 30 years, NASA has invested in life sciences research both in space and on the ground. Data accessibility is an important tool for researchers, and NASA has committed to preserving this vital resource for ongoing use. The Life Sciences Data Archive’s multi-center collaboration between NASA’s Johnson Space Center, Ames Research Center, and Kennedy Space Center is geared toward preserving unique and high-value data from a wide variety of disciplines, data collection methods, and species within NASA’s Life Sciences Portal (NLSP). The data generated by the Human Research Program (HRP) require a systematic approach to data preservation that accounts for diverse data sources, formats, physical storage requirements, and security and privacy protections. This poster presentation will provide a guide to the repositories where the various types of human, non-human animal, plant, and microbial data NASA generates are archived and tips for navigating these data collections. Topics will include where different types of data, metadata, and biospecimens are archived or preserved, how the federated repositories work together as a data preservation ecosystem, and how researchers can access each repository’s collections.
Comparison performance data of compressor blading without slots or vortex generators
High-resolution historical wind data was developed for the entirety of South America using the innovative Super-Resolution for Renewable Resource Data (sup3r) machine learning framework. The publicly available Sup3rWind South America dataset represents a significant advancement in wind resource data generation, leveraging generative machine learning conditioned on near-surface observations from the Meteorological Assimilation Data Ingest System (MADIS) to efficiently and accurately downscale coarse reanalysis data from the European Centre for Medium-Range Weather Forecasts (ERA5). This approach produces fine-scale, spatially and temporally coherent wind and meteorological fields hundreds of times more computationally efficient than traditional numerical weather modeling methods, enabling access to high-fidelity wind information across both continental and offshore regions. Sup3rWind South America builds on the earlier Sup3rWind Ukraine dataset through improvements in model architecture and outputs conditioned on near-surface observation inputs. As with the Ukraine data release, this dataset includes wind speed, wind direction, temperature, relative humidity, and pressure at a horizontal resolution of ~2 km, representing a 15x spatial enhancement relative to the 31 km ERA5 grid. Wind speed and direction are provided at 5-minute resolution, a 12x temporal refinement compared to the hourly ERA5 data, while temperature, relative humidity, and pressure remain at hourly resolution. The data covers all years from 2005 to 2024. Before downscaling, ERA5 inputs were bias-corrected using long-term monthly means and a limited number of quality-controlled observations to align large-scale statistics with regional conditions. The resulting dataset is the first publicly available high-resolution timeseries wind record that provides full spatial coverage of South America. Model validation demonstrates strong agreement with observations across several statistical metrics, consistent with other state-of-the-art high-resolution wind resource datasets. The potential applications of Sup3rWind South America span renewable energy resource assessment, energy system modeling, and grid resilience analysis. The 20-year record and high spatial and temporal resolution support accurate estimation of long-term energy yield and the economic feasibility of potential wind development sites. Continuous coverage across both continental and offshore regions enables comprehensive site prospecting within exclusive economic zones. The 2 km, 5-minute resolution data provide the spatial and temporal variability required for power system simulation, operational planning, and regional risk assessments.
Computer program generates a magnetic output tape containing time and averaged data values of a specified number of major frames over a specified time interval. A decommutation system is used to acquire the raw data, which is then reformatted and averaged.
The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2021 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilizes the appdirs Python dependency (https://pypi.org/project/appdirs/). An overview of the ElectricityLCI data stores may be found on the README (https://github.com/USEPA/ElectricityLCI/blob/v2.0/README.md#data-store). This submission includes the background data used to generate the 2021 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: python -c "import appdirs; print(appdirs.user_data_dir())"). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2021 model run is also included, which contains the statements at the DEBUG level and above.
The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2020 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilizes the appdirs Python dependency (https://pypi.org/project/appdirs/). An overview of the ElectricityLCI data stores may be found on the README (https://github.com/USEPA/ElectricityLCI/blob/v2.0/README.md#data-store). This submission includes the background data used to generate the 2020 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: python -c "import appdirs; print(appdirs.user_data_dir())"). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2020 model run is also included, which contains the statements at the DEBUG level and above.
The ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0) was used to generate the 2022 electricity baseline: a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data. ElectricityLCI implements a local data store for downloading and accessing public data on an individual's computer. The data store follows the folder definition provided by USEPA's esupy Python package (https://github.com/USEPA/esupy), which utilized the appdirs Python dependency (https://pypi.org/project/appdirs/). This submission includes the background data used to generate the 2022 electricity baseline inventory. Each zip archive stores the source files as found in their data stores. Sub-folders in each of the data stores are archived separately. For example, stewi.zip contains the JSON files, while stewi.facility.zip is the 'facility' sub-folder of stewi data store that stores the parquet files. To reproduce the data store, extract each zip file and drag-and-drop sub-folders in to their appropriate root folders to recreate the data stores, then copy the root folders to your data store folder (as returned by running the following on the command line: `python -c "import appdirs; print(appdirs.user_data_dir())"`). The main five data stores include: 'electricitylci', 'facilitymatcher', 'fedelemflowlist', 'stewi', and 'stewicombo'. The log file generated by the 2022 model run is also included, which contains the statements at the DEBUG level and above.
Urban Air Mobility (UAM) defines an environment for managing operations of vertical takeoff and landing (VTOL) and short takeoff and landing (STOL) vehicles in an urban environment. Within a UAM environment, UAM operators manage fleets of vehicles, relying on Providers of Services for UAM (PSUs) for managing flights in a region of airspace. Flight plan deconfliction is primarily performed by the Discovery and Synchronization Service (DSS), and the Federal Aviation Administration (FAA) maintains control over the UAM space via the FAA-Industry Exchange Protocol (FIDXP). UAM is a federated environment with many different entities owning and operating vehicles, PSUs, and other services. These entities often need to interoperate or access data generated by other organizations. This paper demonstrates the feasibility of using blockchain to facilitate a secure data exchange and storage for this flight information in a UAM environment. In particular, this paper is focused on flight plans and telemetry data. A blockchain network was developed with a set of smart contracts for managing relevant flight data. Hyperledger Fabric was chosen as it is performent, scalable, and allows organizations to reuse existing public key infrastructure (PKI) for identity management. A set of simulated UAM services were also developed. These services propose flight plans and negotiate with other UAM services for airspace access. All interactions between UAM services, as well as vehicle telemetry data, is recorded onto the blockchain. Vehicle telemetry data is generated by a vehicle flight simulation service. This paper successfully demonstrates the feasibility of using blockchain as a secure data exchange and storage mechanism in a UAM environment.
Here, we combine US wind generation—a cheap yet intermittent source of electricity—with the latest geothermal resource estimates to understand the technoeconomic possibilities of pairing enhanced geothermal systems (EGSs) with existing wind farms to develop a hybrid energy system. Using observed generation data from 583 wind farms, generation gaps are quantified and geographically paired with the latest EGS estimates. Results demonstrate that EGS potential within a 1 km 2 footprint can supplement wind generation at 56% of onshore wind farms. Each wind farm can be supplemented by EGSs when 10% of its surface-occupying footprint is available. The cost of EGSs at wind farms is lowest in the western US and southern Texas border and highest in the central US. While further experiments are warranted, a wind + EGS hybrid system offers an opportunity to increase power output from the same land footprint while maximizing the use of existing electrical infrastructure.
In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.
Computer-aided corridor selection techniques, utilizing digitized data bases of socio-economic, census, and cadastral data, and developed for highway corridor routing are considered. Land resource data generated from various remote sensing data sources were successfully merged with the ancillary data files of a corridor selection model and prototype highway corridors were designed using the combined data set. Remote sensing derived information considered useful for highway corridor location, special considerations in geometric correction of remote sensing data to facilitate merging it with ancillary data files, and special interface requirements are briefly discussed.
One of the major tasks of a supercomputer center is managing the massive amount of data generated by application codes. A data flow analysis of the San Diego Supercomputer Center is presented that illustrates the hierarchical data buffering/caching capacity requirements and the associated I/O throughput requirements needed to sustain file service and archival storage. Usage paradigms are examined for both tightly-coupled and loosely-coupled file servers linked to the supercomputer by high-speed networks.
One of the major tasks of a supercomputer center is managing the massive amount of data generated by application codes. A data flow analysis of the San Diego Supercomputer Center is presented that illustrates the hierarchical data buffering/caching capacity requirements and the associated I/O throughput requirements needed to sustain file service and archival storage. Usage paradigms are examined for both tightly-coupled and loosely-coupled file servers linked to the supercomputer by high-speed networks.
A thermal model of a single Ultralight Fabric Reflux Tube (UFRT) was constructed and tested against data for an array of such tubes tested in the NASA-JSC facility. Modifications to the single fin model were necessary to accommodate the change in radiation shape factors due to adjacent tubes. There was good agreement between the test data and data generated for the same cases by the thermal model. The thermal model was also used to generate single and linear array data for the lunar environment (the primary difference between the test and lunar data was due to lunar gravity). The model was also used to optimize the linear spacing of the reflux tubes in an array. The optimal spacing of the tubes was recommended to be about 5 tube diameters based on maximizing the heat transfer per unit mass. The model also showed that the thermal conductivity of the Nextel fabric was the major limitation to the heat transfer. This led to a suggestion that the feasibility of jacketing the Nextel fiber bundles with copper strands be investigated. This jacketing arrangement was estimated to be able to double the thermal conductivity of the fabric at a volume concentration of about 12-14%. Doubling the thermal conductivity of the fabric would double the amount of heat transferred at the same steam saturation temperature.