Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “cloud storage”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Cloud Optimized Data Formats

Cloud computing offers the promise of being able to analyze Big Data earth Observations at scale, by allowing scientists to deploy many nodes at once to analyze the data. However, in order to take full advantage of cloud scalability, it is often necessary to reorganize and reformat the data to enable fine-grained, parallel access to the data in Web Object Storage. NASA recently conducted a study of several formats that are optimized for analysis in the cloud: Parquet, zarr, HDF (Hierarchical Data Format) in the Cloud, and Cloud-Optimized GeoTIFF (Tagged Image File Format). They were compared against non-cloud-optimized formats, netCDF (network Common Data Form) and GeoTIFF, with criteria based both on stewardship and analysis performance.

Christopher Lynnes↗

Intelligent Observation Strategies for Geosynchronous Remote Sensing for Natural Hazards

Geosynchronous satellites offer a unique perspective for monitoring environmental factors important to understanding natural hazards and supporting the disasters management life cycle, namely forecast, detection, response, recovery and mitigation. In the NASA decadal survey for Earth science, the GEO-CAPE mission was proposed to address coastal and air pollution events in geosynchronous orbit, complementing similar initiatives in Asia by the South Koreans and by ESA in Europe, thereby covering the northern hemisphere. In addition to analyzing the challenges of identifying instrument capabilities to meet the science requirements, and the implications of hosting the instrument payloads on commercial geosynchronous satellites, the GEO-CAPE mission design team conducted a short study to explore strategies to optimize the science return for the coastal imaging instrument. The study focused on intelligent scheduling strategies that took into account cloud avoidance techniques as well as onboard processing methods to reduce the data storage and transmission loads. This paper expands the findings of that study to address the use of intelligent scheduling techniques and near-real time data product acquisition of both the coastal water and air pollution events. The topics include the use of onboard processing to refine and execute schedules, to detect cloud contamination in observations, and to reduce data handling operations. Analysis of state of the art flight computing capabilities will be presented, along with an assessment of cloud detection algorithms and their performance characteristics. Tools developed to illustrate operational concepts will be described, including their applicability to environmental monitoring domains with an eye to the future. In the geostationary configuration, the payload becomes a networked thing with enough connectivity to exchange data seamlessly with users. This allows the full field of view to be sensed at very high rate under the control of ground infrastructure, resulting in improved efficiencies, accuracy and science benefits. Hence a remote sensing payload and its data may become one of millions of connected objects in the emerging Internet of Things (IoT), and be as easily accessible by a users smart phone as any other smart appliance.

Characterizing Wildfires in Western US.: A Cloud-based Case Study for Interdisciplinary Research using NASA Resources

This presentation will demonstrate a case study of interdisciplinary research done in the Amazon Web Services (AWS) cloud platform, in addition to in the local machine. We conduct data analysis next to data by leveraging various cloud-based data in NASA Earthdata Cloud, which are distributed by different missions/NASA Distributed Active Archive Centers (DAACs), and cloud computing resources at NASA. For instance, we directly access multiple datasets stored in the AWS Simple Storage Service (S3) buckets using a Python Jupyter notebook through a JupyterHub interface hosted in AWS (without having to download data), and conduct data analysis next to data in the cloud. We will also show how to share the research results following Open Source policy. This case study characterizes the change in wildfire events in the western United States during the past 20 years. In particular, we focus on the wildfires in California in 2021, one of the most severe wildfire years occurring in the most recent 20 years in California. We will analyze the possible causes of wildfires, such as drought conditions and climate variability, and examine the impacts of wildfires on air quality and atmospheric composition, and on land cover. We will examine the data distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), including aerosols and meteorological data from the NASA Modern-Era Retrospective analysis for Research and Applications version 2 (MERRA-2), precipitation from the Global Precipitation Measurement (GPM) and Global Precipitation Climate Project (GPCP), and aerosol index from Ozone Monitoring Instrument (OMI). We also utilize the data distributed by the Physical Oceanography (PO) DAAC, such as Sea Surface Temperature (SST) data from the Group for High Resolution Sea Surface Temperature (GHRSST), and the data distributed by Land Processes (LP) DAAC, such as Normalized Difference Vegetation Index (NDVI).

Xiaohua Pan↗

The Future of NASA Earth Science in the Commercial Cloud: Challenges and Opportunities

NASA produces a large volume and variety of data products that are used every day to support research, decision making, and education. The widespread use of NASA’s Earth Science data is enabled by NASA’s Earth Science Data System (ESDS) program, which oversees the archiving and distribution of these data and invests in the development of new data systems and tools. However, NASA’s current approach to Earth Science data distribution — based on distributed institutional archives with individual on-premises high-performance computing capabilities — faces some significant challenges, including massive increases in data volume from upcoming missions, a greater need for transdisciplinary science that synthesizes many different kinds of observations, and a push to make science more open, inclusive, and accessible. To address these challenges, NASA is aggressively migrating its Earth Science data and related tools and services into the commercial cloud. Migration of data into the commercial cloud can significantly improve NASA’s existing data system capabilities by (1) providing more flexible options for storage and compute (including rapid, as-needed access to state-of-the-art capabilities); (2) by centralizing and standardizing data access, which gives all of NASA’s institutional data centers access to all of each other’s datasets; and (3) by facilitating “analysis-in-place”, whereby users can bring their own computational workflows and tools to the data rather than having to maintain their own copies of NASA datasets. However, migration to the commercial cloud also poses some significant challenges, including (1) managing costs under a “pay-as-you-go” model; (2) incompatibility with existing tools and data formats with object-based storage and network access; (3) vendor lock-in; (4) challenges with data access for workflows that mix on-premise and cloud computing; and (5) standardization for highly diverse data as is present in NASA’s data archive. I conclude with two examples of recent NASA activities showcasing capabilities enabled by the commercial cloud: An interactive analysis and development platform for analyzing airborne imaging spectroscopy data, and a new collection of tools and services for data discovery, analysis, publication, and data-driven storytelling (Visualization, Exploration, and Data Analysis, VEDA).

Alexey N Shiklomanov↗

Cultivating an Emergent Earth Observation Analytics Ecosystem in the Cloud

A diverse set of data analytics systems for Earth Observations are sprouting up in the Earth Science community, with a wealth of processing algorithms and analysis methods. There is a similar wealth of data resources available via myriad data providers and clearinghouses, including large institutional systems like the Earth Observing System Data and Information System, Comprehensive Large Scale Array-data Stewardship System, and Federated Earth Observation Missions gateway. With Earth system science driving a need to work with more datasets together, and the community developing more analysis tools (some of them dataset-specific), how can we develop analysis workflows that incorporate far-flung datasets and leverage analysis resources from multiple organizations? Cloud computing points the way toward a solution in two different respects. Firstly, the access to and abstraction of virtually unlimited storage and computing power provides an environment that enables more straightforward means of pulling datasets and analysis resources together. Just as importantly, however, cloud computing serves as an example of an "ecosystem" of interoperating services, since the essence of cloud computing is the presentation of all resources as a service, from hardware to infrastructure to platform to software. This enables the combination of off-the-shelf, diverse services to construct entire systems that emerge out of an equally diverse community of architects and developers. This approach can be similarly applied to the data and analysis resources in the Earth Observation community. By exposing these resources via well understood services, and consuming resources in the same way, different organizations can construct bespoke analysis workflows and systems for their own purposes. The key leap the community needs to make is to develop analysis systems in components that interact with other components via services. The result would be a rich ecosystem of analytics components that can be combined to analyze datasets at scale and in conjunction with other datasets from other sources.

chaos↗

High Resolution Nature Runs and the Big Data Challenge

NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs

big data analysis↗

Understanding Uncertainty of Snow Radiative Transfer Modeling Within a Mixed Deciduous and Evergreen Forest

Satellite-based passive microwave observations provide the best available continuous observational estimates of global snow water storage due to their broad geographic footprint and low sensitivity to clouds and precipitation. However, these observations are subject to substantial uncertainty due to the complex radiative properties of snow and from interference in forested areas. Physical radiative transfer models can be leveraged to improve the fidelity of these observations and as a data-assimilation tool. In this article, the Dense Media Radiative Transfer model with Multiple Layers (DMRT-ML) is used to simulate snow brightness temperatures from data collected from snow pits excavated during a two-day-long field study performed a temperate forest in the Northeast United States. The simulations are evaluated against surface-based radiometer observations collected at the snow pits. The DMRT-ML is configured with varying complexity to determine the snowpack characteristics most essential toward simulating brightness temperature within a temperate forest with complicated snow stratigraphy. In general, the single-layer configurations were not sufficiently complex to accurately simulate snow brightness temperature without significant tuning. The most accurate simulation was a two-layer configuration with a prescribed ice layer separating the snow layers. This simulation had a root-mean-square error <;15 K for the 37-GHz frequency. More complicated snowpack stratigraphy configurations did not substantively improve the results over the two-layer model configuration. The DMRT-ML was also used to examine differences between redundant datasets of density and grain size. It was determined that similar snow data collection and radiative transfer model configuration techniques are critical to ensure cross-study comparability.

Theodore Letcher↗

High Performance Access to Archival Data Stored in HDF4 and HDF5 on Cloud Object Stores Without Reformatting the Files

Cloud computing offers numerous advantages for users of extensive Earth science data collections. These benefits encompass direct online access to data files and granules from any location, scalable access supporting parallel computing workflows, and flexible computing tools enabling innovative experimentation with processing techniques. However, older archival file formats designed for distinct computing systems hinder efficient access to decade-long time-series data when compared to data stored in modern cloud-optimized formats like Web Object Stores (WOS), exemplified by Amazon Web Services’ Simple Storage Service (S3). We describe DMR++ (Dataset Metadata Response plus plus), a technology facilitating efficient access to HDF5 (Hierarchical Data Format, version 5) and HDF4 files stored on WOS systems without requiring data reformatting. DMR++ achieves performance comparable to technologies like Zarr while preserving the original file structure, a substantial benefit considering the vast quantity of archival files held by organizations such as NASA. Moreover, DMR++ typically outperforms cloud-optimized versions of HDF5. Essentially an XML (Extensible Markup Language) document usually stored alongside the described data, DMR++ can also be generated on-the-fly but is generally created during data staging to the WOS. Archival files that use HDF4/5 often store large arrays of numerical data. The data in these files is often compressed, typically reducing their size by a factor of four or more. To achieve efficient access to portions of those arrays, they are 'chunked' into smaller sub-arrays, each individually compressed. The chunk size is a compromise, where spinning disks can efficiently access data in smaller chunks while S3 favors larger chunks. A simple optimization of aggregating smaller chunks that are stored adjacently, transferring them in a single access and then individually decompressing them will improve performance. NASA data pose an additional challenge: special Application Programmer Interface (API) libraries are often needed to compute some variables. These libraries are incompatible with WOS environments. Our solution involves storing computed values in the DMR++ document or a companion file, making them accessible like other variables and eliminating the need for specialized APIs. We outline specific optimizations for both satellite grid and swath data stored in HDF4-EOS2 (Earth Observing System).

James Gallagher↗

Processing NASA Earth Science Data on Nebula Cloud

Three applications were successfully migrated to Nebula, including S4PM, AIRS L1/L2 algorithms, and Giovanni MAPSS. Nebula has some advantages compared with local machines (e.g. performance, cost, scalability, bundling, etc.). Nebula still faces some challenges (e.g. stability, object storage, networking, etc.). Migrating applications to Nebula is feasible but time consuming. Lessons learned from our Nebula experience will benefit future Cloud Computing efforts at GES DISC.

Chen, Aijun↗

Huygens Probe Relay Data Subsystem Anomaly and Recovery

European Space Agency Mission is designed to study the atmosphere and surface of Saturn's largest satellite, Titan carried by the Cassini spacecraft which provides: a) Power for support equipment; b) S-band antenna system; and c) Data storage and playback. Instruments/investigations include: 1) Aerosol Collector Pyrolyzer (ACP). Study of clouds and aerosols in the Titan atmosphere. 2) Descent Imager and Spectral Radiometer (DISR). Aerosol and cloud optical properties and spectroscopy measurements of Titan's atmosphere and surface. 3) Doppler Wind Experiment (DWE). Study of winds from their effect on the Probe during Titan descent. 4) Gas Chromatograph and Mass Spectrometer (GCMS). Chemical composition of gases and aerosols in Titan's atmosphere. 5) Huygens Atmospheric Structure Instrument (HASI). In-situ study of Titan atmospheric physical and electrical properties. 6) Surface Science Package (SSP). Physical properties of Titan's surface and related atmospheric properties.

Huygens↗

Atmosphere-biosphere exchange of CO2 and O3 in the Central Amazon Forest

An eddy correlation measurement of O3 deposition and CO2 exchange at a level 10 m above the canopy of the Amazon forest, conducted as part of the NASA/INPE ABLE2b mission during the wet season of 1987, is presented. It was found that the ecosystem exchange of CO2 undergoes a well-defined diurnal variation driven by the input of solar radiation. A curvilinear relationship was found between solar irradiance and uptake of CO2, with net CO2 uptake at a given solar irradiance equal to rates observed over forests in other climate zones. The carbon balance of the system appeared sensitive to cloud cover on the time scale of the experiment, suggesting that global carbon storage might be affected by changes in insolation associated with tropical climate fluctuations. The forest was found to be an efficient sink for O3 during the day, and evidence indicates that the Amazon forests could be a significant sink for global ozone during the nine-month wet period and that deforestation could dramatically alter O3 budgets.

Fan, Song-Miao↗

Cloud-Based Orchestration of a Model-Based Power and Data Analysis Toolchain

The proposed Europa Mission concept contains many engineering and scientific instruments that consume varying amounts of power and produce varying amounts of data throughout the mission. System-level power and data usage must be well understood and analyzed to verify design requirements. Numerous cross-disciplinary tools and analysis models are used to simulate the system-level spacecraft power and data behavior. This paper addresses the problem of orchestrating a consistent set of models, tools, and data in a unified analysis toolchain when ownership is distributed among numerous domain experts. An analysis and simulation environment was developed as a way to manage the complexity of the power and data analysis toolchain and to reduce the simulation turnaround time. A system model data repository is used as the trusted store of high-level inputs and results while other remote servers are used for archival of larger data sets and for analysis tool execution. Simulation data passes through numerous domain-specific analysis tools and end-to-end simulation execution is enabled through a web-based tool. The use of a cloud-based service facilitates coordination among distributed developers and enables scalable computation and storage needs, and ensures a consistent execution environment. Configuration management is emphasized to maintain traceability between current and historical simulation runs and their corresponding versions of models, tools and data.

Post, Ethan↗

Designing the User Experience for Earth Observation Data Services in the Cloud

NASA's Earth Observation (EO) inventory is projected to grow by an order of magnitude over the next 5-6 years. The current mode for working with EO data of downloading the data to a local machine (laptop, desktop or server) will be difficult to sustain for these upcoming volumes. Therefore, NASA is in the process of developing a capability to host large volume data in commercial cloud, with an eye toward encouraging data analysis in the cloud. However, in order for the user community to take advantage of this new mode of data interaction, the user experience must be redesigned. Cloud-hosted data brings new challenges, such as managing the costs of data egress and working with data in Web Object Storage instead of a Posix filesystem. However, it also brings new opportunities. Scaling data transformation processes may permit more synchronous data services with near-immediate response vs. cumbersome ordering systems with latencies of hours or days. Data co-location in the cloud can facilitate data integration and fusion. Highly scalable filesystems and databases in the cloud support data reorganization to facilitate analysis at scale. In the course of NASA's reimagining of the User Experience for EO data usage relies on end user input gathered through surveys, workshops and meetings (such as this). At the same time, we have embarked on a course of pedagogy and capacity building to help the user community evolve to cloud-based analysis.

Lynnes, Christopher↗

The evolution of organic mantles on interstellar grains

By laboratory simulation of the chemical processes on dust grains it was investigated how solid organic materials can be produced in the interstellar medium. The ice mantles that accrete on grains in molecular clouds, consisting primarily of H2O, CO, H2CO, NH3, and O2, are irradiated by the internal UV field, resulting in the storage of radicals upon photodissociation of the original molecules. Transient heating events lead to the production of oxygen-rich organic species by recombination reactions. The experiments indicated that in this way the observed amount of organic material can be produced if a grain passes a few times through a molecular cloud during its life. After the destruction of the cloud the grains enter a more diffuse medium. Here they are subjected to the interstellar UV field as well as to collisions with atomic hydrogen. Experiments show that the intense photoprocessing results in the removal of small species like H2O and NH3 as well as in carbonization of the organic molecules. Contrary to this, the atomic H flux will maintain a certain hydrogen level in the mantle. These processes likely convert the original, oxygen-rich organics into an unsaturated hydrocarbon type material such as that observed towards IRS 7 and in Comet Halley grains.

Schutte, Willem A.↗

Thermal buffering of receivers for parabolic dish solar thermal power plants

A parabolic dish solar thermal power plant comprises a field of parabolic dish power modules where each module is composed of a two-axis tracking parabolic dish concentrator which reflects sunlight (insolation) into the aperture of a cavity receiver at the focal point of the dish. The heat generated by the solar flux entering the receiver is removed by a heat transfer fluid. In the dish power module, this heat is used to drive a small heat engine/generator assembly which is directly connected to the cavity receiver at the focal point. A computer analysis is performed to assess the thermal buffering characteristics of receivers containing sensible and latent heat thermal energy storage. Parametric variations of the thermal inertia of the integrated receiver-buffer storage systems coupled with different fluid flow rate control strategies are carried out to delineate the effect of buffer storage, the transient response of the receiver-storage systems and corresponding fluid outlet temperature. It is concluded that addition of phase change buffer storage will substantially improve system operational characteristics during periods of rapidly fluctuating insolation due to cloud passage.

Manvi, R.↗

Propulsion Research at the Propulsion Research Center of the NASA Marshall Space Flight Center

The Propulsion Research Center of the NASA Marshall Space Flight Center is engaged in research activities aimed at providing the bases for fundamental advancement of a range of space propulsion technologies. There are four broad research themes. Advanced chemical propulsion studies focus on the detailed chemistry and transport processes for high-pressure combustion, and on the understanding and control of combustion stability. New high-energy propellant research ranges from theoretical prediction of new propellant properties through experimental characterization propellant performance, material interactions, aging properties, and ignition behavior. Another research area involves advanced nuclear electric propulsion with new robust and lightweight materials and with designs for advanced fuels. Nuclear electric propulsion systems are characterized using simulated nuclear systems, where the non-nuclear power source has the form and power input of a nuclear reactor. This permits detailed testing of nuclear propulsion systems in a non-nuclear environment. In-space propulsion research is focused primarily on high power plasma thruster work. New methods for achieving higher thrust in these devices are being studied theoretically and experimentally. Solar thermal propulsion research is also underway for in-space applications. The fourth of these research areas is advanced energetics. Specific research here includes the containment of ion clouds for extended periods. This is aimed at proving the concept of antimatter trapping and storage for use ultimately in propulsion applications. Another activity in this involves research into lightweight magnetic technology for space propulsion applications.

Blevins, John↗

Geonex: A NASA-NOAA Collaboration for Producing Land Surface Products from Geostationary Sensors Using Cloud Computing

The latest generation of geostationary satellites carry sensors such as the Advanced Baseline Imager (GOES-16/17) and the Advanced Himawari Imager (Himawari-8/9) that closely mimic the spatial and spectral characteristics of MODIS and VIIRS, useful for monitoring land surface conditions. The NASA Earth Exchange (NEX) team at Ames Research Center has embarked on a collaborative effort among scientists from NASA and NOAA exploring the feasibility of producing operational land surface products similar to those from MODIS/VIIRS. The team built a processing pipeline called GEONEX that is capable of converting raw geostationary data into routine products of Fires, surface reflectances, vegetation indices, LAI/FPAR, ET and GPP/NPP using algorithms adapted from both NASA/EOS and NOAA/GOES-R programs. The GEONEX pipeline has been deployed on Amazon Web Services cloud platform and it currently leverages near-realtime geostationary data hosted in AWS public datasets under a NOAA-AWS agreement.Initial analyses of various products from ABI/AHI sensors suggest that they are comparable to those from MODIS in representing the spatio-temporal dynamics of land conditions. Cloud computing offers a variety of options for deploying the GEONEX pipeline including choice CPUs, storage media, and automation. We estimate the cost of deploying GEONEX to be $400 - 750 a month for processing data (every 30 minutes) and producing products over the conterminous US. For products such as Fire, latency can be as little as 10 minutes from the time of data acquisition.

Geostationary↗

Benchmark Comparison of Cloud Analytics Methods Applied to Earth Observations

Earth Observation data are a vital resource for studying long term changes, but the large data volumes can be challenging to analyze. Time series analysis in particular is hampered by the typical thin-time-slice file organization. We examine several potential solutions inspired in large part by the data-parallel methods that have arisen with cloud computing. These solutions include various combinations of data re-organization, spatial indexing, distributed storage and pre-computation that we term "Analytics Optimized Data Stores" (AODS). We find that even simple solutions (such as a data cube) produce more than an order of magnitude improvement; the best provide two to three orders of magnitude improvement. The most performant solutions have tradeoffs in terms of generality or storage footprint, but may nonetheless be useful components in data analytics frameworks where performance is critical.

parallel processing (computers)↗