Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Big lunar data visualization and analysis

NASA's earth and planetary spacecraft return large amounts of remote sensing data, such as imagery and raw science measurements, in support of remarkable research. Not only does the data lead to new scientific discoveries about our planet and the solar system, it provides a wealth of information to educate, inspire, and engage the public at large. To leverage this rich data for mission planning, scientific research, public outreach and education, it is essential to make it accessible and understandable, analyzable, all while appealing to their interests. This presentation will highlight web-based capabilities that showcase NASA's large volume of lunar data collected from past and current Moon missions. It is particularly relevant as the new Administration has more plans for the Moon. We will illustrate big data visualization and analysis in easy-touse and interactive mediums for diverse use.

Malhotra, Shan↗

Return of the Big Glitcher: NICER Timing and Glitches of PSR J0537-6910

PSR J0537−6910, also known as the Big Glitcher, is the most prolific glitching pulsar known, and its spin-induced pulsations are only detectable in X-ray. We present results from analysis of 2.7 yr of NICER timing observations, from 2017 August to 2020 April. We obtain a rotation phase-connected timing model for the entire time span, which overlaps with the third observing run of LIGO/Virgo, thus enabling the most sensitive gravitational wave searches of this potentially strong gravitational wave-emitting pulsar. We find that the short-term braking index between glitches decreases towards a value of 7 or lower at longer times since the preceding glitch. By combining NICER and RXTE data, we measure a long-term braking index n = −1.25 ± 0.01. Our analysis reveals eight new glitches, the first detected since 2011, near the end of RXTE, with a total NICER and RXTE glitch activity of 8.88 × 10−7 yr−1. The new glitches follow the seemingly unique time-to-next-glitch–glitch-size correlation established previously using RXTE data, with a slope of 5 d μHz−1. For one glitch around which NICER observes 2 d on either side, we search for but do not see clear evidence of spectral nor pulse profile changes that may be associated with the glitch.

Wynn C G Ho↗

Large-Scale High-Resolution Coastal Mangrove Forests Mapping Across West Africa With Machine Learning Ensemble and Satellite Big Data

Coastal mangrove forests provide important ecosystem goods and services, including carbon sequestration, biodiversity conservation, and hazard mitigation. However, they are being destroyed at an alarming rate by human activities. To characterize mangrove forest changes, evaluate their impacts, and support relevant protection and restoration decision making, accurate and up-to-date mangrove extent mapping at large spatial scales is essential. Available large-scale mangrove extent data products use a single machine learning method commonly with 30 m Landsat imagery, and significant inconsistencies remain among these data products. With huge amounts of satellite data involved and the heterogeneity of land surface characteristics across large geographic areas, finding the most suitable method for large-scale high-resolution mangrove mapping is a challenge. The objective of this study is to evaluate the performance of a machine learning ensemble for mangrove forest mapping at 20 m spatial resolution across West Africa using Sentinel-2 (optical) and Sentinel-1 (radar) imagery. The machine learning ensemble integrates three commonly used machine learning methods in land cover and land use mapping, including Random Forest (RF), Gradient Boosting Machine (GBM), and Neural Network (NN). The cloud-based big geospatial data processing platform Google Earth Engine (GEE) was used for pre-processing Sentinel-2 and Sentinel-1 data. Extensive validation has demonstrated that the machine learning ensemble can generate mangrove extent maps at high accuracies for all study regions in West Africa (92%–99% Producer’s Accuracy, 98%–100% User’s Accuracy, 95%–99% Overall Accuracy). This is the first-time that mangrove extent has been mapped at a 20 m spatial resolution across West Africa. The machine learning ensemble has the potential to be applied to other regions of the world and is therefore capable of producing high-resolution mangrove extent maps at global scales periodically.

coastal environment↗

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) was conceived to address the existing ground testing data management of the NASA Ames arc jet facilities (e.g., manually entered Excel files and USB drive data transfers). These data management practices were seen as a choke point for future thermal protection system (TPS) development as they limit statistical tracking, resolution of diagnostics, coordination between video/time series, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) is a facility data management application developed for the NASA Ames arc jet facilities. The current decentralized data management practices limit statistical tracking, synchronization between video/time series, search capability, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Big-data Efficient Automated Science Transfer (BEAST): an open-source software architecture for arc jet data management, modeling, and automation

Big-data Efficient and Automated Science Transfer (BEAST) was conceived to address the existing ground testing data management of the NASA Ames arc jet facilities (e.g., manually entered Excel files and USB drive data transfers). These data management practices were seen as a choke point for future thermal protection system (TPS) development as they limit statistical tracking, resolution of diagnostics, coordination between video/time series, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Transforming NASA Earth Science Data Systems: A Journey from Big Earth Data Initiative (BEDI) to Open-Source Science Initiative (OSSI)

NASA's Earth Science Data and Information Systems (ESDIS) have undergone a significant evolution, particularly with the introduction of the Big Earth Data Initiative (BEDI) and the Open-Source Science Initiative (OSSI). In this talk, I will provide an overview of NASA's Earth Science Data and Information systems, highlighting key components such as EOSDIS, ESDIS, and ESDS. Moving forward, I will delve into the BEDI initiative, discussing its objectives, key players, and lessons learned. The second part of the talk will cover the OSSI initiative, exploring its objectives, strategy, and innovative solutions. Throughout the presentation, I will provide insights into the requests, strategies, and solutions behind both BEDI and OSSI. By the end, you will gain a comprehensive understanding of how NASA's Earth Science Data Systems have evolved over the years and witness the organization's commitment to advancing an open-source and collaborative approach to data science. Join me for an enlightening exploration into the future of Earth science data and the pivotal role played by NASA in shaping this transformative landscape.

Jennifer Wei↗

Big Data Challenges at CCMC

Like other research centers, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) at NASA Goddard Space Flight Center (GSFC) is also experiencing the big data challenges. CCMC hosts over 80 space weather models for Runs On Request (ROR), Continuous Runs and Instant Runs simulation services for the research community. In addition, CCMC has started to support simulation output onboarding in response to the Open Science initiative. Overall, we have accumulated over petabytes of simulation output data and are rapidly growing. In this presentation, we will discuss our data and storage challenges. We will present our attempts to address our challenges and any associated lessons learned. CCMC uses Apache Airflow to ensure data transfer is consistent. We will give a brief overview on how we leverage Apache Airflow to enhance our environment.

space weather↗

Big Data Challenges at the Community Coordinated Modeling Center (CCMC)

Like other research centers, the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) at NASA Goddard Space Flight Center (GSFC) is also experiencing the big data challenges. CCMC hosts over 80 space weather models for Runs On Request (ROR), Continuous Runs and Instant Runs simulation services for the research community. In addition, CCMC has started to support simulation output onboarding in response to the Open Science initiative. Overall, we have accumulated over petabytes of simulation output data and are rapidly growing. In this presentation, we will discuss our data and storage challenges. We will present our attempts to address our challenges and any associated lessons learned. CCMC uses Apache Airflow to ensure data transfer is consistent. We will give a brief overview on how we leverage Apache Airflow to enhance our environment.

space weather↗

Taming Big Data Variety in the Earth Observing System Data and Information System

Although the volume of the remote sensing data managed by the Earth Observing System Data and Information System is formidable, an oft-overlooked challenge is the variety of data. The diversity in satellite instruments, science disciplines and user communities drives cost as much or more as the data volume. Several strategies are used to tame this variety: data allocation to distinct centers of expertise; a common metadata repository for discovery, data format standards and conventions; and services that further abstract the variations in data.

Information Systems↗

An Integrated Gate Turnaround Management Concept Leveraging Big Data/Analytics for NAS Performance Improvements

The Integrated Gate Turnaround Management (IGTM) concept was developed to improve the gate turnaround performance at the airport by leveraging relevant historical data to support optimization of airport gate operations, which include: taxi to the gate, gate services, push back, taxi to the runway, and takeoff, based on available resources, constraints, and uncertainties. By analyzing events of gate operations, primary performance dependent attributes of these events were identified for the historical data analysis such that performance models can be developed based on uncertainties to support descriptive, predictive, and prescriptive functions. A system architecture was developed to examine system requirements in support of such a concept. An IGTM prototype was developed to demonstrate the concept using a distributed network and collaborative decision tools for stakeholders to meet on time pushback performance under uncertainties.

Big Data and Net-enabled ATM↗

NASA's Big Earth Data Initiative Accomplishments

The goal of NASA's effort for BEDI is to improve the usability, discoverability, and accessibility of Earth Observation data in support of societal benefit areas. Accomplishments: In support of BEDI goals, datasets have been entered into Common Metadata Repository(CMR), made available via the Open-source Project for a Network Data Access Protocol (OPeNDAP), have a Digital Object Identifier (DOI) registered for the dataset, and to support fast visualization many layers have been added in to the Global Imagery Browse Services (GIBS).

CMR↗

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage↗

Investigating Access Performance of Long Time Series with Restructured Big Model Data

Data sets generated by models are substantially increasing in volume, due to increases in spatial and temporal resolution, and the number of output variables. Many users wish to download subsetted data in preferred data formats and structures, as it is getting increasingly difficult to handle the original full-size data files. For example, application research users such as those involved with wind or solar energy, or extreme weather events are likely only interested in daily or hourly model data at a single point (or for a small area) for a long time period, and prefer to have the data downloaded in a single file. With native model file structures, such as hourly data from NASA Modern-Era Retrospective analysis for Research and Applications Version-2 (MERRA-2), it may take over 10 hours for the extraction of parameters-of-interest at a single point for 30 years. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is exploring methods to address this particular user need. One approach is to create value-added data by reconstructing the data files. Taking MERRA-2 data as an example, we have tested converting hourly data from one-day-per-file into different data cubes, such as one-month, or one-year. Performance is compared for reading local data files and accessing data through interoperable services, such as OPeNDAP. Results show that, compared to the original file structure, the new data cubes offer much better performance for accessing long time series. We have noticed that performance is associated with the cube size and structure, the compression method, and how the data are accessed. An optimized data cube structure will not only improve data access, but also may enable better online analysis services

reanalysis↗

Advanced Analytics and Big Earth Data

NASA's Earth Science Data Systems process, archive and distribute petabytes of Earth Observation data to a variety of end users. These end users will face dramatically increased data size in the near future, bringing about new challenges and opportunities in analyzing those data. One area of particular ferment currently is Machine Learning. Many Machine Learning methods are black boxes, limiting direct insight into the data's properties. However, they can be used for a variety of data enhancement purposes, such as parameter retrieval, data fusion and image classification and segmentation. The Earth Observing System Data and Information System is also evolving to host large data volumes in the cloud, enabling data proximal analysis. As part of this effort, an Analytics framework is being developed to support and enhance user analysis of the data. By using standards based services in the framework, diverse user communities can be served, while also allowing inter-system collaboration in the analysis process.

Cloud Computing↗