Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

SWIPE: Spectral Water Inversion Processor and Emulator

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will be discussing the progress made developing SWIPE: Spectral Water Inversion Processor and Emulator. SWIPE is a platform for advanced modeling of coastal and inland aquatic habitats. The goal is create a comprehensive and cohesive system to leverage recent advancements in computation and machine learning to develop a synthetic training ground for sensitivity studies and algorithm development. The four principal facets of SWIPE include: 1. Advanced two-layer coated sphere bio-optical modeling and GPU radiative transfer modeling, 2. Big Data involving massive synthetic spectral libraries of optical properties of various global aquatic particles, surface reflectance, and top-of-atmosphere reflectance, all at hyperspectral resolution leveraging high-end computing systems at NASA Ames Research Center, 3. Deep Learning for algorithm development for water quality inversion of concentrations of common biogeophysical variables as well as optics, full uncertainty characterization by water type, and forward emulation, and lastly, 4. Image Processing for application of developed retrieval algorithms for both hyperspectral and multispectral sensors with experimental corrections for global adjacency, noise, sunglint, and benthic reflectance. This presentation will demonstrate the Equivalent Algal Populations (EAP) two-layer coated sphere scattering model which has been used develop spectral libraries of hyperspectral inherent optical properties of roughly 80 species of phytoplankton, covering 15 different classes and nine taxonomic functional types. The EAP model was also used to derive spectral properties of 10 different non-algal particle functional types. Examples of how the SMART-G (Speed-up Monte-carlo Advanced Radiative Transfer using GPU) radiative transfer code is used to model optically complex aquatic signals will be presented and discussed in the context of creating a massive synthetic database which can leverage the full power of next generation machine learning techniques and high end computing for water quality inversion. We will discuss our active investigation in things like appropriate model architectures, dimensionality reduction techniques such as PCA and autoencoders, uncertainty quantification and abstaining, and which variables actually benefit most from hyperspectral information versus multispectral resolution. We are also curious about questions relating to cost/benefit analysis in terms of computation resources, neural network complexity, and data volumes. Answers to these questions will hopefully elaborate on cost efficiency for potential future sensor design considerations.

SWIPE↗

Machine Intelligence for Radiation Science: Summary of the Radiation Research Society 67th Annual Meeting Symposium

The era of high-throughput techniques created big data in the medical field and research disciplines. Machine intelligence (MI) approaches can overcome critical limitations on how those large-scale data sets are processed, analyzed, and interpreted. The 67 th Annual Meeting of the Radiation Research Society featured a symposium on MI approaches to highlight recent advancements in the radiation sciences and their clinical applications. This article summarizes three of those presentations regarding recent developments for metadata processing and ontological formalization, data mining for radiation outcomes in pediatric oncology, and imaging in lung cancer.

radiation↗

Coevolution of Machine Learning and Process-Based Modelling to Revolutionize Earth and Environmental Sciences: A Perspective

Machine learning (ML) applications in Earth and environmental sciences (EES) have gained incredible momentum in recent years. However, these ML applications have largely evolved in ‘isolation’ from the mechanistic, process-based modelling (PBM) paradigms, which have historically been the cornerstone of scientific discovery and policy support. In this perspective, we assert that the cultural barriers between the ML and PBM communities limit the potential of ML, and even its ‘hybridization’ with PBM, for EES applications. Fundamental, but often ignored, differences between ML and PBM are discussed as well as their strengths and weaknesses in light of three overarching modelling objectives in EES, (1) nowcasting and prediction, (2) scenario analysis, and (3) diagnostic learning. The paper ponders over a ‘coevolutionary’ approach to model building, shifting away from a borrowing to a co-creation culture, to develop a generation of models that leverage the unique strengths of ML such as scalability to big data and high-dimensional mapping, while remaining faithful to process-based knowledge base and principles of model explainability and interpretability, and therefore, falsifiability.

Saman Razavi↗

Development of an Airspace Simulation and Modeling Tool for Enhanced Spectrum Management

The emergence of new aerial vehicles into the National Airspace System creates an increased demand for aeronautical communications to support aviation operations. However, the issue of spectrum scarcity remains an ever-present concern, and the growing demand cannot be supported using existing spectrum allocation strategies. As a result, a new spectrum management approach is required, and the National Aeronautics and Space Administration (NASA) is investigating advanced concepts to modernize the management and use of aviation spectrum by leveraging the latest advancements in wireless communications, big data and machine learning. This research proposes an autonomous spectrum allocation concept, which allocates communications resources, such as spectrum and power, based on the predicted communications and air traffic demands throughout the airspace, as opposed to the use of fixed allocations as is done today. This approach will result in improved spectrum utilization efficiency and enhanced airspace capacity. The autonomous spectrum allocation concept decomposes into three research areas: demand prediction, resource allocation, and use case evaluation. As part of the use case evaluation effort, a modeling and simulation capability is currently under development. This simulation capability includes the implementation of various features, including visualization of both live or virtually-generated airspace traffic, simulation scenario development, simulation management with data collection, and flight plan creation with corresponding trajectory generation. This modeling and simulation capability will continue to evolve as new and advanced airspace applications are introduced into existing and emerging operational environments.

Eric J. Knoblock↗

Scientific Content Curation in an Open Science Era

Today’s open science environment, in combination with the Big Data era, means more scientific data, software, tools, documentation, publications and other resources are available than ever. The promise of the open science era is that scientists will spend less time reinventing the wheel and more time doing actionable research. Yet navigating this vast and complex information landscape can feel overwhelming to scientists trying to get their bearings. In this presentation, we define and discuss the importance of scientific content curation for enhancing discovery and use of scientific data and information. We also share two examples of scientific content curation in action: the Catalog of Archived Suborbital Earth Science Investigations (CASEI) and the Science Discovery Engine (SDE).

Kaylin Bugbee↗

Improvements on Low-Density Parity-Check (LDPC) Codes and High-Performance Neuromorphic Engineering for Communication Systems

Belief propagation (BP) on LDPC codes is an iterative decoding algorithm that performs information transfer on the Tanner graph, which represents the code. In each iteration, the algorithm exchanges information (LLR) between variable nodes and check nodes through the edges of the graph. LLR values represent the probability that a given bit in a transmitted codeword equals 0 or 1, given a received word. In the hardware part, recent advancements in intelligent technologies, such as artificial intelligence, big data analytics, autonomous vehicles, and speech/image recognition, have heightened the demand for faster calculations and reduced energy consumption.

Danilo Barrionuevo↗

End-to-End Mission Design & Trajectory Optimization

Need: A need exists for a generalized, robust, user-friendly and accessible end-to-end mission design optimization tool. Solution: Our solution to developing this capability was to interface two JSC tools—Copernicus and Genesis. Each of these tools has a specific area of the mission design process that it excels at. By utilizing them both, we can gain performance benefits not seen by either on their own. - Copernicus is a trajectory design and optimization software used for in-space trajectories around multiple bodies. - Genesis is a flight mechanics tool used to model ascent, entry, descent, and landing trajectories around a single planetary body. Year 1 was focused on combining these 2 software packages—allowing Copernicus to incorporate the ascent/descent capabilities of Genesis into the optimization problem—and developing this end-to-end mission design capability. Year 2 we focused on increasing the robustness of this capability by building the initial guess generator (IGG), which produces initial guesses based on simplifying assumptions and the physics of the problem. Year 3 of our project focused on utilizing the end-to-end mission design and optimization capabilities developed in the previous 2 years to analyze specific mission scenarios—scaling up from proof of concept to real analyses—capturing any resulting performance benefits, as well as addressing the Big Data challenges we’re faced with—namely, how we’re going to manage and interpret all the data that’s generated.

Kristin Nichols↗

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

2024 Update of Comprehensive Review of Multi-arm Caliper Data for the Big Hill SPR Site

The Big Hill SPR site has a rich data set consisting of multi-arm caliper (MAC) logs collected from the cavern wells. This data set provides insight into the on-going casing deformation at the Big Hill site. This report summarizes the MAC surveys for each well and presents well longevity estimates where possible. Included in the report is an examination of the well twins for each cavern and a discussion on what may or may not be responsible for the different levels of deformation between some of the well twins. The report also takes a systematic view of the MAC data presenting spatial patterns of casing deformation and deformation orientation in an effort to better understand the underlying causes. The conclusions present a hypothesis suggesting the small-scale variations in casing deformation are attributable to similar scale variations in the character of the salt-caprock interface. These variations do not appear directly related to shear zones or faults. In addition, the deformation orientation shows no preferred directionality. This 2024 edition of this report represents an update to the original, 2023 edition. The updates primarily focus on the inclusion of MAC log data run since the December 2021 threshold date for the original report, but some new analyses are also included.

58 GEOSCIENCES↗

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data↗

Combining Large Datasets - Cancer Moonshot Task Group Final Summary

In February 2022, President Biden re-ignited the Cancer Moonshot with bold new goals: to reduce the cancer death rate by half within 25 years and improve the lives of people with cancer and cancer survivors. To achieve these ambitious goals, the White House convened the first-ever Cancer Cabinet, bringing together departments and agencies from across the federal government to end cancer as we know it.The Cancer Cabinet convened three task forces and supporting task groups, including the Data and Innovation Task Force, which supported the Cancer Moonshot priority to “Deliver innovation to patients and communities.” In early 2023, the “Combining Large Datasets” (CoLD) Task Group was created within the Data and Innovation Task Force. The scope of the CoLD Task Group was how federal agencies combine large datasets for broad applications across cancer prevention and control, including nutrition, epidemiology, and military/Veteran health. Within this scope, the group sought to better leverage the immense potential of data and power of data tools to increase our understanding of cancer incidence, causes, mortality, treatments, prevention, outcomes, costs, and all other aspects of the burden of cancer.

data integration↗

Earth Observing Data System Data and Information System (EOSDIS) Overview

The National Aeronautics and Space Administration (NASA) acquires and distributes an abundance of Earth science data on a daily basis to a diverse user community worldwide. The NASA Big Earth Data Initiative (BEDI) is an effort to make the acquired science data more discoverable, accessible, and usable. This presentation will provide a brief introduction to the Earth Observing System Data and Information System (EOSDIS) project and the nature of advances that have been made by BEDI to other Federal Users.

Earth Science↗

Metadata Evaluation and Improvement: Evolving Analysis and Reporting

ESIP Community members create and manage a large collection of environmental datasets that span multiple decades, the entire globe, and many parts of the solar system. Metadata are critical for discovering, accessing, using and understanding these data effectively and ESIP community members have successfully created large collections of metadata describing these data. As part of the White House Big Earth Data Initiative (BEDI), ESDIS has developed a suite of tools for evaluating these metadata in native dialects with respect to recommendations from many organizations. We will describe those tools and demonstrate evolving techniques for sharing results with data providers.

metadata recommendations↗