Engineering PapersSearch

SEARCH · Engineering Papers

Results for “big data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

NASA GES DISC New Data Service and Data Management for the Air Quality Community

President Obama's Big Data Research and Development Initiative seeks to improve our ability to acquire knowledge and discover insights into large and complex collections of digital data. The Big Earth Data Initiative (BEDI) Invests in standardizing and optimizing the collection, management and delivery of U.S. Government's civil Earth observation data.

air pollution

Climate Analytics as a Service

Climate science is a big data domain that is experiencing unprecedented growth. In our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). CAaaS combines high-performance computing and data-proximal analytics with scalable data management, cloud computing virtualization, the notion of adaptive analytics, and a domain-harmonized API to improve the accessibility and usability of large collections of climate data. MERRA Analytic Services (MERRA/AS) provides an example of CAaaS. MERRA/AS enables MapReduce analytics over NASA's Modern-Era Retrospective Analysis for Research and Applications (MERRA) data collection. The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of key climate variables. The effectiveness of MERRA/AS has been demonstrated in several applications. In our experience, CAaaS is providing the agility required to meet our customers' increasing and changing data management and data analysis needs.

big data

A Web of Data Analytics Services

Cloud Computing has become the ubiquitous approach to our Big Data challenge. However, one will quickly discover that moving (a.k.a. forklifting) existing on-premise data analytics solutions to the Cloud doesn’t always translate to costing saving and performance boost. The Cloud’s elasticity, its availability, and its wide selection of computing options and selections of costing models making Cloud an attractive environment to tackle our Big Data challenge. The fact is Cloud, on its own, is not the silver bullet to our daunting challenge need for analyze and derive scientific inferences through vast collections of multi-sensor measurements. We would like to have all scientific data in one easy to access environment, but getting the world of scientific data in one analytic system is immensely difficult to achieve. This paper describes the data analytics web architecture NASA is developing by infusing instances of Integrated Data Analytics systems next to the data. The goal is to minimize unnecessary data movement through collection of data access and analytics webservices for researchers to interact with and analyze measurements without have to download data to their local computer. These services are RESTful and provisioned by the data centers with the help from subject matter and science experts. These services encapsulate the physical computing infrastructure, which could local computing cluster, on-premise or public Cloud environment.

Huang, Thomas

2015 Space Radiation Standing Review Panel

The 2015 Space Radiation Standing Review Panel (from here on referred to as the SRP) met for a site visit in Houston, TX on December 8 - 9, 2015. The SRP met with representatives from the Space Radiation Element and members of the Human Research Program (HRP) to review the updated research plan for the Risk of Radiation Carcinogenesis Cancer Risk. The SRP also reviewed the newly revised Evidence Reports for the Risk of Acute Radiation Syndromes Due to Solar Particle Events (SPEs) (Acute Risk), the Risk of Acute (In-flight) and Late Central Nervous System Effects from Radiation Exposure (CNS Risk), and the Risk of Cardiovascular Disease and Other Degenerative Tissue Effects from Radiation (Degen Risk), as well as a status update on these Risks. The SRP would like to commend Dr. Simonsen, Dr. Huff, Dr. Nelson, and Dr. Patel for their detailed presentations. The Space Radiation Element did a great job presenting a very large volume of material. The SRP considers it to be a strong program that is well-organized, well-coordinated and generates valuable data. The SRP commended the tissue sharing protocols, working groups, systems biology analysis, and standardization of models. In several of the discussed areas the SRP suggested improvements of the research plans in the future. These include the following: It is important that the team has expanded efforts examining immunology and inflammation as important components of the space radiation biological response. This is an overarching and important focus that is likely to apply to all aspects of the program including acute, CVD, CNS, cancer and others. Given that the area of immunology/inflammation is highly complex (and especially so as it relates to radiation), it warrants the expansion of investigators expertise in immunology and inflammation to work with the individual research projects and also the NASA Specialized Center of Research (NSCORs). Historical data on radiation injury to be entered into the Watson “big data” study must be used with caution. The general scientific issues of reproducibility, details of experimental methods and data analysis from preclinical and basic research laboratories have been raised broadly over the last few years (not specific to this work) and indicate that caution must be applied in the ways these data are used. This pertains to preclinical data and also to phase 3 clinical trials in radiation oncology and medical oncology. Of course, appropriate use and analysis of these “big-data” sets also offer the potential of pinpointing limitations and extracting remaining useful information. Emphasis should be placed on the latter possibility. A key target is risk reduction from radiation exposure. Progress of the entire space program, now moving towards the Mars mission, requires timely answers to key components of human risk, which are known to be complex. Periodic review of progress should be conducted with additional resources directed into achieving critical milestones. Turning the long red bars to yellow and green (or for some risks such as CNS possibly to grey) must be high priority. That such progress will require new science and not engineering means that it should be viewed in a knowledge-based light. The technology-based aspects of engineering issues are certainly as important, however, science and knowledge-based problems are solved in a different way than engineering. Timelines for engineering are more predictable, while for science, progress can be methodical with occasional major incremental findings that can rapidly change the rate of progress. As opportunities for rapid incremental changes arise, periodic enhancement of investment is strongly recommended to enable such new knowledge to be quickly and efficiently exploited. Collaborations and linkages with National Institute of Allergy and Infectious Diseases (NIAID), the Biomedical Advanced Research and Development Authority (BARDA) and the Department of Defense (DoD) are in place and more are encouraged, where possible, with the radiation injury and medical countermeasure studies. This could include utilizing some of their animal model testing contracts to facilitate obtaining results using common platforms. Such approach will facilitate the comparison of results among laboratories, and will facilitate and accelerate the development of medical countermeasures. It is particularly noteworthy that the NASA Space Radiation Element is reaching out to the Multidisciplinary European Low Dose Initiative (MELODI) platform coordinating low dose radiation risk research, and to other international agencies that are studying low dose radiation effects in an effort to fill the void generated by the cancelation of the Department of Energy (DOE) low dose radiation program. While NASA is working actively with NIAID and BARDA to integrate their relevant findings of radiation mitigator investigations to NASA programs, the committee notes its disappointment that the United States currently lacks a dedicated low dose radiation program with clear mechanistic orientation and aimed at the quantification and mitigation of human radiation risk on Earth. This void gives to the NASA Space Radiation Program Element special societal value, but also makes its overall design more challenging.

Steinberg, Susan

Analyzing a 35-Year Hourly Data Record: Why So Difficult?

At the Goddard Distributed Active Archive Center, we have recently added a 35-Year record of output data from the North American Land Assimilation System (NLDAS) to the Giovanni web-based analysis and visualization tool. Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) offers a variety of data summarization and visualization to users that operate at the data center, obviating the need for users to download and read the data themselves for exploratory data analysis. However, the NLDAS data has proven surprisingly resistant to application of the summarization algorithms. Algorithms that were perfectly happy analyzing 15 years of daily satellite data encountered limitations both at the algorithm and system level for 35 years of hourly data. Failures arose, sometimes unexpectedly, from command line overflows, memory overflows, internal buffer overflows, and time-outs, among others. These serve as an early warning sign for the problems likely to be encountered by the general user community as they try to scale up to Big Data analytics. Indeed, it is likely that more users will seek to perform remote web-based analysis precisely to avoid the issues, or the need to reprogram around them. We will discuss approaches to mitigating the limitations and the implications for data systems serving the user communities that try to scale up their current techniques to analyze Big Data.

computational performance

Efficient First-Order Algorithms for Large-Scale, Non-Smooth Maximum Entropy Models with Application to Wildfire Science

Maximum entropy (MaxEnt) models are a class of statistical models that use the maximum entropy principle to estimate probability distributions from data. Due to the size of modern data sets, MaxEnt models need efficient optimization algorithms to scale well for big data applications. State-of-the-art algorithms for MaxEnt models, however, were not originally designed to handle big data sets; these algorithms either rely on technical devices that may yield unreliable numerical results, scale poorly, or require smoothness assumptions that many practical MaxEnt models lack. In this paper, we present novel optimization algorithms that overcome the shortcomings of state-of-the-art algorithms for training large-scale, non-smooth MaxEnt models. Our proposed first-order algorithms leverage the Kullback–Leibler divergence to train large-scale and non-smooth MaxEnt models efficiently. For MaxEnt models with discrete probability distribution of n elements built from samples, each containing m features, the stepsize parameter estimation and iterations in our algorithms scale on the order of O(mn) operations and can be trivially parallelized. Moreover, the strong ℓ1 convexity of the Kullback–Leibler divergence allows for larger stepsize parameters, thereby speeding up the convergence rate of our algorithms. To illustrate the efficiency of our novel algorithms, we consider the problem of estimating probabilities of fire occurrences as a function of ecological features in the Western US MTBS-Interagency wildfire data set. Our numerical results show that our algorithms outperform the state of the art by one order of magnitude and yield results that agree with physical models of wildfire occurrence and previous statistical analyses of wildfire drivers.

Physics

NASA's Big Earth Data Initiative Accomplishments

The goal of NASA's effort for BEDI is to improve the usability, discoverability, and accessibility of Earth Observation data in support of societal benefit areas. Accomplishments: In support of BEDI goals, datasets have been entered into Common Metadata Repository(CMR), made available via the Open-source Project for a Network Data Access Protocol (OPeNDAP), have a Digital Object Identifier (DOI) registered for the dataset, and to support fast visualization many layers have been added in to the Global Imagery Browse Services (GIBS).

CMR

Smallholder Crop Area Mapped with Wall-To-Wall WorldView Sub-Meter Panchromatic Image Texture: A Test Case for Tigray, Ethiopia

Global food production in the developing world occurs within sub-hectare fields that are difficult to identify with moderate resolution satellite imagery. Knowledge about the distribution of these fields is critical in food security programs. We developed a semi-automated image segmentation approach using wall-to-wall sub-meter imagery with high-performance computing to map crop area (CA) throughout Tigray, Ethiopia that encompasses over 41,000 km (exp 2). Multiple processing streams were tested to minimize mapping error while applying five unique smoothing kernels to capture differences in land surface texture associated to CA. Typically, very-small fields (mean < 2 ha) have a smooth image roughness compared to natural scrub/shrub woody vegetation at the ~1m scale and these features can be segmented in panchromatic imagery with multi-level histogram thresholding. Multi-temporal very-high resolution (VHR) panchromatic imagery with multi-spectral VHR are sufficient in extracting critical CA information needed in food security programs. A 2011 to 2015 CA map was produced, using over 3000 WorldView-1 panchromatic images wall-to-wall in 1/2 deg mosaics for Tigray, Ethiopia. CA was evaluated with nearly 3000 WorldView-2 2m multispectral 250 X 250 m image subsets by seven expert interpretations, and with in-situ global positioning system photography. CA estimates ranged from 32 to 41% in sub regions of Tigray with median maximum per bin commission and omission errors of 11% and 1% respectively, with most of the error occurring in bins <15%. This empirical, simple, and low direct cost approach via U.S. government license agreement to access commercial VHR data, could be a viable big-data high-performance computing methodology to extract wall-to-wall CA for other regions of the world that have very-small agriculture fields with similar image texture."

Ethiopia

Using Machine Learning to Predict Core Sizes of High-Efficiency Turbofan Engines

With the rise in big data and analytics, machine learning is transforming many industries. It is being increasingly employed to solve a wide range of complex problems, producing autonomous systems that support human decision-making. For the aircraft engine industry, machine learning of historical and existing engine data could provide insights that help drive for better engine design. This work explored the application of machine learning to engine preliminary design. Engine core-size prediction was chosen for the first study because of its relative simplicity in terms of number of input variables required (only three). Specifically, machine-learning predictive tools were developed for turbofan engine core-size prediction, using publicly available data of two hundred manufactured engines and engines that were studied previously in NASA aeronautics projects. The prediction results of these models show that, by bringing together big data, robust machine-learning algorithms and automation, a machine learning-based predictive model can be an effective tool for turbofan engine core-size prediction. The promising results of this first study paves the way for further exploration of the use of machine learning for aircraft engine preliminary design.

Tong, Michael T.

Using Machine Learning to Predict Core Sizes of High-Efficiency Turbofan Engines

With the rise in big data and analytics, machine learning is transforming many industries. It is being increasingly employed to solve a wide range of complex problems, producing autonomous systems that support human decision-making. For the aircraft engine industry, machine learning of historical and existing engine data could provide insights that help drive for better engine design. This work explored the application of machine learning to engine preliminary design. Engine core-size prediction was chosen for the first study because of its relative simplicity in terms of number of input variables required (only three). Specifically, machine-learning predictive tools were developed for turbofan engine core-size prediction, using publicly available data of two hundred manufactured engines and engines that were studied previously in NASA aeronautics projects. The prediction results of these models show that, by bringing together big data, robust machine-learning algorithms and data science, a machine learning-based predictive model can be an effective tool for turbofan engine core-size prediction. The promising results of this first study paves the way for further exploration of the use of machine learning for aircraft engine preliminary design.

Core Size

LandScan Mosaic

The LandScan program at Oak Ridge National Laboratory (ORNL), in collaboration with the National Geospatial-Intelligence Agency (NGA), continues to deliver the most accurate and up to date global, high resolution gridded population data. Additionally, the latest advancements in the LandScan HD methodology led to reduced latency in development of rapid updates for geopolitical events. With momentum towards reporting more up to date population estimates, feedback from the user community expressed interest in reporting population estimates in ranges - whether to express a level of uncertainty or confirm to leadership and stakeholders the modeled data are estimates. Building upon the need to understand uncertainty or confidence in the modeled data and report ranges at the global scale, LandScan Mosaic was developed. LandScan Mosaic represents the next generation of high-resolution population modeling, building upon the established success of previous LandScan HD iterations. While LandScan HD employed a deterministic big data fusion approach, LandScan Mosaic enhances this methodology by integrating advanced machine learning techniques to impute missing, yet crucial, population model parameters. This advancement allows for probabilistic modeling of building occupancy and population distribution, incorporating uncertainty quantification through Monte Carlo sampling methods. By combining big data fusion with machine learning-driven imputation and stochastic modeling, LandScan Mosaic provides a more comprehensive and robust representation of population dynamics. LandScan Mosaic will be following the in the footsteps of its longstanding counterpart LandScan Global and releasing a global gridded population raster, at the 3-arcsecond resolution. This technical report documents the current stage of development of LandScan Mosaic, detailing the methodologies and data sources behind the modeling. Stakeholders are encouraged to use this document as an authoritative reference for insight into Mosaic’s data development processes. However, readers should note that LandScan Mosaic remains in a late-stage research and development phase, and methodologies and data presented here are subject to refinements ahead of the anticipated global release in Summer 2025. Feedback and inquiries from users and stakeholders are welcomed as we continue to refine and enhance this important population resource.

97 MATHEMATICS AND COMPUTING

Machine Learning Technologies and Their Applications for Science and Engineering Domains Workshop -- Summary Report

The fields of machine learning and big data analytics have made significant advances in recent years, which has created an environment where cross-fertilization of methods and collaborations can achieve previously unattainable outcomes. The Comprehensive Digital Transformation (CDT) Machine Learning and Big Data Analytics team planned a workshop at NASA Langley in August 2016 to unite leading experts the field of machine learning and NASA scientists and engineers. The primary goal for this workshop was to assess the state-of-the-art in this field, introduce these leading experts to the aerospace and science subject matter experts, and develop opportunities for collaboration. The workshop was held over a three day-period with lectures from 15 leading experts followed by significant interactive discussions. This report provides an overview of the 15 invited lectures and a summary of the key discussion topics that arose during both formal and informal discussion sections. Four key workshop themes were identified after the closure of the workshop and are also highlighted in the report. Furthermore, several workshop attendees provided their feedback on how they are already utilizing machine learning algorithms to advance their research, new methods they learned about during the workshop, and collaboration opportunities they identified during the workshop.

Ambur, Manjula

NASA EOSDIS Evolution in the BigData Era

NASA's EOSDIS system faces several challenges in the Big Data Era. Although volumes are large (but not unmanageably so), the variety of different data collections is daunting. That variety also brings with it a large and diverse user community. One key evolution EOSDIS is working toward is to enable more science analysis to be performed close to the data.

Information Systems

Evaluating the Impact of Data Placement to Spark and SciDB with an Earth Science Use Case

We investigate the impact of data placement for two Big Data technologies, Spark and SciDB, with a use case from Earth Science where data arrays are multidimensional. Simultaneously, this investigation provides an opportunity to evaluate the performance of the technologies involved. Two datastores, HDFS and Cassandra, are used with Spark for our comparison. It is found that Spark with Cassandra performs better than with HDFS, but SciDB performs better yet than Spark with either datastore. The investigation also underscores the value of having data aligned for the most common analysis scenarios in advance on a shared nothing architecture. Otherwise, repartitioning needs to be carried out on the fly, degrading overall performance.

Spark

Introduction to Big Earth Data Applications

Climate and weather modeling generate enormous volumes that make iterative analysis challenging, spurring the development of new ways to work with the data. A theme going across applications is the need to identify and highlight "interesting" data for the scientist to focus on. Operational applications often scale up from small, local studies to larger spatial scales with more analysis targets.

parallel processing (computers)

Introduction to Big Earth Data Applications

Climate and weather modeling generate enormous volumes that make iterative analysis challenging, spurring the development of new ways to work with the data. At the same time in the Earth Observation area, technology advances are enabling new sensors and satellites that will increase data volume, velocity and application variety. Scaling up can also be seen when operational applications expand from small, local studies to larger spatial scales with more analysis targets.

Christopher Lynnes