Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Big Data and AI at DoE's Legacy Sites - 20546

More than 30 years have passed since DOE started the decommissioning of nuclear weapon complexes and the clean-up of soil and groundwater. All the sites have been collecting and archiving soil and groundwater monitoring datasets; particularly contaminant concentration time-series. These datasets provide unparalleled opportunities to understand the system behavior (including more fundamental hydrological and geochemical processes, the response to various perturbations, the long-term trend and environmental decay rate towards the regulatory limit). This understanding is critical for providing multiple lines of evidences that can support site closure. In this study, we explore the machine learning (ML) and artificial intelligence (AI) applications to the long-term soil and groundwater management at DoE's legacy sites. ML can improve our understanding of the subsurface systems, which is critical for long-term monitoring and management of the sites, while AI can automate or support some of decision-making processes (e.g., anomaly detection, monitoring well placements). The particular focuses are to develop general algorithms to: (1) to identify distinct spatiotemporal patterns and to identify several groups that have similar temporal behaviors, using unsupervised clustering methods, (2) identify the different temporal scales of hydrological responses to climate perturbations by time-series analysis, and (3) reduce the number of monitoring wells by identifying the minimum sufficient number of wells to capture the heterogeneity of the groundwater contaminant plume and concentration distribution, using the Gaussian Process model. We demonstrate our methodology at the Savannah River Site F-Area. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Learning to Correct Climate Projection Biases

The fidelity of climate projections is often undermined by biases in climate models due to their simplification or misrepresentation of unresolved climate processes. While various bias correction methods have been developed to post-process model outputs to match observations, existing approaches usually focus on limited, low-order statistics, or break either the spatiotemporal consistency of the target variable, or its dependency upon model resolved dynamics. We develop a Regularized Adversarial Domain Adaptation (RADA) methodology to overcome these deficiencies, and enhance efficient identification and correction of climate model biases. Instead of pre-assuming the spatiotemporal characteristics of model biases, we apply discriminative neural networks to distinguish historical climate simulation samples and observation samples. The evidences based on which the discriminative neural networks make distinctions are applied to train the domain adaptation neural networks to bias correct climate simulations. We regularize the domain adaptation neural networks using cycle-consistent statistical and dynamical constraints. An application to daily precipitation projection over the contiguous United States shows that our methodology can correct all the considered moments of daily precipitation at approximately $1^\circ$ resolution, ensures spatiotemporal consistency and inter-field correlations, and can discriminate between different dynamical conditions. Our methodology offers a powerful tool for disentangling model parameterization biases from their interactions with the chaotic evolution of climate dynamics, opening a novel avenue toward big-data enhanced climate predictions.

58 GEOSCIENCES↗

Ocean Wave Studies with Applications to Ocean Modeling and Improvement of Satellite Altimeter Measurements

Combining analysis of satellite data (altimeter, scatterometer, high-resolution visible and infrared images, etc.) with mathematical modeling of non-linear wave processes, we investigate various ocean wave fields (on scales from capillary to planetary), their role in ocean dynamics and turbulent transport (of heat and biogeochemical quantities), and their effects on satellite altimeter measuring accuracy. In 1998 my attention was focused on long internal gravity waves (10 to 1000 km), known also as baroclinic inertia-gravity (BIG) waves. We found these waves to be a major factor of altimeter measurements "noise," resulting in a greater uncertainty [up to 10 cm in terms of sea surface height (SSH) amplitude] in the measured SSH signal than that caused by the sea state bias variations (up to 5 cm or so). This effect still remains largely overlooked by the satellite altimeter community. Our studies of BIG waves address not only their influence on altimeter measurements but also their role in global ocean dynamics and in transport and turbulent diffusion of biogeochemical quantities. In particular, in collaboration with Prof Peter Weichman, Caltech, we developed a theory of turbulent diffusion caused by wave motions of most general nature. Applied to the problem of horizontal turbulent diffusion in the ocean, the theory yielded the effective diffusion coefficient as a function of BIG wave parameters obtainable from satellite altimeter data. This effort, begun in 1997, has been successfully completed in 1998. We also developed a theory that relates spatial fluctuations of scalar fields (such as sea surface temperature, chlorophyll concentration, drifting ice concentration, etc.) to statistical characteristics of BIG waves obtainable from altimeter measurements. A manuscript is in the final stages of preparation. In order to verify the theoretical predictions and apply them to observations, we are now analyzing Sea-viewing Wide Field of view Sensor (SeaWiFS) and Field of view Sensor (SeaWiFS) and Advanced Very High-Resolution Radiometer (AVHRR) data on sea surface temperature (SST) and chlorophyll concentration jointly with TOPEX/POSEIDON data on SSH variations.

Glazman, Roman E.↗

Big PanDa Workflow Management on Titan for High Energy and Nuclear Physics and for Future Extreme Scale Scientific Application

Over a three year period, from 2016-2019, this project demonstrated the scientific benefits of integrating the Titan supercomputer at Oak Ridge Leadership Computing Facility into traditional high throughput grid based distributed computing systems managed by PanDA, the workflow management system used for the execution of all distributed computing applications by the ATLAS experiment at the Large Hadron Collider. PanDA manages millions of batch jobs daily at hundreds of clusters worldwide on request by thousands of physicist users, and processes more than an exabyte of data annually using grid middleware. High levels of operational use of Titan was sustained by PanDA in order to meet the physics goals of ATLAS. The success of this project led to the use of other supercomputers worldwide by ATLAS, and to the adoption of PanDA by other experiments and other scientists. Multiple innovative operational and computer science research goals were achieved supporting the use of supercomputers for scientific domains with large scale distributed data and distributed processing needs.

97 MATHEMATICS AND COMPUTING↗

Application of a Machine Learning Algorithm in Generating an Evapotranspiration Data Product From Coupled Thermal Infrared and Microwave Satellite Observations

Land surface evapotranspiration (ET) is one of the main energy sources for atmospheric dynamics and a critical component of the local, regional, and global water cycles. Consequently, accurate measurement or estimation of ET is one of the most active topics in hydro-climatology research. With massive and spatially distributed observational data sets of land surface properties and environmental conditions being collected from the ground, airborne or space-borne platforms daily over the past few decades, many research teams have started to use big data science to advance the ET estimation methods. The Geostationary satellite Evapotranspiration and Drought (GET-D) product system was developed at the National Oceanic and Atmospheric Administration (NOAA) in 2016 to generate daily ET and drought maps operationally. The primary inputs of the current GET-D system are the thermal infrared (TIR) observations from NOAA GOES satellite series. Because of the cloud contamination to the TIR observations, the spatial coverage of the daily GET-D ET product has been severely impacted. Based on the most recent advances, we have tested a machine learning algorithm to estimate all-weather land surface temperature (LST) from TIR and microwave (MW) combined satellite observations. With the regression tree machine learning approach, we can combine the high accuracy and high spatial resolution of GOES TIR data with the better spatial coverage of passive microwave observations and LST simulations from a land surface model (LSM). The regression tree model combines the three LST data sources for both clear and cloudy days, which enables the GET-D system to derive an all-weather ET product. This paper reports how the all-weather LST and ET are generated in the upgraded GET-D system and provides an evaluation of these LST and ET estimates with ground measurements. The results demonstrate that the regression tree machine learning method is feasible and effective for generating daily ET under all weather conditions with satisfactory accuracy from the big volume of satellite observations.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning for Synchrophasor Analysis

The report presents results from the development of a cloud-based, Big Data analysis framework for power systems. The computational pipeline uses the Apache Spark framework running in an OpenStack cloud infrastructure. A real-world phasor measurement unit (PMU) dataset has been used to carry out the analysis. Several Machine Learning (ML) methods have been developed and implemented for event and anomaly detection and classification. Actual examples of power system events detection and analysis using synchrophasor data are presented. It has been shown that applications of the cloud-based computing environment and the Apache Spark framework enable a significant increase in the computational efficiency of large-scale PMU data analysis.

20 FOSSIL-FUELED POWER PLANTS↗

NASA GES DISC's Customized Services for Climatology and Meteorology

At the NASA Goddard Earth Sciences (GES) Data and Information Service Center (DISC), we have archived and distributed more than 2,400 Earth science data products, from different missions or projects containing more than 100 M data files/granules with a total volume size nearly 2 PB that broadly serve user needs in science areas such as Atmospheric Composition, Water & Energy Cycles and Climate Variability. To date, GES DISC has developed many pertinent services to facilitate the usage of data products by our research communities, represented by approximately 24,000 registered users. We are facing the big data with increasingly archival volume and data types, moreover, we also encounter increasing users' demands and the demands are more diversified. It is still a challenge for us to better understand exactly what our users' needs are, even after developing more than 70 services, including well-known online tools such as Giovanni and MERRA subsetter. In this presentation, we will try to address how we can accommodate the users' needs from two applicational user communities, Air Quality and Wind Energy, from data or service discovery to guide them properly utilize the data and services to fit their needs.

customizable services for climate and meteorology↗

SWIPE: Spectral Water Inversion Processor and Emulator

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will be discussing the progress made developing SWIPE: Spectral Water Inversion Processor and Emulator. SWIPE is a platform for advanced modeling of coastal and inland aquatic habitats. The goal is create a comprehensive and cohesive system to leverage recent advancements in computation and machine learning to develop a synthetic training ground for sensitivity studies and algorithm development. The four principal facets of SWIPE include: 1. Advanced two-layer coated sphere bio-optical modeling and GPU radiative transfer modeling, 2. Big Data involving massive synthetic spectral libraries of optical properties of various global aquatic particles, surface reflectance, and top-of-atmosphere reflectance, all at hyperspectral resolution leveraging high-end computing systems at NASA Ames Research Center, 3. Deep Learning for algorithm development for water quality inversion of concentrations of common biogeophysical variables as well as optics, full uncertainty characterization by water type, and forward emulation, and lastly, 4. Image Processing for application of developed retrieval algorithms for both hyperspectral and multispectral sensors with experimental corrections for global adjacency, noise, sunglint, and benthic reflectance. This presentation will demonstrate the Equivalent Algal Populations (EAP) two-layer coated sphere scattering model which has been used develop spectral libraries of hyperspectral inherent optical properties of roughly 80 species of phytoplankton, covering 15 different classes and nine taxonomic functional types. The EAP model was also used to derive spectral properties of 10 different non-algal particle functional types. Examples of how the SMART-G (Speed-up Monte-carlo Advanced Radiative Transfer using GPU) radiative transfer code is used to model optically complex aquatic signals will be presented and discussed in the context of creating a massive synthetic database which can leverage the full power of next generation machine learning techniques and high end computing for water quality inversion. We will discuss our active investigation in things like appropriate model architectures, dimensionality reduction techniques such as PCA and autoencoders, uncertainty quantification and abstaining, and which variables actually benefit most from hyperspectral information versus multispectral resolution. We are also curious about questions relating to cost/benefit analysis in terms of computation resources, neural network complexity, and data volumes. Answers to these questions will hopefully elaborate on cost efficiency for potential future sensor design considerations.

SWIPE↗

Machine learning in materials science: From explainable predictions to autonomous design

The advent of big data and algorithmic developments in the field of machine learning (and artificial intelligence, in general) have greatly impacted the entire spectrum of physical sciences, including materials science. Materials data, measured or computed, combined with various techniques of machine learning have been employed to address a myriad of challenging problems, such as, development of efficient and predictive surrogate models for a range of materials properties, screening and down-selection of novel candidate materials for targeted applications, new methodologies to improve and further expedite molecular and atomistic simulations, with likely many more important developments to come in the foreseeable future. While the applications thus far have provided a glimpse of the true potential data-enabled routes have to offer, it has also become clear that further progress in this direction hinges on our ability to understand, explain and rationalize findings of a machine learning model in light of the domain-knowledge. This focused review provides an overview of the main areas where machine learning has been widely and successfully used in materials science. Subsequently, a brief discussion of several techniques that have been helpful in extracting physically-meaningful insights, causal relationships and design-centric knowledge from materials data is provided. Finally, we identify some of the imminent opportunities and challenges that materials community faces in this exciting and rapidly growing field.

36 MATERIALS SCIENCE↗

Analyzing a 35-Year Hourly Data Record: Why So Difficult?

At the Goddard Distributed Active Archive Center, we have recently added a 35-Year record of output data from the North American Land Assimilation System (NLDAS) to the Giovanni web-based analysis and visualization tool. Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) offers a variety of data summarization and visualization to users that operate at the data center, obviating the need for users to download and read the data themselves for exploratory data analysis. However, the NLDAS data has proven surprisingly resistant to application of the summarization algorithms. Algorithms that were perfectly happy analyzing 15 years of daily satellite data encountered limitations both at the algorithm and system level for 35 years of hourly data. Failures arose, sometimes unexpectedly, from command line overflows, memory overflows, internal buffer overflows, and time-outs, among others. These serve as an early warning sign for the problems likely to be encountered by the general user community as they try to scale up to Big Data analytics. Indeed, it is likely that more users will seek to perform remote web-based analysis precisely to avoid the issues, or the need to reprogram around them. We will discuss approaches to mitigating the limitations and the implications for data systems serving the user communities that try to scale up their current techniques to analyze Big Data.

computational performance↗

Coevolution of Machine Learning and Process-Based Modelling to Revolutionize Earth and Environmental Sciences: A Perspective

Machine learning (ML) applications in Earth and environmental sciences (EES) have gained incredible momentum in recent years. However, these ML applications have largely evolved in ‘isolation’ from the mechanistic, process-based modelling (PBM) paradigms, which have historically been the cornerstone of scientific discovery and policy support. In this perspective, we assert that the cultural barriers between the ML and PBM communities limit the potential of ML, and even its ‘hybridization’ with PBM, for EES applications. Fundamental, but often ignored, differences between ML and PBM are discussed as well as their strengths and weaknesses in light of three overarching modelling objectives in EES, (1) nowcasting and prediction, (2) scenario analysis, and (3) diagnostic learning. The paper ponders over a ‘coevolutionary’ approach to model building, shifting away from a borrowing to a co-creation culture, to develop a generation of models that leverage the unique strengths of ML such as scalability to big data and high-dimensional mapping, while remaining faithful to process-based knowledge base and principles of model explainability and interpretability, and therefore, falsifiability.

Saman Razavi↗

Multi-Functional Distributed Fiber Sensors for Pipeline Monitoring and Methane Detections. Final Report

As an abundant and cheap fossil energy source, natural gas has become a significant energy supply to support the United States’ economy. However, the large-scale extraction and utilization of natural gas also impose significant challenges on methane leakage. This problem is exacerbated by aging gas utility delivery systems, including interstate high-pressure pipelines, storage, and transmission facilities. This project aims to develop a cost-effective fiber optical sensing method that can perform multi-parameter real-time measurements of natural gas pipelines across long interrogation distances up to 100 km with 1-meter spatial resolution. This sensing tool can evaluate overall pipeline efficiency and reduce methane emissions for mid-stream methane infrastructures. To accomplish this objective, research and development efforts funded by this project have resulted in the following accomplishments: This project successfully has developed new functional sensory polymer materials using Metal-Organic Frameworks (MOFs) that can be coated on optical fiber through the reel-to-reel coating process. Functional polymer-coated optical fibers can perform sensitive methane detection through evanescence wave interaction and strain-based measurements to achieve 1% detection sensitivities. The new sensors fibers support both distributed measurements and multiplexed fiber sensors array for multi-point measurements. This project developed and optimized a new multi-core optical fiber that supports simultaneous and distributed measurements of strain and temperatures with 1-meter spatial resolutions across up to 100-km interrogation distance. This new fiber, combined with sensory polymercoated fiber, could perform both distributed temperature and methane detections. This project developed a new artificial intelligence big data algorithm approach that can effectively analyze high-resolution data harnessed by distributed fiber sensors to protect natural gas pipelines against external threats and detect internal defects induced by corrosion. Working with our industry partner, this project developed new optical fibers that support fiber sensor fabrications through polymer coating after the fibers are drawn. These new fibers eliminate the need for direct sensor fabrication when the fiber is fabricated on a fiber draw tower, which drastically expands fiber sensors' applicability. This research project has significantly advanced the distributed fiber sensing technology. It will dramatically increase the applicability and adaptability of distributed fiber sensors for a wide array of applications in energy, sustainability, and environmental science, including structural health monitoring of natural gas pipelines, oil infrastructures, hydrogen facilities, and environmental monitoring of carbon storage sites, water supply systems, and others.

03 NATURAL GAS↗

Decision Manifold Approximation for Physics-Based Simulations

With the recent surge of success in big-data driven deep learning problems, many of these frameworks focus on the notion of architecture design and utilizing massive databases. However, in some scenarios massive sets of data may be difficult, and in some cases infeasible, to acquire. In this paper we discuss a trajectory-based framework that quickly learns the underlying decision manifold of binary simulation classifications while judiciously selecting exploratory target states to minimize the number of required simulations. Furthermore, we draw particular attention to the simulation prediction application idealized to the case where failures in simulations can be predicted and avoided, providing machine intelligence to novice analysts. We demonstrate this framework in various forms of simulations and discuss its efficacy.

Wong, Jay Ming↗

Mining Twitter Data to Augment NASA GPM Validation

The Twitter data stream is an important new source of real-time and historical global information for potentially augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. There have been other similar uses of Twitter, though mostly related to natural hazards monitoring and management. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. Twitter provides a large source of crowd for crowdsourcing. During a 24-hour period in the middle of the snow storm this past March in the U.S. Northeast, we collected more than 13,000 relevant precipitation tweets with exact geolocation. The overall objective of our project is to determine the extent to which processed tweets can provide additional information that improves the validation of GPM data. Though our current effort focuses on tweets and precipitation, our approach is general and applicable to other social media and other geophysical measurements. Specifically, we have developed an operational infrastructure for processing tweets, in a format suitable for analysis with GPM data; engaged with potential participants, both passive and active, to "enrich" the Twitter stream; and inter-compared "precipitation" tweet data, ground station data, and GPM retrievals. In this presentation, we detail the technical capabilities of our tweet processing infrastructure, including data abstraction, feature extraction, search engine, context-awareness, real-time processing, and high volume (big) data processing; various means for "enriching" the Twitter stream; and results of inter-comparisons. Our project should bring a new kind of visibility to Twitter and engender a new kind of appreciation of the value of Twitter by the science research communities.

validatio↗

Process Image Analysis using Big Data, Machine Learning, and Computer Vision

The development of algorithms for machine learning and data analysis for the 3013 MIS corrosion surveillance program is a collaborative effort by SRNL, USC and GT. For corrosion detection, LCM image data is extracted from large binary files, with software written to convert the data to physical attributes (i.e. height, color and grayscale values; all as functions of a location in a plane projection). The user interface for the software permits selective downloading of binary data and interrogation of attributes. User input thresholds are used to flag attributes of interest. Machine learning algorithms, developed for this application, are used to determine whether the features are the result of corrosion. To address the fundamental mechanisms of corrosion, machine learning algorithms are being developed to derive interatomic potential force-fields from ab-initio DFT calculations. The goal is to apply molecular modeling on a large enough scale to guide the design of resistant materials.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

New generation lidar systems for eye safe full time observations

The traditional lidar over the last thirty years has typically been a big pulse low repetition rate system. Pulse energies are in the 0.1 to 1.0 J range and repetition rates from 0.1 to 10 Hz. While such systems have proven to be good research tools, they have a number of limitations that prevent them from moving beyond lidar research to operational, application oriented instruments. These problems include a lack of eye safety, very low efficiency, poor reliability, lack of ruggedness and high development and operating costs. Recent advances in solid state laser, detectors and data systems have enabled the development of a new generation of lidar technology that meets the need for routine, application oriented instruments. In this paper the new approaches to operational lidar systems will be discussed. Micro pulse lidar (MPL) systems are currently in use, and their technology is highlighted. The basis and current development of continuous wave (CW) lidar and potential of other technical approaches is presented.

Spinhirne, James D.↗

NeuralCubes: Deep Representations for Visual Data Exploration

Visual exploration of large multi-dimensional datasets has seen tremendous progress in recent years, allowing users to express rich data queries that produce informative visual summaries, all in real time. Techniques based on data cubes are some of the most promising approaches. However, these techniques usually require a large memory footprint for large datasets. To tackle this problem, we present NeuralCubes: neural networks that predict results for aggregate queries, similar to data cubes. NeuralCubes learns a function that takes as input a given query, for instance, a geographic region and temporal interval, and outputs the result of the query. The learned function serves as a real-time, low-memory approximator for aggregation queries. Our models are small enough to be sent to the client side (e.g. the web browser for a web-based application) for evaluation, enabling data exploration of large datasets without database/network connection. Here, we demonstrate the effectiveness of NeuralCubes through extensive experiments on a variety of datasets and discuss how NeuralCubes opens up opportunities for new types of visualization and interaction.

97 MATHEMATICS AND COMPUTING↗

Scalable Adaptive Graphics Environment (SAGE) Software for the Visualization of Large Data Sets on a Video Wall

The use of collaborative scientific visualization systems for the analysis, visualization, and sharing of "big data" available from new high resolution remote sensing satellite sensors or four‐dimensional numerical model simulations is propelling the wider adoption of ultra‐resolution tiled display walls interconnected by high speed networks. These systems require a globally connected and well‐integrated operating environment that provides persistent visualization and collaboration services. This abstract and subsequent presentation describes a new collaborative visualization system installed for NASA's Shortterm Prediction Research and Transition (SPoRT) program at Marshall Space Flight Center and its use for Earth science applications. The system consists of a 3 x 4 array of 1920 x 1080 pixel thin bezel video monitors mounted on a wall in a scientific collaboration lab. The monitors are physically and virtually integrated into a 14' x 7' for video display. The display of scientific data on the video wall is controlled by a single Alienware Aurora PC with a 2nd Generation Intel Core 4.1 GHz processor, 32 GB memory, and an AMD Fire Pro W600 video card with 6 mini display port connections. Six mini display‐to‐dual DVI cables are used to connect the 12 individual video monitors. The open source Scalable Adaptive Graphics Environment (SAGE) windowing and media control framework, running on top of the Ubuntu 12 Linux operating system, allows several users to simultaneously control the display and storage of high resolution still and moving graphics in a variety of formats, on tiled display walls of any size. The Ubuntu operating system supports the open source Scalable Adaptive Graphics Environment (SAGE) software which provides a common environment, or framework, enabling its users to access, display and share a variety of data‐intensive information. This information can be digital‐cinema animations, high‐resolution images, high‐definition video‐teleconferences, presentation slides, documents, spreadsheets or laptop screens. SAGE is cross‐platform, community‐driven, open‐source visualization and collaboration middleware that utilizes shared national and international cyberinfrastructure for the advancement of scientific research and education.

Jedlovec, Gary↗