Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

15 GEOTHERMAL ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.

accessibility↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Introduction to the special issue on smart transportation

Transportation is getting smarter and smarter, with the prominence of connected automated vehicle technologies in the global auto industry’s near-term growth strategies, of big data analytics and unprecedented access to sensing data of mobility, and of integration of this analytics into the optimization of mobility and transport. Further, these developments are setting off a wave of smart transportation innovations, which are featured by new methods and applications driven by various forms of sensor data such as GPS, CAN bus, LIDA, images, etc. At the same time, complexities surrounding the use, conflation, and processing of disparate data in near real-time is of essence for the design and development of futuristic smart transportation.

33 ADVANCED PROPULSION SYSTEMS↗

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data↗

PanDA: Production and Distributed Analysis System

The Production and Distributed Analysis (PanDA) system is a data-driven workload management system engineered to operate at the LHC data processing scale. The PanDA system provides a solution for scientific experiments to fully leverage their distributed heterogeneous resources, showcasing scalability, usability, flexibility, and robustness. The system has successfully proven itself through nearly two decades of steady operation in the ATLAS experiment, addressing the intricate requirements such as diverse resources distributed worldwide at about 200 sites, thousands of scientists analyzing the data remotely, the volume of processed data beyond the exabyte scale, dozens of scientific applications to support, and data processing over several billion hours of computing usage per year. PanDA’s flexibility and scalability make it suitable for the High Energy Physics community and wider science domains at the Exascale. Beyond High Energy Physics, PanDA’s relevance extends to other big data sciences, as evidenced by its adoption in the Vera C. Rubin Observatory and the sPHENIX experiment. As the significance of advanced workflows continues to grow, PanDA has transformed into a comprehensive ecosystem, effectively tackling challenges associated with emerging workflows and evolving computing technologies. The paper discusses PanDA’s prominent role in the scientific landscape, detailing its architecture, functionality, deployment strategies, project management approaches, results, and evolution into an ecosystem.

97 MATHEMATICS AND COMPUTING↗

Operation and performance of VRF systems: Mining a large-scale dataset

The energy consumption of air-conditioning systems has gained increasing attention as it contributes significantly to the global building energy use. The variable refrigerant flow (VRF) system is a common air-conditioning system applied widely in residential and office buildings in China. Understanding the actual operation and performance of VRF systems is fundamental for the energy-efficient design and operation of VRF systems. Previous research on VRF system operation used either limited field data covering certain building types and climate zones or used a questionnaire to obtain a larger dataset. However, they did not capture the wide applications of VRF systems quantitatively across all building types, climate zones, and operating conditions. To fill this gap, statistical and clustering analysis was conducted on the newly proposed key performance indicators of approximately 287,000 VRF systems for residential and commercial buildings in all five climate zones in China. In this work, the main findings are: (1) VRF systems are mainly used for cooling in all climate zones in China; (2) among all building types, the duration of use is lowest in residential buildings and highest in hotels and medical buildings; (3) the distribution of the ideal VRF cooling coefficient of performance (COP) is similar across all climate zones and building types; whereas the COPs of ideal VRF heating in the Severe Cold region and Cold regions are lower than those in other climate zones; and (4) partial load operations for VRF systems are common in residential buildings and office buildings due to the part-time-part-space operation mode. These findings can inform the actual application of VRF systems in China, supporting the design, operation, industry standard development, and performance optimization of VRF systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Applications of Anomaly Detection and Precursor Identification in Airspace Operations

As we continue to advance the U.S. National Airspace into the next generation of air traffic, we face challenges in both increase in complexity, as well as, a significant growth in traffic volume. Addressing these challenges, while maintaining the same level of safety is an important application of data mining. Because of these significant shifts in airspace design and usage there is a need to identify current and emergent safety risks along with their potential precursors. In recent years NASA has made advancements in developing scalable methods to address this effort in the Big Data paradigm. Multiple kernel anomaly detection approaches have been employed on both surveillance radar data and flight operational quality assurance data to identify operationally significant safety risks. Additionally, events have been explored with a recently developed precursor identification tool to discover states that reveal an increased probability of a safety event. These tools can be used to discover emerging safety risks that may not be currently monitored, which allows for mitigation tactics to be employed and ultimately make the overall airspace safer. This talk will discuss an overview of these methods and a discussion of the findings.

anomaly detection↗

Western energy related overhead monitoring project. Phase 2: Summary

Assistance by NASA to EPA in the establishment and maintenance of a fully operational energy-related monitoring system included: (1) regional analysis applications based on LANDSAT and auxiliary data; (2) development of techniques for using aircraft MSS data to rapidly monitor site specific surface coal mine activities; and (3) registration of aircraft MSS data to a map base. The coal strip mines used in the site specific task were in Campbell County, Wyoming; Big Horn County, Montana; and the Navajo mine in San Juan County, New Mexico. The procedures and software used to accomplish these tasks are described.

Anderson, J. E.↗

GOOML Big Kahuna Forecast Modeling and Genetic Optimization Files

This submission includes example files associated with the Geothermal Operational Optimization using Machine Learning (GOOML) Big Kahuna fictional power plant, which uses synthetic data to model a fictional power plant. A forecast was produced using the GOOML data model framework and fictional input data, and a genetic optimization is included which determines optimal flash plant parameters. The inputs and outputs associated with the forecast and genetic optimization are included. The input and output files consist of data, configuration files, and plots. A link to the Physics-Guided Neural Networks (phygnn) GitHub repository is also included, which augments a traditional neural network loss function with a generic loss term that can be used to guide the neural network to learn physical or theoretical constraints. phygnn is used by the GOOML framework to help integrate its machine learning models into the relevant physics and engineering applications. Note that the data included in this submission are intended to provide a demonstration of GOOML's capabilities. Additional files that have not been released to the public are needed for users to run these models and reproduce these results. Units can be found in the readme data resource.

15 GEOTHERMAL ENERGY↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Bayesian Analysis of the Power Spectrum of the Cosmic Microwave Background

There is a wealth of cosmological information encoded in the spatial power spectrum of temperature anisotropies of the cosmic microwave background. The sky, when viewed in the microwave, is very uniform, with a nearly perfect blackbody spectrum at 2.7 degrees. Very small amplitude brightness fluctuations (to one part in a million!!) trace small density perturbations in the early universe (roughly 300,000 years after the Big Bang), which later grow through gravitational instability to the large-scale structure seen in redshift surveys... In this talk, I will discuss a Bayesian formulation of this problem; discuss a Gibbs sampling approach to numerically sampling from the Bayesian posterior, and the application of this approach to the first-year data from the Wilkinson Microwave Anisotropy Probe. I will also comment on recent algorithmic developments for this approach to be tractable for the even more massive data set to be returned from the Planck satellite.

Bayesian inference↗

Using the SPoRT POES/GOES Hybrid Product in OCONUS Forecasting

The SPoRT (Short-term Prediction and Research Transition) Program at the NASA/Marshall Space Flight Center has been providing unique NASA and NOAA data and techniques to partner Weather Forecast Offices (WFOs) for ten years. Data are provided in the Decision Support System used by WFO forecasters: AWIPS. For the last couple of years, SPoRT has been producing the POES/GOES Hybrid. This suite of products combines the strength ofl5- minute animations of GOES imagery - providing temporal continuity, with the higher resolution, relatively random availability, of polar orbiting (POES) imagery data. The product was first introduced with only MODIS data from NASA's Terra and Aqua satellites, but recently the VIIRS instrument onboard the Suomi-NPP satellite was added, providing better high-resolution coverage. These products represent SPoRT's efforts to prepare for higher resolution, higher frequency GOES-R imagery - as well as helping to move VIIRS (JPSS) data into the mainstream of weather forecasting. SPoRT generates 5 products for this dataset: Visible, Longwave Infrared (11 micrometers), Shortwave IR (3.7 micrometers), Water Vapor (6.7 micrometers), and Fog (Difference of 11 micrometer and 3.7 micrometer channels). The Water Vapor hybrid product has a Red-Blue-Green image from MODIS inlaid, since it provides even more qualitative information than water vapor alone. Animated examples of the products will be shown in this presentation. While the resolution at nadir of GOES imagery is nominally Han (4km for IR channels), the inlaid polar orbiter imagery has a resolution of 250m (lkm for IR channels). This has tremendous application in the continental US. However, in high latitudes, since the usefulness of GOES degrades poleward rapidly, the contrast of GOES and POES data is stark. The consistent temporal nature of GOES, even though at a reduced resolution at high latitudes, provides basic situational awareness, but the introduction of polar data is very helpful in seeing the big picture with clarity - even if only briefly. This presentation will offer real situations where these products helped forecasters make better informed decisions quickly. Plans to augment the product further with the addition of data from several A VHRR instruments will be described.

Smith, Matt↗

Machine Intelligence for Radiation Science: Summary of the Radiation Research Society 67th Annual Meeting Symposium

The era of high-throughput techniques created big data in the medical field and research disciplines. Machine intelligence (MI) approaches can overcome critical limitations on how those large-scale data sets are processed, analyzed, and interpreted. The 67 th Annual Meeting of the Radiation Research Society featured a symposium on MI approaches to highlight recent advancements in the radiation sciences and their clinical applications. This article summarizes three of those presentations regarding recent developments for metadata processing and ontological formalization, data mining for radiation outcomes in pediatric oncology, and imaging in lung cancer.

radiation↗

Near 40 Years MERRA-2 Data at NASA GES DISC -Opportunity and Challenge to Support Extremes Study

To the end of 2019, 40 years NASA climate reanalysis data sets from the Modern Era Retrospective-analysis for Research and Applications, Version 2 (MERRA-2) will be available at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). MERRA-2 consists atmosphere, land, and ocean data, which may be used for the studies ranging from the short scale weather events to the large scale decadal vulnerabilities. The hourly products, such as precipitation, soil moisture, temperature, and aerosols etc., have been used widely to study extreme events.In supporting users from broad communities, GES DISC have developed various data access services, including subsetter for downloading only data of interest with preferred format; OPeNDAP - for machine-to-machine data access; and Giovanni- for online visualization and analysis, etc. A big challenge for extreme study is to downloading and processing long-term hourly or daily data. The data downloading performance is not very satisfied by many users with current services and the native archived data structure. Late June 2019, many people in Europe had experienced extreme heat waves. The temperatures in several countries exceeded 40°C (104°F). For example, MERRA-2 shows that the near surface daily maximum temperature of June 28 2019 over Marseille, a city in southern France, reached 41.1 °C (106°F), which is the record breaking temperature in the last 40 years. GES DISC is working together with domain science experts to improve the performance of long time series access, making analysis ready data sets in supporting application researches, such as extreme study. In this presentation, using Europe heat wave as an example, we will show prototype of the in developing service for finding extremes from near 40 years MERRA-2 data at a given location. MERRA-2 data can be accessed from NASA GES DISC(https://disc.gsfc.nasa.gov/ ) by search keyword "MERRA-2".

Shen, Suhung↗

Open Source Application of Fusing Aerosol Products from GEO and LEO Satellites

Retrieving aerosol optical depths (AODs) from sun-synchronous polar orbiting (aka low earth orbit, LEO) satellites, such as MODISs, and VIIRSs, OMI, TROPOMI, etc, has become well-established as a tool for extracting information on particulate matter (PM) and related processes in the atmosphere. However, with recently launched geostationary satellites (GEO), such as GOES-16/17/18, and Himawari-8/9, and Meteosat Third Generation (MTG) they provide a much higher temporal resolution (order of 10 minutes), typically an image once or more per hour during daylight compared to LEO once per day. By combining these observations, we may be able to characterize the diurnal cycle of global AOD at the local, regional and global scale. While the science community is still exploring the new data from GEO observations, we have been thinking about how to properly combine/merge/fuse those data considering differences in their spatial and temporal resolutions. However, this poses a “Big Data” challenge. The big data challenge is not just about data storage, but also about data discoverability, and accessibility, and even more, about data migration/mirroring in the cloud-computing environment. This paper is merely showing some of the efforts and approaches we have attempted in fusing six satellites’ Level 2 aerosol data (three are from GEO (GOES-16/17 and Himawari-8), and the other three are from LEO (TERRA/MODIS, AQUA/MODIS, SNPP-VIIRS) from Dark Target (DT) aerosol retrieval algorithm. Having the on-demand capability of fusing remote sensing products onto the desired temporal and spatial domain enables researchers and application practitioners to better manipulate and work with satellite and sensor data. It is our hopeWe hope that by making such an open-source package, and the accompanying functionality, the scientific community will be granted easier access to aerosol data processing resources. The MEaSUREs Program (Making Earth System Data Records for Use in Research Environments) expands our understanding of the Earth's current system through atmospheric and surface measurements. In an effort to aid the scientific research component and improve open source methods, this project developed Python code for fusing six satellite Level 2 aerosol data (three are from geostationary satellites (GEO), and the other three are from low earth orbital satellites (LEO)) from Dark Target Aerosol Retrieval Algorithm.

Jennifer Wei↗

Raman Microscopy Analysis of Wyoming CarbonSAFE Pilot Well Thin Sections for Mineralogy and Organic Matter Characterization

Scanning confocal Raman microspectroscopy (RMS) was used to analyze thin sections for inorganic mineral content and for evaluation of organic material (OM). Thin sections were selected from a stratigraphic test well core in the Powder River Basin, Wyoming under the Wyoming CarbonSAFE project. The well was analyzed with a variety of rock and fluid characterization techniques to determine the feasibility of a commercial-scale CO2 storage site. RMS complements other analyses, including traditional petrography, SEM, porosity and permeability, and XRD For this aspect of the study, RMS is especially important in the evaluation of OM in sealing lithologies. At prospective geologic CO2 storage sites, it is imperative to assess the unconventional oil and gas potential of seals to ensure that any future development would not compromise the integrity of the seals. Whole-slide mineralogical surveys were performed on thin sections from various shale and sandstone formations. Surveys were analyzed with Direct Classical Least Squares to identify and quantify minerals. The location and concentration of minerals was color-coded and overlaid on optical images for visualization of the distribution of minerals. Dense hyperspectral Raman mapping of OM was performed on five thin sections. Eleven spectral parameters diagnostic of organic type and thermal maturity were used to train a Partial Least Squares (PLS) calibration against a set of artificially matured samples spanning the pre- to mid-oil window. The PLS was applied to the study set and a post-mature set. Additionally, the PLS was applied to each point in hyperspectral maps for visualization of trends in maturity across sample sets and discrimination of organic matter types within a given map. In inorganic surveys on thin sections, a total of 14 unique inorganic minerals were identified in Raman spectra including quartz, dolomite, calcite, hematite and anhydrite. Shale thin sections tended to be dominated by organic material. OM was often observed mixed with inorganic minerals. Sand- and mudstones were dominated by inorganic minerals. The PLS extrapolated the post-mature set to reflectances >1.2%. The study set ranged from very immature to postmature in the median of map fit-peak parameters. However, point maturity maps indicate that matrix OM in all study samples is immature and that discreet organic particles selected for mapping, which may be inertinites, bias medians towards more-mature. The work demonstrates the capabilities of RMS to perform both whole-slide mineralogy and OM analysis with applications to formation evaluation in oil & gas, carbon sequestration and mining. Here, analysis of sealing formations in the well indicates high levels of immature organic matter that would not be a viable target for future oil production that could compromise the CO2 storage site. The work brings together diverse disciplines from geology and petrography to analytical chemistry, big data and microscopy.

Myers, Grant↗

Air Traffic Management TestBed Simulation Architect: User's Guide

The Air Traffic Management (ATM) TestBed is a Platform as a Service that is being developed by the National Aeronautics and Space Administration (NASA) to help design, configure, integrate, run, and monitor air traffic simulations. The platform provides cloud services including back-end big-data analytics tools, on-demand computing resource management, data storage, and communication middleware. The ATM TestBed reduces the time to test concepts and technologies, supports interactions among various concepts such as human-in-the-loop and automation-in-the-loop simulations, and enables collaborative simulations by sharing technologies and tools in the ATM community. The Simulation Architect application provides a graphical user interface tool for designing traffic scenarios and simulations using blocks representing components and links representing message channels linking them. This guide describes a high-level user interface design of Simulation Architect and provides information for a new user to compose traffic scenarios and simulations.

Software User Guide↗