Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Virtual Log-Structured Storage for High-Performance Streaming

Over the past decade, given the higher number of data sources (e.g., Cloud applications, Internet of things) and critical business demands, Big Data transitioned from batch-oriented to real-time analytics. Stream storage systems, such as Apache Kafka, are well known for their increasing role in real-time Big Data analytics. For scalable stream data ingestion and processing, they logically split a data stream topic into multiple partitions. Stream storage systems keep multiple data stream copies to protect against data loss while implementing a stream partition as a replicated log. This architectural choice enables simplified development while trading cluster size with performance and the number of streams optimally managed. This paper introduces a shared virtual log-structured storage approach for improving the cluster throughput when multiple producers and consumers write and consume in parallel data streams. Stream partitions are associated with shared replicated virtual logs transparently to the user, effectively separating the implementation of stream partitioning (and data ordering) from data replication (and durability). We implement the virtual log technique in the KerA stream storage system. When comparing with Apache Kafka, KerA improves the cluster ingestion throughput by up to 4x when multiple producers write over hundreds of data streams.

consistent stream ordering↗

Genome-Scale Metabolic Modeling Enables In-Depth Understanding of Big Data

Genome-scale metabolic models (GEMs) enable the mathematical simulation of the metabolism of archaea, bacteria, and eukaryotic organisms. GEMs quantitatively define a relationship between genotype and phenotype by contextualizing different types of Big Data (e.g., genomics, metabolomics, and transcriptomics). In this review, we analyze the available Big Data useful for metabolic modeling and compile the available GEM reconstruction tools that integrate Big Data. We also discuss recent applications in industry and research that include predicting phenotypes, elucidating metabolic pathways, producing industry-relevant chemicals, identifying drug targets, and generating knowledge to better understand host-associated diseases. In addition to the up-to-date review of GEMs currently available, we assessed a plethora of tools for developing new GEMs that include macromolecular expression and dynamic resolution. Finally, we provide a perspective in emerging areas, such as annotation, data managing, and machine learning, in which GEMs will play a key role in the further utilization of Big Data.

59 BASIC BIOLOGICAL SCIENCES↗

4th Big Data for Nuclear Power Plants Workshop 2023

The Ohio State University and Idaho National Laboratory organized the 4 th Big Data for Nuclear Power Plants Workshop in November, 2023 in Columbus, Ohio. Workshop topics were chosen to understand the challenges and gaps that need to be addressed to maximize the impact of data on the nuclear industry, as well as the associated applications and risks. Discussions were focused around six specific application areas: Operation and Maintenance; Machine Learning in Nuclear Materials and Advanced Manufacturing; Cybersecurity; High-Performance Computing and Massive Computation; Big Data and Digital Twins; and Nuclear Non-Proliferation. The opportunities, challenges, and risks identified in the six focus areas explored in this workshop are diverse, but some common themes emerge, such as the importance of data integrity, quality, coverage, privacy, and traceability. Big data and AI/ML tools can be leveraged to reduce costs, optimize human tasking, and reduce human error across various application areas. In order for the nuclear industry to benefit from big data and advanced analytic capabilities, it is essential to address challenges and risks, such as data privacy, model reliability, and computational resource availability. Learning from other industries that have successfully implemented big data and AI/ML technologies, like the aerospace industry, can help the nuclear industry successfully integrate these technologies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Nuclear Data: What is the big deal? [Slides]

This presentation begins with a look at the practical application: differential and integral data. It also includes a nuclear data evaluation: cross section formalisms, data processing and cross section library generation, computer tools, and criticality safety assessment: nuclear data. The presentation ends with some concluding remarks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Use of Schema on Read in Earth Science Data Archives

Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.

cloud applications↗

Development of a Unified Taxonomy for HVAC System Faults

Detecting and diagnosing HVAC faults is critical for maintaining building operation performance, reducing energy waste, and ensuring indoor comfort. An increasing deployment of commercial fault detection and diagnostics (FDD) software tools in commercial buildings in the past decade has significantly increased buildings’ operational reliability and reduced energy consumption. A massive amount of data has been generated by the FDD software tools. However, efficiently utilizing FDD data for ‘big data’ analytics, algorithm improvement, and other data-driven applications is challenging because the format and naming conventions of those data are very customized, unstructured, and hard to interpret. This paper presents the development of a unified taxonomy for HVAC faults. A taxonomy is an orderly classification of HVAC faults according to their characteristics and causal relations. The taxonomy includes fault categorization, physical hierarchy, fault library, relation model, and naming/tagging scheme. The taxonomy employs both a physical hierarchy of HVAC equipment and a cause-effect relationship model to reveal the root causes of faults in HVAC systems. A structured and standardized vocabulary library is developed to increase data representability and interpretability. The developed fault taxonomy can be used for HVAC system ‘big data’ analytics such as HVAC system fault prevalence analysis or the development of an HVAC FDD software standard. A common type of HVAC equipment-packaged rooftop unit (RTU) is used as an example to demonstrate the application of the developed fault taxonomy. Two RTU FDD software tools are used to show that after mapping FDD data according to the taxonomy, the meta-analysis of the multiple FDD reports is possible and efficient.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

On Performance Prediction of Big Data Transfer in High-performance Networks

Big data generated by large-scale scientific and industrial applications need to be transferred between different geographical locations for remote storage, processing, and analysis. High-speed dedicated connections provisioned in High-performance Networks (HPNs) are increasingly utilized to carry out such big data transfer. HPN management highly relies on an important capability of performance (mainly throughput) prediction to reserve sufficient bandwidth and meanwhile avoid over-provisioning that may result in unnecessary resource waste. This capability is critical to improving the resource (mainly bandwidth) utilization of dedicated connections and meeting various user requests for data transfer. Conventional methods conduct performance prediction by fitting prior observed transfer history with predefined loss functions, without considering unobservable latent factors such as competing loads on end hosts. Such latent factors also have a significant impact on the application-level data transfer performance, which may result in an inaccurate prediction model. In this paper, we first investigate the impact of latent factors and propose a clustering-based method to eliminate their negative impact on performance prediction. We then develop a robust machine learning-based performance predictor by: i) incorporating the proposed latent factor elimination method into data preprocessing, and ii) adopting a customized domain guided loss function. Extensive experimental results show that our predictor achieves significantly higher prediction accuracy than several other state-of-the-art methods.

Liu, Wuji↗

RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific Data

In modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead.In this paper, we propose RAPIDS, a hybrid approach that combines the multigrid-based error-bounded lossy compression with erasure coding, to significantly reduce the storage and network overhead required for maintaining high data availability. Our experiments show that RAPIDS reduces the storage overhead by up to 7.5x and network overhead by up to 3x to achieve the same level of availability compared to the regular EC method. We improve RAPIDS by building two models to optimize the fault tolerance configurations and data gathering strategy. We demonstrate that RAPIDS significantly improves performance when running on many CPU cores in parallel or on GPUs.

Wan, Lipeng↗

TURBO: Terabits/s Using Reconfigurable Bandwidth Optics (Final Report)

Large scale, big science applications are generating petabytes of data that need to be shared among multiple locations. Moving a petabyte of data at 100 Gb/s takes 22 hours. Reducing this to single digit hours or below is clearly attractive and with future applications anticipating exabyte scale transfers, multi-terabit per second capacities are needed. This scale of capacity is available using optical systems. Core optical systems have aggregate per fiber capacities on the order of 10 Tb/s and more than 100 Tb/s system experiments have been realized in the lab. But large optical capacities are not just reserved for the core as similar system capacities are available in metro networks right up to the enterprise, campus, or data center. In general, the same technology is used today in the metro area as in the long-haul networks—reconfigurable optical add drop multiplexing (ROADM) node-based wavelength division multiplexing (WDM) systems. Data centers already support massive internal capacities on the scale of 100’s of Tb/s. As new photonic integrated technologies mature, for example through initiatives such as the Integrated Photonics Innovative Manufacturing Institute (IP-IMI), the cost of optical interfaces is expected to decrease, and higher capacity and higher performance interfaces will become more affordable. This set of circumstances creates the potential that Tb/s capacities will be available at the enterprise and campus level to support large scale science network applications.

42 ENGINEERING↗

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)↗

Exploratory analysis and performance prediction of big data transfer in High-performance Networks

Big data transfer in large-scale scientific and business applications is increasingly carried out over connections with guaranteed bandwidth provisioned in High-performance Networks (HPNs) via advance bandwidth reservation. Provisioning agents need to carefully schedule data transfer requests, compute network paths, and allocate appropriate bandwidths. Such reserved bandwidths, if not fully utilized, could be simply wasted due to the exclusive access during the approved time window, and cause extra overhead and complexity for resource management. This calls for accurate performance prediction to reserve bandwidths that match actual needs and avoid over-provisioning. We employ machine learning algorithms to predict big data transfer performance based on extensive performance measurements collected in the past several years from data transfer tests using different protocols and toolkits between various end sites on several real-life physical or emulated testbeds. We first analyze the performance patterns in response to a comprehensive list of parameters in end-host systems, network connections, and data transfer applications, which motivate the use of machine learning and also help us identify the effects of latent factors. We then propose threshold- and clustering-based methods to eliminate negative effects of latent factors in data preprocessing and build a robust performance predictor based on customized domain-oriented loss functions. The performance of the proposed methods is verified by extensive experiments using SVR and RFR as well as theoretical analysis of the general performance bound.

97 MATHEMATICS AND COMPUTING↗

NASA’s Prototype Spectral Water Inversion Processor and Emulator (SWIPE): Towards Global Coastal and Inland Water Quality and Algal Biodiversity Monitoring

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will provide updates on NASA’s prototype open-source aquatic modeling platform, Spectral Water Inversion Processor and Emulator (SWIPE), which is a comprehensive, multi-faceted modeling platform for both forward and inverse modeling of diverse aquatic ecosystems from the benthos to top-of-atmosphere (TOA). SWIPE provides a cohesive application which leverages recent advancements in particle modeling, Big Data analytics, and machine learning to develop a high-fidelity synthetic training ground for sensitivity studies and algorithm development for multispectral or upcoming hyperspectral missions. Some of the prominent features of SWIPE to be discussed include: 1. Advanced hyperspectral modeling of globally diverse algal and non-algal particles using a novel two-layer coated sphere scattering model and radiative transfer modeling, 2. Massive, highly detailed synthetic spectral libraries of Analysis-Ready-Data (ARD) which include spectral libraries of particle microphysics, water biogeophysical and optical properties, as well as surface and TOA reflectances at 1 nm resolution, 3. An ensemble of pre-built analytic, machine learning, and deep learning inversion algorithms for various water quality and biodiversity related retrieval parameters and uncertainty quantification, 4. Sensor-agnostic water quality inversion at wide ranging spatial and spectral resolutions including a codebase for seamless application in the Google Earth Engine and NASA Earth Exchange (NEX) for planetary scale analysis. SWIPE will be a fully open-source platform based in python with comprehensive documentation, tutorials, and options for distributed computing on high performance computing clusters or on single, local machines. Further, we will discuss how we envision SWIPE contributing towards a global analysis of coastal and inland water quality dynamics.

top-of-atmosphere (TOA)↗

A Survey on Error-Bounded Lossy Compression for Scientific Datasets

Error-bounded lossy compression has been effective in significantly reducing the data storage/transfer burden while preserving the reconstructed data fidelity very well. Many error-bounded lossy compressors have been developed for a wide range of parallel and distributed use cases for years. They are designed with distinct compression models and principles, such that each of them features particular pros and cons. In this article, we provide a comprehensive survey of emerging error-bounded lossy compression techniques. The key contribution is fourfold. (1) We summarize a novel taxonomy of lossy compression into six classic models. (2) We provide a comprehensive survey of 10 commonly used compression components/modules. (3) We summarized pros and cons of 47 state-of-the-art lossy compressors and present how state-of-the-art compressors are designed based on different compression techniques. (4) We discuss how customized compressors are designed for specific scientific applications and use-cases. We believe this survey is useful to multiple communities including scientific applications, high-performance computing, lossy compression, and big data.

Error-Bounded Lossy Compression↗

Functional Data Analysis for Extracting the Intrinsic Dimensionality of Spectra: Application to Chemical Homogeneity in the Open Cluster M67

High-resolution spectroscopic surveys of the Milky Way have entered the Big Data regime and have opened avenues for solving outstanding questions in Galactic archeology. However, exploiting their full potential is limited by complex systematics, whose characterization has not received much attention in modern spectroscopic analyses. In this work, we present a novel method to disentangle the component of spectral data space intrinsic to the stars from that due to systematics. Using functional principal component analysis on a sample of 18,933 giant spectra from APOGEE, we find that the intrinsic structure above the level of observational uncertainties requires ≈10 functional principal components (FPCs). Our FPCs can reduce the dimensionality of spectra, remove systematics, and impute masked wavelengths, thereby enabling accurate studies of stellar populations. To demonstrate the applicability of our FPCs, we use them to infer stellar parameters and abundances of 28 giants in the open cluster M67. We employ Sequential Neural Likelihood, a simulation-based Bayesian inference method that learns likelihood functions using neural density estimators, to incorporate non-Gaussian effects in spectral likelihoods. By hierarchically combining the inferred abundances, we limit the spread of the following elements in M67: Fe ≲ 0.02 dex; C ≲ 0.03 dex; O, Mg, Si, Ni ≲ 0.04 dex; Ca ≲ 0.05 dex; N, Al ≲ 0.07 dex (at 68% confidence). Our constraints suggest a lack of self-pollution by core-collapse supernovae in M67, which has promising implications for the future of chemical tagging to understand the star formation history and dynamical evolution of the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

Explainable and trustworthy artificial intelligence for correctable modeling in chemical sciences

Data science has primarily focused on big data, but for many physics, chemistry, and engineering applications, data are often small, correlated and, thus, low dimensional, and sourced from both computations and experiments with various levels of noise. Typical statistics and machine learning methods do not work for these cases. Expert knowledge is essential, but a systematic framework for incorporating it into physics-based models under uncertainty is lacking. Here, we develop a mathematical and computational framework for probabilistic artificial intelligence (AI)–based predictive modeling combining data, expert knowledge, multiscale models, and information theory through uncertainty quantification and probabilistic graphical models (PGMs). We apply PGMs to chemistry specifically and develop predictive guarantees for PGMs generally. Our proposed framework, combining AI and uncertainty quantification, provides explainable results leading to correctable and, eventually, trustworthy models. The proposed framework is demonstrated on a microkinetic model of the oxygen reduction reaction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Restructuring Big Data to Improve Data Access and Performance in Analytic Services Making Research More Efficient for the Study of Extreme Weather Events and Application User Communities

By developing and enhancing various services and tools, the GES DISC provides users with the capability to access and visualize data, and to make comparisons of data from multiple sensor and models via a number of cross-discipline projects. Discovering Data via Faceted Web Interface Web interface to data products and services Search and Download mechanisms Dataset Landing Pages Accessing Data through Interoperable Services: GDS – GrADS Data Server OPeNDAP - Open-source Project for a Network Data Access Protocol WMS – OGC service GIS connector – allowing IS tools to access data easier (coming soon) HTTPS -- direct online access Downloading Data Basics: Subset and egridding Service – Parameter, Spatial, Time, Vertical, Mean averaging, format conversion, and regridding for L3/L4 gridded data Swath Data Subsetter – Parameter, spatial subset of L2 /L1 data. Visualizing Data Online: Giovanni –Visualization and Analysis L3/L4 gridded data AIRS NRT Viewer – AIRS near-real-time DQVis – L2 data quality visualization

data cube↗

High Resolution Nature Runs and the Big Data Challenge

NASA's Global Modeling and Assimilation Office at Goddard Space Flight Center is undertaking a series of very computationally intensive Nature Runs and a downscaled reanalysis. The nature runs use the GEOS-5 as an Atmospheric General Circulation Model (AGCM) while the reanalysis uses the GEOS-5 in Data Assimilation mode. This paper will present computational challenges from three runs, two of which are AGCM and one is downscaled reanalysis using the full DAS. The nature runs will be completed at two surface grid resolutions, 7 and 3 kilometers and 72 vertical levels. The 7 km run spanned 2 years (2005-2006) and produced 4 PB of data while the 3 km run will span one year and generate 4 BP of data. The downscaled reanalysis (MERRA-II Modern-Era Reanalysis for Research and Applications) will cover 15 years and generate 1 PB of data. Our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS), a specialization of the concept of business process-as-a-service that is an evolving extension of IaaS, PaaS, and SaaS enabled by cloud computing. In this presentation, we will describe two projects that demonstrate this shift. MERRA Analytic Services (MERRA/AS) is an example of cloud-enabled CAaaS. MERRA/AS enables MapReduce analytics over MERRA reanalysis data collection by bringing together the high-performance computing, scalable data management, and a domain-specific climate data services API. NASA's High-Performance Science Cloud (HPSC) is an example of the type of compute-storage fabric required to support CAaaS. The HPSC comprises a high speed Infinib and network, high performance file systems and object storage, and a virtual system environments specific for data intensive, science applications. These technologies are providing a new tier in the data and analytic services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. In our experience, CAaaS lowers the barriers and risk to organizational change, fosters innovation and experimentation, and provides the agility required to meet our customers' increasing and changing needs

big data analysis↗