Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Where Is the Provenance? Ethical Replicability and Reproducibility in GIScience and Its Critical Applications

As replicability and reproducibility (R&R) crises develop within emerging convergent inquiry, ethical use of provenance information is central to the establishment and preservation of trust in critical applications of GIScience and geospatial technologies. Today large volumes of geospatial data are generated at high velocity from satellite sensors and unmanned aircraft systems, citizen sensors, geolocation-based data services, global navigation satellite systems, and so on. The extensive use of these data for applications such as disaster and humanitarian response raises the issue of R&R from competing perspectives of location privacy and geospatial data quality. Although geospatial data can be integrated and linked with contextual information to identify individuals’ movements, steps taken to ensure privacy can complicate the multiuser development of high-quality geospatial workflows. Provenance information as digital records of historical (retrospective) and potential future (prospective) geospatial processes is often overlooked, misunderstood, or inadequately addressed. We explore the relationship between provenance information, location privacy, and geospatial data quality in the context of R&R with a focus on disaster analytics. Here, we argue that in the era of big data and deep learning, GIScientists and associated institutions bear greater responsibility both for geospatial workflow quality and for location privacy. Given vastly heterogenous computational landscapes, we provide practical recommendations for ethically driven provenance and R&R research and development within the GIScience community and beyond.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints

Vast volumes of data are produced by today's scientific simulations and advanced instruments. These data cannot be stored and transferred efficiently because of limited I/O bandwidth, network speed, and storage capacity. Error-bounded lossy compression can be an effective method for addressing these issues: not only can it significantly reduce data size, but it can also control the data distortion based on user-defined error bounds. In practice, many scientific applications have specific requirements or constraints for lossy compression, in order to guarantee that the reconstructed data are valid for post hoc analysis. For example, some datasets contain irrelevant data that should be isolated in particular and users often have intuition regarding value ranges, geospatial regions, and other data subsets that are crucial for subsequent analysis. Existing state-of-the-art error-bounded lossy compressors, however, do not consider these constraints during compression, resulting in inferior compression ratios with respect to user's post hoc analysis, due to the fact that the data itself provides little or no value for post hoc analysis. In this work we address this issue by proposing an optimized framework that can preserve diverse constraints during the error-bounded lossy compression, e.g., cleaning the irrelevant data, efficiently preserving different precision for multiple value intervals, and allowing users to set diverse precision over both regular and irregular regions. We perform our evaluation on a supercomputer with up to 2,100 cores. Experiments with six real-world applications show that our proposed diverse constraints based error-bounded lossy compressor can obtain a higher visual quality or data fidelity on reconstructed data with the same or even higher compression ratios compared with the traditional state-of-the-art compressor SZ. Furthermore, our experiments also demonstrate very good scalability in compression performance compared with the I/O throughput of the parallel file system.

97 MATHEMATICS AND COMPUTING↗

Physics-informed machine learning

Despite great progress in simulating multiphysics problems using the numerical discretization of partial differential equations (PDEs), one still cannot seamlessly incorporate noisy data into existing algorithms, mesh generation remains complex, and high-dimensional problems governed by parameterized PDEs cannot be tackled. Moreover, solving inverse problems with hidden physics is often prohibitively expensive and requires different formulations and elaborate computer codes. Machine learning has emerged as a promising alternative, but training deep neural networks requires big data, not always available for scientific problems. Instead, such networks can be trained from additional information obtained by enforcing the physical laws (for example, at random points in the continuous space-time domain). Such physics-informed learning integrates (noisy) data and mathematical models, and implements them through neural networks or other kernel-based regression networks. Moreover, it may be possible to design specialized network architectures that automatically satisfy some of the physical invariants for better accuracy, faster training and improved generalization. Furthermore, we review some of the prevailing trends in embedding physics into machine learning, present some of the current capabilities and limitations and discuss diverse applications of physics-informed learning both for forward and inverse problems, including discovering hidden physics and tackling high-dimensional problems.

97 MATHEMATICS AND COMPUTING↗

Evaluating E3SM Global Storm‐Resolving Model Simulations of Deep Convection: Insights From DP‐SCREAM During TRACER

Global Storm-Resolving Models (GSRMs) are becoming increasingly vital for advancing climate modeling and improving the prediction of extreme weather events. Houston, a coastal region frequently affected by deep convective storms, offers an ideal setting to evaluate the ability of GSRMs to simulate deep convection. This study assesses the performance of the Doubly Periodic Simple Cloud-Resolving E3SM (Energy Exascale Earth System Model) Atmosphere Model (DP-SCREAM) using observations from the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign. DP-SCREAM effectively reproduces the diurnal cycles of clouds and precipitation, demonstrating much greater skill than the E3SM single column model. The DP-SCREAM is demonstrated to be applicable to coastal regions, partially due to the forcing data sets already capturing the influence of breezes. DP-SCREAM also replicates biases persistent in the global version of SCREAM: the underrepresentation of boundary layer shallow clouds, a lack of mid-level congestus clouds, and the popcorn convection, characterized by small and disorganized convective cells generating the strongest precipitation. To investigate these issues, two sensitivity experiments were conducted: increasing the mixing length and scaling up the buoyancy flux within the Simplified Higher Order Closure scheme. Increasing the mixing length improved mid-level congestus representation and reduced unrealistic early morning fog occurrence. Enhancing buoyancy flux only marginally improved the bias of underproduced big convective cells. In conclusion, an additional resolution sensitivity test at 0.5 km grid spacing demonstrated that a refined horizontal resolution alone is insufficient to resolve these biases.

54 ENVIRONMENTAL SCIENCES↗

Enabling discovery data science through cross-facility workflows

Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.

Antypas, Katerina B.↗

Train small, model big: Scalable physics simulators via reduced order modeling and domain decomposition

Numerous cutting-edge scientific technologies originate at the laboratory scale, but transitioning them to practical industry applications is a formidable challenge. Traditional pilot projects at intermediate scales are costly and time-consuming. An alternative, the pilot-scale model, relies on high-fidelity numerical simulations, but even these simulations can be computationally prohibitive at larger scales. To overcome these limitations, we propose a scalable, physics-constrained reduced order model (ROM) method. The ROM identifies critical physics modes from small-scale unit components, projecting governing equations onto these modes to create a reduced model that retains essential physics details. We also employ Discontinuous Galerkin Domain Decomposition (DG-DD) to apply ROM to unit components and interfaces, enabling the construction of large-scale global systems without data at such large scales. Here this method is demonstrated on the Poisson and Stokes flow equations, showing that it can solve equations about 15–40 times faster with only ~1% relative error. Furthermore, ROM takes one order of magnitude less memory than the full order model, enabling larger scale predictions at a given memory limitation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Multi-Branch Decoder Network Approach to Adaptive Temporal Data Selection and Reconstruction for Big Scientific Simulation Data

A key challenge in scientific simulation is that the simulation outputs often require intensive I/O and storage space to store the results for effective post hoc analysis. This article focuses on a quality-aware adaptive temporal data selection and reconstruction problem where the goal is to adaptively select simulation data samples at certain key timesteps in situ and reconstruct the discarded samples with quality assurance during post hoc analysis. This problem is motivated by the limitation of current solutions that a significant amount of simulation data samples are either discarded or aggregated during the sampling process, leading to inaccurate modeling of the simulated phenomena. Two unique challenges exist: 1) the sampling decisions have to be made in situ and adapted to the dynamics of the complex scientific simulation data; 2) the reconstruction error must be strictly bounded to meet the application requirement. To address the above challenges, we develop DeepSample , an error-controlled convolutional neural network framework, that jointly integrates a set of coherent multi-branch deep decoders to effectively reconstruct the simulation data with rigorous quality assurance. The results on two real-world scientific simulation applications show that DeepSample significantly outperforms other state-of-the-art methods on both sampling efficiency and reconstructed simulation data quality.

Zhang, Yang↗

W2VPCA: A Machine Learning Method for Measuring Attitudes With Natural Language

Company strategy influences many decisions in freight transportation. Behavioral models of company decision-making therefore could benefit from including strategy variables. However, strategy is difficult to observe and quantify. Attitudinal surveys of company executives can be used to collect measurements of latent strategy to use in quantitative models. However, surveys are costly and burdensome. Text mining methods to collect measurements overcome these issues somewhat, but typically require manual intervention and ignore the context of words, which can be problematic. This study introduces a new machine learning method to generate strategy measurement data from existing big text data. The new method, called W2VPCA, combines Natural Language Processing and Principal Components Analysis. W2VPCA produces measurement data that serve as quantitative indicators of latent strategy in behavioral models. W2VPCA is unsupervised, data-driven, and uses information on word context. We apply W2VPCA to generate measurements of latent strategies using readily available, large-scale text data: annual company reports. The empirical measurements are used successfully to associate two latent strategies, one focusing on distribution and the other on products, with truck fleet and distribution center outsourcing decisions. The main empirical outcome is that the W2VPCA measurements outperform Bag-of-Words measurements in a psychometric analysis of latent firm strategies. While this study focuses on freight behavioral models, W2VPCA may also have applications in behavioral modeling in other domains.

97 MATHEMATICS AND COMPUTING↗

Benchmarking Demand Flexibility in Commercial Buildings and Flattening the Duck – Addressing Baseline and Commissioning Challenges

With the transition from our traditional electric grid to a cleaner grid with renewable power generation, there is a need to enable building loads to be flexible. Load shedding and shifting will be essential for flattening the “Duck” for decarbonization. This paper explored the trend in the timing of DR events as a reflection of the grid’s needs using recent four years of event data from 203 retail stores in 11 states. The events are becoming significantly shorter with 2-hour duration being the most popular; shifting to late afternoon and early evening is another trend beyond California. Benchmarking will be essential for accounting DF as a reliable grid resource. This paper addresses a challenging aspect of benchmarking – inaccuracies in counterfactual baseline methods can introduce significant DF metrics variations in addition to weather and building characteristics related factors. The conventional “10/10” with adjustment baseline method has inherent limitation by design for load shifting applications. Therefore, it is imperative to identify alternative methods. This study compared three hourly regression baseline methods with “10/10” methods using two groups of commercial buildings that participated in DR programs: (1) 121 big-box retail stores, and (2) 11 office buildings in CA. The 14-day hourly outdoor temperature regression method was found to produce least error in the tested datasets and is promising for load shifting. The paper also pointed out that commissioning issues can also be a significant barrier for achieving consistent DF performance, which building managers and utilities should be aware of.

Liu, Jingjing↗

Mitigating Catastrophic Forgetting in Deep Learning in a Streaming Setting Using Historical Summary

Recent advancements in scientific equipment and the adaptation of electronics and the Internet of Things (IoT) in our everyday lives resulted in large and complex data production at a high rate. Making meaningful and timely knowledge discovery at a modest cost from this big data is difficult for computing power and storage limitations. Training deep learning models incrementally in a streaming setting can help us with overcoming these limitations. However, in a well-known phenomenon named catastrophic forgetting, incrementally trained models increasingly perform poorly on the past data. To mitigate catastrophic forgetting in training in a streaming setting, we propose constructing a historical summary over time and use the summary with newly arrived data during incremental training. We propose various data summarization techniques such as random sampling, micro clustering, coreset computation, and Auto Encoders to counteract catastrophic forgetting. We built a pipeline for incremental training with a historical summary for training deep learning models for streaming data. We demonstrate the effectiveness of historical summary in mitigating catastrophic forgetting using three case studies involving three different deep learning applications: an Artificial Neural Network (ANN) for classification task on MNIST dataset, a language model (RNN-LM) on the WikiText2 dataset, and a Convolutional Neural Network (CNN), ResNet50 to classify the ImageNet dataset. Through the training of the models, we observe that catastrophic forgetting is evident in ANN and CNN but not in an RNN. For the first task, our method recovers up to 47.9% lost accuracy due to catastrophic forgetting. For the third task, the historical summary recovers classification accuracy by up to 25%. For the second task, though there is not proof of catastrophic forgetting, the training performance (PPL) improves by up to 26% with historical summary.

Dash, Sajal↗

2021 GeoAI Workshop Report: The Trillion Pixel Challenge

The convergence of geospatial big data with advancements from artificial intelligence, cloud infrastructure, and high-performance computing continues to revolutionize mapping and analysis of Earth's surface in unprecedented detail. Rapid innovations in sensing technologies will soon collect geospatial data in even higher resolution and throughput. These developments offer the potential for breakthroughs in science, policy, and national security via end-to-end GeoAI systems that can provide fresh insights into how humans occupy and alter their environment over time. At the 2021 GeoAI Trillion Pixel workshop, international subject matter experts from government, academia, industry, and nonprofit organizations gathered virtually to discuss the Trillion Pixel GeoAI Challenge. The event focused on six major themes currently influencing scientific innovation and breakthroughs. Particular focus was paid to societal impacts. As an additional takeaway message, the gathering identified remaining application gaps and challenges that are in need of stronger community partnerships and collaborations.

58 GEOSCIENCES↗

Data-driven computational prediction and experimental realization of exotic perovskite-related polar magnets

Rational design of technologically important exotic perovskites is hampered by the insufficient geometrical descriptors and costly and extremely high-pressure synthesis, while the big-data driven compositional identification and precise prediction entangles full understanding of the possible polymorphs and complicated multidimensional calculations of the chemical and thermodynamic parameter space. Here we present a rapid systematic data-mining-driven approach to design exotic perovskites in a high-throughput and discovery speed of the A 2 BB ’O 6 family as exemplified in A 3 TeO 6 . The magnetoelectric polar magnet Co 3 TeO 6 , which is theoretically recognized and experimentally realized at 5 GPa from the six possible polymorphs, undergoes two magnetic transitions at 24 and 58 K and exhibits helical spin structure accompanied by magnetoelastic and magnetoelectric coupling. We expect the applied approach will accelerate the systematic and rapid discovery of new exotic perovskites in a high-throughput manner and can be extended to arbitrary applications in other families.

36 MATERIALS SCIENCE↗

Tsdat: An Open-Source Data Standardization Framework for Marine Energy and Beyond: Preprint

Many organizations are tasked with the collection and processing of large quantities of data from various measurement devices. Data reported from these sources are often not interoperable with datasets and software used by analysts and other organizations in the same field, introducing barriers for collaboration on large-scale projects. This poses a particular problem for cross-device comparisons and machine learning applications. To address these challenges, the open source Time-Series Data Pipelines (Tsdat) software was developed by a joint collaboration between Pacific Northwest National Laboratory, the National Renewable Energy Laboratory, and Sandia National Laboratories to facilitate collaboration and accelerate advancements in the Marine Energy domain through the development of an open-source ecosystem of tools. This paper will describe the Tsdat software and the data standards within which the framework operates. A beta version of the framework has been released and is currently being used by several projects in marine energy, wind energy, and building energy systems.

big data↗

US Department of Energy, Office of Science High Performance Computing Facility Operational Assessment 2019 Oak Ridge Leadership Computing Facility

Oak Ridge National Laboratory's (ORNL's) Leadership Computing Facility (OLCF) continues to surpass its operational target goals: supporting users; delivering fast, reliable computational ecosystems; creating innovative solutions for high performance computing (HPC) needs; and managing risks, safety, and security associated with operating some of the most powerful computers in the world. The results can be seen in the cutting-edge science conducted by users and the praise from the research community. Calendar year (CY) 2019 was a big year as OLCF staff ran five world-class resources (the leadershipclass computers Titan and Summit, the large analysis cluster called Eos, and the massive parallel filesystems called Atlas and Alpine)) and also began power and cooling upgrades for a 2021 exascale system called Frontier. While continuing exceptional operation of Titan, Eos, and Rhea, the OLCF released the Summit supercomputer for production on January 1, 2019. Summit debuted as the most capable and efficient system in its class and has been recognized as the most powerful system in the world for its performance on both the high performance linpack (HPL) and conjugate gradient (HPCG) benchmark applications since June 2018 according to TOP500. Summit represents the culmination of a multiyear effort between the OLCF, IBM, NVIDIA, and Mellanox to deliver a system that is unmatched for modeling, simulation, data analysis, and learning. To hit the ground running with science-ready applications on day one, application teams worked closely with the OLCF through the Center for Accelerated Application Readiness (CAAR) program for years in advance of the Summit deployment. CY 2019 was filled with outstanding results and accomplishments: a very high rating from users on overall satisfaction for the sixth year in a row; a tremendous amount of core-hours delivered to researchers from two leadership-class systems; and success in delivering on the allocation split of roughly 60%, 30%, and 10% of core-hours offered for the Innovative and Novel Computational Impact on Theory and Experiment (INCITE), Advanced Scientific Computing Research Leadership Computing Challenge (ALCC), and Director's Discretionary (DD) programs, respectively (see Operational Performance section). These accomplishments, coupled with the high utilization rates (overall and capability usage), represent the fulfillment of the promise of both leadership-class machines: efficient facilitation of leadership-class computational applications. Table ES.1 presents a summary of the 2019 OLCF metric targets and the associated results. More information can be found in the Operational Performance section for each OLCF resource. The scientific accomplishments of OLCF users are a strong indication of long-term operational success, with publications this year in such notable journals and publications as Nature, Nature Physics, Nature Plants, Physical Review X, Journal of the American Physical Society, Cell, Nano Letters, and Trends in Biotechnology. Crucial domain-specific discoveries facilitated by resources at the OLCF are described in the High Performance Computing Facility Operational Assessment 2019 Oak Ridge Leadership Computing Facility (OAR) Strategic Results section. For example, researchers used Summit to pinpoint and understand the production of proteins from genetic information, including mutations and the functional expression of disease (Section 8.2).

97 MATHEMATICS AND COMPUTING↗

Gaussian Process Regression for Aggregate Baseline Load Forecasting

Demand response (DR) is one of the most effective ways to maintain the reliability and improve the flexibility of power systems. Accurate forecasts of baseline loads are essential for DR programs. In the era of big data, machine learning-based approaches present a unique opportunity for baseline load forecasting. Thus, this paper presents a machine learning-based approach using a relatively less explored algorithm, Gaussian process regression (GPR), to forecast aggregate baseline loads. As such, a dataset was generated using a set of EnergyPlus simulations. Using the generated dataset, a GPR-based forecasting model was developed. In addition, support vector regression (SVR)-, artificial neural network (ANN)-, and averaging-based models were developed as baseline models for comparison. These models were compared in terms of accuracy, simplicity, and integrity. The prediction performance of the models showed that the GPR-based model is more accurate and reliable than the others. Such high performance shows the potential of the GPR in baseline load forecasting. GPR, therefore, can be used for DR applications.

Amasyali, Kadir↗

Tsdat: An Open-Source Data Standardization Framework for Marine Energy and Beyond

Many organizations are tasked with the collection and processing of large quantities of data from various measurement devices. Data reported from these sources are often not interoperable with datasets and software used by analysts and other organizations in the same domain, introducing barriers for collaboration on large-scale projects. This poses a particular problem for cross-device comparisons and machine learning applications, which rely on large quantities of data from multiple sources. To address these challenges, the open-source Time-Series Data Pipelines (Tsdat) Python framework was developed by Pacific Northwest National Laboratory, with strategic guidance and direction provided by the National Renewable Energy Laboratory and Sandia National Laboratories to facilitate collaboration and accelerate advancements in the marine energy domain through the development of an open-source ecosystem of tools. This paper will describe the Tsdat framework and the data standards within which it operates. A beta version of Tsdat has been released and is being used by several projects in marine energy, wind energy, and building energy systems.

big data↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

Globus service enhancements for exascale applications and facilities

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications, particularly when leadership computing facilities are involved, can touch upon edge cases (e.g., terabyte files) that had not been a focus of previous Globus optimization work, which had emphasized rather the movement of many smaller (megabyte to gigabyte) files. We report here on how automated client-driven chunking can be used to accelerate both the movement of large files and the integrity checking operations that have proven to be essential for large data transfers. In conclusion, we present detailed performance studies that provide insights into the benefits of these modifications in a range of file transfer scenarios.

97 MATHEMATICS AND COMPUTING↗