Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data applications”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Few measurement shots challenge generalization in learning to classify entanglement

The ability to extract general laws from a few known examples depends on the complexity of the problem and on the amount of training data. In the quantum setting, the learner's generalization performance is further challenged by the destructive nature of quantum measurements that, together with the no-cloning theorem, limits the amount of information that can be extracted from each training sample. In this paper we focus on hybrid quantum learning techniques where classical machine-learning methods are paired with quantum algorithms and show that, in some settings, the uncertainty coming from a few measurement shots can be the dominant source of errors. We identify an instance of this possibly general issue by focusing on the classification of maximally entangled vs. separable states, showing that this toy problem becomes challenging for learners unaware of entanglement theory. Finally, we introduce an estimator based on classical shadows that performs better in the big data, few copy regime. Our results show that the naive application of classical machine-learning methods to the quantum setting is problematic, and that a better theoretical foundation of quantum learning is required.

97 MATHEMATICS AND COMPUTING↗

A Review of Recent and Emerging Machine Learning Applications for Climate Variability and Weather Phenomena

Abstract Climate variability and weather phenomena can cause extremes and pose significant risk to society and ecosystems, making continued advances in our physical understanding of such events of utmost importance for regional and global security. Advances in machine learning (ML) have been leveraged for applications in climate variability and weather, empowering scientists to approach questions using big data in new ways. Growing interest across the scientific community in these areas has motivated coordination between the physical and computer science disciplines to further advance the state of the science and tackle pressing challenges. During a recently held workshop that had participants across academia, private industry, and research laboratories, it became clear that a comprehensive review of recent and emerging ML applications for climate variability and weather phenomena that can cause extremes was needed. This article aims to fulfill this need by discussing recent advances, challenges, and research priorities in the following topics: sources of predictability for modes of climate variability, feature detection, extreme weather and climate prediction and precursors, observation–model integration, downscaling, and bias correction. This article provides a review for domain scientists seeking to incorporate ML into their research. It also provides a review for those with some ML experience seeking to broaden their knowledge of ML applications for climate variability and weather.

54 ENVIRONMENTAL SCIENCES↗

X-composer: enabling cross-environments in-situ workflows between HPC and cloud

As large-scale scientific simulations and big data analyses become more popular, it is increasingly more expensive to store huge amounts of raw simulation results to perform post-analysis. To minimize the expensive data I/O, "in-situ" analysis is a promising approach, where data analysis applications analyze the simulation generated data on the fly without storing it first. However, it is challenging to organize, transform, and transport data at scales between two semantically different ecosystems due to the distinct software and hardware difference. To tackle these challenges, we design and implement the X-Composer framework. X-Composer connects cross-ecosystem applications to form an "in-situ" scientific workflow, and provides a unified approach and recipe for supporting such hybrid in-situ workflows on distributed heterogeneous resources. X-Composer reorganizes simulation data as continuous data streams and feeds them seamlessly into the Cloud-based stream processing services to minimize I/O overheads. For evaluation, we use X-Composer to set up and execute a cross-ecosystem workflow, which consists of a parallel Computational Fluid Dynamics simulation running on HPC, and a distributed Dynamic Mode Decomposition analysis application running on Cloud. Our experimental results show that X-Composer can seamlessly couple HPC and Big Data jobs in their own native environments, achieve good scalability, and provide high-fidelity analytics for ongoing simulations in real-time.

Wang, Dali↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine

03 NATURAL GAS↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data (Final Report)

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components: (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine.

20 FOSSIL-FUELED POWER PLANTS↗

Spatiotemporal Pattern Recognition in the PMU Signals in the WECC system

Phasor measurement unit (PMU) data has been used by multiple power system applications, including state estimation, post event analysis, oscillation detection, model validation, and many others. Still, due to its big data nature and availability to general research institutions, comprehensive understanding of the spatiotemporal patterns and underlying mechanisms are incomplete. This study applies a set of signal processing and machine learning approaches aiming at deciphering the characteristic behaviors of multiple phasor measurement units (PMUs) attributes (e.g., voltage, frequency, rate of change of frequency, phase angle), including their auto-correlation, cross-dependence, similarities and discrepancies across units and temporal scales, and distributions of anomalies and their linkages to potential external factors such as weather events. Data analytics are applied to PMUs from the U.S. Western Electricity Coordinating Council (WECC) system. The PMU measurements, recorded events, outages, and weather extremes are all from real world datasets. The findings from the study and mechanistic understanding of the PMU dynamics help provide guidance on system control or preventing blackouts. The derived metrics can be directly used for adjusting or filtering simulated PMU data used for advanced algorithm development.

Hou, Zhangshuan↗

Parallel I/O Evaluation Techniques and Emerging HPC Workloads: A Perspective

Emerging workloads such as artificial intelligence, big data analytics and complex multi-step workflows alongside future exascale applications are anticipated future HPC workloads, which will result in a more diverse I/O system workload and even less predictable I/O behavior and access patterns. Along with the ever increasing gap between the compute and storage performance capabilities, the in-depth understanding of extreme-scale I/O behavior and the I/O performance modeling and prediction are essential tools of the large-scale I/O evaluation process for addressing the needs of extreme-scale hybrid workloads. In this survey article, we focus on the state-of-the-art of the I/O behavior and performance analysis process for HPC systems in a 5-year time window and identify future research challenges.

Neuwirth, Sarah↗

Big-data Efficient and Automated Science Transfer (BEAST): An Open-Source Software Architecture for Arc Jet Data Management, Modeling, and Automation

Big-data Efficient and Automated Science Transfer (BEAST) is a facility data management application developed for the NASA Ames arc jet facilities. The current decentralized data management practices limit statistical tracking, synchronization between video/time series, search capability, data throughput, and data processing speed/efficiency. Consequently, BEAST was developed to provide a new data infrastructure with streamlined data collection, processing, transfer, and analysis. This new framework also seeks to implement the FAIR principles of data stewardship: Findable, Accessible, Interoperable, and Reusable. The BEAST framework is based on a combination of the Python Django web framework and the Python data stack to provide a monolithic, open-source platform for data management, automation, and machine learning. This architecture was chosen for maintainability and scalability for a small, in-house development team. This paper will describe the application framework, deployment, and discuss the benefits and future plans for the system.

Data management↗

Strategies for Integrating Deep Learning Surrogate Models with HPC Simulation Applications

The emerging trend of the convergence of high performance computing (HPC), machine learning/deep learning (ML/DL), and big data analytics presents a host of challenges for large-scale computing campaigns that seek best practices to interleave traditional scientific simulation-based workloads with ML/DL models. A portfolio of systematic approaches to incorporate deep learning into modeling and simulation serves a vital need when we support AI for science at a computing facility. In this paper, we evaluate several strategies for deploying deep learning surrogate models in a representative physics application on supercomputers at the Oak Ridge Leadership Computing Facility (OLCF). We discuss a set of recommended deployment architectures and implementation approaches. We analyze and evaluate these alternatives and show their performance and scalability up to 1000 GPUs on two mainstream platforms equipped with different deep learning hardware and software stacks.

Yin, Junqi↗

Bundle Data Approach at GES DISC Targeting Natural Hazards

Severe natural phenomena such as hurricane, volcano, blizzard, flood and drought have the potential to cause immeasurable property damages, great socioeconomic impact, and tragic loss of human life. From searching to assessing the Big, i.e., massive and heterogeneous scientific data (particularly, satellite and model products) in order to investigate those natural hazards, it has, however, become a daunting task for Earth scientists and applications researchers, especially during recent decades. The NASA Goddard Earth Sciences Data and Information Service Center (GES DISC) has served Big Earth science data, and the pertinent valuable information and services to the aforementioned users of diverse communities for years. In order to help and guide our users to online readily (i.e., with a minimum effort) acquire their requested data from our enormous resource at GES DISC for studying their targeted hazard event, we have thus initiated a Bundle Data approach in 2014, first targeting the hurricane event topic. We have recently worked on new topics such as volcano and blizzard. The bundle data of a specific hazard event is basically a sophisticated integrated data package consisting of a series of proper datasets containing a group of relevant (knowledge--based) data variables readily accessible to users via a system-prearranged table linking those data variables to the proper datasets (URLs). This online approach has been developed by utilizing a few existing data services such as Mirador as search engine; Giovanni for visualization; and OPeNDAP for data access, etc. The online Data Cookbook site at GES DISC is the current host for the bundle data. We are now also planning on developing an Automated Virtual Collection Framework that shall eventually accommodate the bundle data, as well as further improve our management in Big Data.

GES DISC↗

MERRA Analytic Services: Meeting the Big Data Challenges of Climate Science Through Cloud-enabled Climate Analytics-as-a-service

Climate science is a Big Data domain that is experiencing unprecedented growth. In our efforts to address the Big Data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). We focus on analytics, because it is the knowledge gained from our interactions with Big Data that ultimately produce societal benefits. We focus on CAaaS because we believe it provides a useful way of thinking about the problem: a specialization of the concept of business process-as-a-service, which is an evolving extension of IaaS, PaaS, and SaaS enabled by Cloud Computing. Within this framework, Cloud Computing plays an important role; however, we it see it as only one element in a constellation of capabilities that are essential to delivering climate analytics as a service. These elements are essential because in the aggregate they lead to generativity, a capacity for self-assembly that we feel is the key to solving many of the Big Data challenges in this domain. MERRA Analytic Services (MERRAAS) is an example of cloud-enabled CAaaS built on this principle. MERRAAS enables MapReduce analytics over NASAs Modern-Era Retrospective Analysis for Research and Applications (MERRA) data collection. The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of 26 key climate variables. It represents a type of data product that is of growing importance to scientists doing climate change research and a wide range of decision support applications. MERRAAS brings together the following generative elements in a full, end-to-end demonstration of CAaaS capabilities: (1) high-performance, data proximal analytics, (2) scalable data management, (3) software appliance virtualization, (4) adaptive analytics, and (5) a domain-harmonized API. The effectiveness of MERRAAS has been demonstrated in several applications. In our experience, Cloud Computing lowers the barriers and risk to organizational change, fosters innovation and experimentation, facilitates technology transfer, and provides the agility required to meet our customers' increasing and changing needs. Cloud Computing is providing a new tier in the data services stack that helps connect earthbound, enterprise-level data and computational resources to new customers and new mobility-driven applications and modes of work. For climate science, Cloud Computing's capacity to engage communities in the construction of new capabilies is perhaps the most important link between Cloud Computing and Big Data.

Data Analytics↗

Machine Learning Lifecycle for Earth Science Application: A Practical Insight into Production Deployment

Earth science domain presents unique sets of problems that are increasingly being solved using data driven approaches. The availability of big Earth science data offers immense potential for Machine learning (ML) as evident from numerous research publications lately. However, many of these publications are not ending up as production applications mainly because the data scientists who develop the ML models are now expected to complete the ML lifecycle by deploying and scaling the models in production. We introduce ML lifecycle to the Earth science community including the opportunities and challenges that lie ahead in each phase of the lifecycle. We demonstrate the lifecycle using an Earth science problem that we used ML to address and transitioned to production.

Maskey, Manil↗

Topological Optimization with Big Steps

Using persistent homology to guide optimization has emerged as a novel application of topological data analysis. Existing methods treat persistence calculation as a black box and backpropagate gradients only onto the simplices involved in particular pairs. We show how the cycles and chains used in the persistence calculation can be used to prescribe gradients to larger subsets of the domain. In particular, we show that in a special case, which serves as a building block for general losses, the problem can be solved exactly in linear time. This relies on another contribution of this paper, which eliminates the need to examine a factorial number of permutations of simplices with the same value. Here, we present empirical experiments that show the practical benefits of our algorithm: the number of steps required for the optimization is reduced by an order of magnitude.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Big Data Meets Geothermal Exploration (CRADA Final Report)

As part of the Cyclotron Road program, Zanskar Geothermal & Minerals, Inc. investigated the application of micro-earthquake and ambient noise seismology methods to imaging and characterizing the structural characteristics and hydrothermal flux of subsurface faults. Significant advances in what could be resolved were enabled by two major developments in seismology: 1) the availability of large-n arrays of low-cost seismometers, and 2) the availability of increased computational power and semi-automated data reduction algorithms. In tandem, these advances may improve the signal-to-noise ratio and spatial precision of the data collected and enable higher-resolution characterization of subsurface fracture systems and their spatio-temporal evolution. These tools supported efforts to reduce dry-hole risk and to improve wellfield productivity for geothermal resource development. In particular, two applications of these advances were evaluated: 1) fracture-seismic imaging, which was used to detect ambient emissions from fluid-filled fractures, and 2) reservoir tomography, which used information about travel paths, source locations, and source parameters of micro-earthquakes to identify areas of enhanced permeability. Integration of these methods provided guidance for siting wells and served as prior constraints for reservoir models, informing forecasts of power potential and production and injection strategies aimed at minimizing temperature decline and improving overall resource productivity.

15 GEOTHERMAL ENERGY↗

Machine Learning Technologies and Their Applications for Science and Engineering Domains Workshop -- Summary Report

The fields of machine learning and big data analytics have made significant advances in recent years, which has created an environment where cross-fertilization of methods and collaborations can achieve previously unattainable outcomes. The Comprehensive Digital Transformation (CDT) Machine Learning and Big Data Analytics team planned a workshop at NASA Langley in August 2016 to unite leading experts the field of machine learning and NASA scientists and engineers. The primary goal for this workshop was to assess the state-of-the-art in this field, introduce these leading experts to the aerospace and science subject matter experts, and develop opportunities for collaboration. The workshop was held over a three day-period with lectures from 15 leading experts followed by significant interactive discussions. This report provides an overview of the 15 invited lectures and a summary of the key discussion topics that arose during both formal and informal discussion sections. Four key workshop themes were identified after the closure of the workshop and are also highlighted in the report. Furthermore, several workshop attendees provided their feedback on how they are already utilizing machine learning algorithms to advance their research, new methods they learned about during the workshop, and collaboration opportunities they identified during the workshop.

Ambur, Manjula↗

An overview of the operation architectures and energy management system for multiple microgrid clusters

The emerging novel energy infrastructures, such as energy communities, smart building-based microgrids, electric vehicles enabled mobile energy storage units raise the requirements for a more interconnective and interoperable energy system. It leads to a transition from simple and isolated microgrids to relatively large-scale and complex interconnected microgrid systems named multi-microgrid clusters. In order to efficiently, optimally, and flexibly control multi-microgrid clusters, cross-disciplinary technologies such as power electronics, control theory, optimization algorithms, information and communication technologies, cyber-physical, and big-data analysis are needed. This paper introduces an overview of the relevant aspects for multi-microgrids, including the outstanding features, architectures, typical applications, existing control mechanisms, as well as the challenges.

24 POWER TRANSMISSION AND DISTRIBUTION↗

How Bayesian methods can improve R -matrix analyses of data: The example of the d t reaction

The 3 H(d, n) 4 He reaction is of significant interest in nuclear astrophysics and nuclear applications. It is an important, early step in big-bang nucleosynthesis and a key process in nuclear fusion reactors. We use one- and two-level R-matrix approximations to analyze data on the cross section for this reaction at center-of-mass energies below 215 keV. We critically examine the data sets using a Bayesian statistical model that allows for both common-mode and additional point-to-point un- certainties. We use Markov Chain Monte Carlo sampling to evaluate this R-matrix-plus-statistical model and find two-level R-matrix results that are stable with respect to variations in the channel radii. The S factor at 40 keV evaluates to 25.36(19) MeV b (68% credibility interval). We discuss our Bayesian analysis in detail and provide guidance for future applications of Bayesian methods to R-matrix analyses. We also discuss possible paths to further reduction of the S-factor uncertainty.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗