Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “big data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

GOOML: Geothermal Operational Optimization with Machine Learning

Geothermal Operational Optimization with Machine Learning (GOOML) is a project focused on maximizing increased availability and capacity from existing industrial-scale geothermal generation assets. The GOOML project will develop a suite of machine learning-based algorithms that analyze historical production datasets and provide predictive setpoints for geothermal field operations. Historical datasets from New Zealand and the US will provide the input to develop digital geothermal system twins which allow prediction of market conditions, maintenance operations and steamfield optimization. The algorithms will identify key parameters within fields and suggest setpoints for components of the system to maintain optimal generation. Set-points can be instructed to follow mass flow restrictions, generation maximization and optimal field/reservoir balance and give field operators a guide by which generation can be optimized. The datasets that will be used to develop GOOML are sourced from operating geothermal fields in New Zealand and the United States with varying degrees of complexity. This will ensure that most geothermal systems can utilize the GOOML tool to assist in optimizing operations. GOOML aims to achieve a step-change in geothermal operations by developing state-of-the-art machine learning algorithms, comprehensive data analytics, and a first-of-its-kind automated, intelligent geothermal system model.

algorithms↗

Time and Measurement Days

Questions in data analysis involving the concepts of time and measurement are often pushed into the background or reserved for a philosophical discussion. Some examples are: a) Is causality a consequence of the laws of physics, or can the arrow of time be reversed? b) Can we determine the arrow of time of an event? c) Do we need the continuum hypothesis for the underlying function in any measurement process? d) Can we say anything about the analyticity of the underlying process of an event? e) Would it be valid to model a non-analytical process as function of time? f) What are the implications of all these questions for classical Fourier techniques? However, in the age of big data gathered either from space missions supplying ultra-precise long time series, or e.g. LIGO data from the ground, the moment to bring these questions to the foreground seems arrived. The limitations of our understanding of some fundamental processes is emphasized by the lack of solution for problems open for more than 2 decades, such as the non-detection of solar g-modes, or the modal identification of main sequence stellar pulsators like delta Scuti stars. Flicker noise or 1/f noise, for example, attributed in solar-like stars to granulation, is analyzed mostly only to apply noise reduction techniques, neither considering the classical problem of 1/f noise that was introduced a 100 years ago, nor taking into account ergodic or non-ergodic solutions that make inapplicable spectral analysis techniques in practice. This topic was discussed by Nicholas W. Watkins during the ITISE meeting held in Granada in 2016. There he presented preliminary results of his research on Mandelbrot's related work. We reproduce here his quotation of Mandelbrot (1999) "There is a sharp contrast between a highly anomalous ("non-white") noise that proceeds in ordinary clock time and a noise whose principal anomaly is that it is restricted to fractal time", suggesting a connection with the above proposed topics that could be phrased as the following additional questions:a) Is self-organized criticality (SOC) frequent in astrophysical phenomena? b) Could all fractals in nature be considered stochastic? c) Could we establish mathematical/physical relationships between chaotic and fractal behaviors in time series? d) Could the differences between fractals and chaos in terms of analyticity be used to understand the residuals of the fitting of stellar light curves? In this meeting we would like to approximate these problems from a holistic and multidisciplinary perspective, taking into account not only technical issues but also the deeper implications. In particular the concept of connectivity (introduced in Pascual-Granado et al. A&A, 2015) could be used to implement, within the framework of ARMA processes, an "arrow of time" (see attached document), and so studying the possible implications in the concept of time as envisaged by Watkins.

data analysis↗

Sherlock Data Warehouse

This slide deck provides an overview of the data and resources available in the Sherlock Data Warehouse. Sherlock was developed and is currently maintained by the Aviation Systems Division at NASA Ames Research Center. Sherlock contains a valuable collection of flight, air traffic management, and weather data. But Sherlock is not just a data archive. Sherlock also includes tools and resources to access, download, and visualize data, as well as resources to process the data. This overview summarizes Sherlock data sources, demonstrates data analytics and visualization with MicroStrategy, illustrates disparate data integration using the ATM Knowledge graph, and presents a machine learning use case using the Big Data system.

data warehouse↗

FaceBase 3: analytical tools and FAIR resources for craniofacial and dental research

ABSTRACT The FaceBase Consortium was established by the National Institute of Dental and Craniofacial Research in 2009 as a ‘big data’ resource for the craniofacial research community. Over the past decade, researchers have deposited hundreds of annotated and curated datasets on both normal and disordered craniofacial development in FaceBase, all freely available to the research community on the FaceBase Hub website. The Hub has developed numerous visualization and analysis tools designed to promote integration of multidisciplinary data while remaining dedicated to the FAIR principles of data management (findability, accessibility, interoperability and reusability) and providing a faceted search infrastructure for locating desired data efficiently. Summaries of the datasets generated by the FaceBase projects from 2014 to 2019 are provided here. FaceBase 3 now welcomes contributions of data on craniofacial and dental development in humans, model organisms and cell lines. Collectively, the FaceBase Consortium, along with other NIH-supported data resources, provide a continuously growing, dynamic and current resource for the scientific community while improving data reproducibility and fulfilling data sharing requirements.

60 APPLIED LIFE SCIENCES↗

Open Energy Data Initiative (OEDI) FY22-24 (Final Technical Report)

Final technical report for the Open Energy Data Initiative (OEDI) project covering fiscal years FY22 through FY24. The DOE Open Energy Data Initiative (OEDI) is a partnership between the National Renewable Energy Laboratory (NREL), the U.S. Department of Energy (DOE), and major cloud providers including Amazon, Microsoft, and Google to provide universal access to big data in the cloud. At the heart of OEDI is a centralized repository of high-value energy research datasets aggregated from the U.S. Department of Energy's Program Offices, National Laboratories and other collaborators. It aggregates smaller, domain-specific repositories, allows direct data submissions, and includes support for big data through its energy data lakes. OEDI's data lakes make high-value data universally accessible and help researchers, collaborators and the general public overcome many of the obstacles to accessing and using big data.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Integrated Analysis of Multiple User Metrics - A “Sequel”; and Introducing the Google Analytic

For decades, the Goddard Earth Sciences Data and Information Services Center (GES DISC) has archived and distributed enormous volumes of NASA Earth science data (accompanied with many developed tools and services) to various research/applications communities and the general public. Being “immersed” in the Big Data era, we have inevitably faced the challenges of our continually increasing archived data in both volume and variety, as well as enhanced user needs and demands. In recent years, we have actively analyzed different types of user metrics, such as operational distribution metrics (recording numbers of distinct users and downloaded data files, size of distributed data volume): user publication metrics (mining info from our Giovanni users’ publications): and Bugzilla metrics (collecting info from user questions or feedback from user assistance tickets). Such metrics have helped us achieve a better understanding of user needs, demands, characteristics, and behaviors, which has then helped us improve our user services. Now we will present a “Sequel” of integrated analysis of multiple metrics at the GES DISC by introducing and adding one new kind of metrics acquired via utilizing our recently implemented Google Analytic 360 suite. Several “newer” reports, e.g., “What web site features and links are the most popular (and least)?” and “What are the top 25 dataset Keyword searches?” retrieved from this new metrics set will be presented, along with the aforementioned “traditional” metrics results.

Shie, Chung-Lin↗

Thermally driven phase transition of halide perovskites revealed by big data-powered in situ electron microscopy

Halide perovskites are promising light-absorbing materials for high-efficiency solar cells, while the crystalline phase of halide perovskites may influence the device’s efficiency and stability. In this work, we investigated the thermally driven phase transition of perovskite (CsPbIxBr3—x), which was confirmed by electron diffraction and high-resolution transmission electron microscopy results. CsPbIxBr3—x transitioned from δ phase to α phase when heated, and the γ phase was obtained when the sample was cooled down. The γ phase was stable as long as it was isolated from humidity and air. A template matching-based data analysis method enabled visualization of the thermally driven phase evolution of perovskite during heating. Here, we also proposed a possible atomic movement in the process of phase transition based on our in situ heating experimental data. The results presented here may improve our understanding of the thermally driven phase transition of perovskite as well as provide a protocol for big-data analysis of in situ experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Comparing Model Representations of Physiological Limits on Transpiration at a Semi-arid Ponderosa Pine Site

Mechanistic representations of biogeochemical processes in ecosystem models are rapidly advancing, requiring advancements in model evaluation approaches. Here we quantify multiple aspects of model functional performance to evaluate improved process representations in ecosystem models. We compare semi-empirical stomatal models with hydraulic constraints against more mechanistic representations of stomatal and hydraulic functioning at a semi-arid pine site using a suite of metrics and analytical tools. We find that models generally perform similarly under unstressed conditions, but performance diverges under atmospheric and soil drought. The more empirical models better capture synergistic information flows between soil water potential and vapor pressure deficit to transpiration, while the more mechanistic models are overly deterministic. Although models can be parameterized to yield similar functional performance, alternate parameterizations could not overcome structural model constraints that underestimate the unique information contained in soil water potential about transpiration. Additionally, both multilayer canopy and big-leaf models were unable to capture the magnitude of canopy temperature divergence from air temperature, and we demonstrate that errors in leaf temperature can propagate to considerable error in simulated transpiration. This study demonstrates the value of merging underutilized observational data streams with emerging analytical tools to characterize ecosystem function and discriminate among model process representations.

54 ENVIRONMENTAL SCIENCES↗

A Machine-Learning Approach to Assess Aircraft Engine System Performance

Artificial intelligence (AI)/machine learning, and big data are transforming the global business environment. They have become the most disruptive technologies for organizations to improve workplace efficiency and productivity. This work explored the application of machine learning-based predictive analytics that would enable aircraft engine designers to estimate engine system performance quickly during the conceptual design stage. Supervised machine-learning algorithm was employed to study patterns in an existing database of production and research turbofan engines, and built predictive analytics for use in predicting system performance of new turbofan designs. Specifically, the author developed deep-learning analytics to predict turbofan system weight, using turbofan design parameters as the input. The predictive analytics were trained and deployed in Keras, an open-source neural networks API (application program interface) written in Python, with TensorFlow (an open-source artificial AI library developed by Google) serving as the backend engine. The current engine-weight prediction results, together with those for the TSFC (thrust specific fuel consumption) and core-size predictions that were studied previously by the author, show that machine learning-based predictive analytics can be an effective, time-saving tool for aircraft engine design-space exploration during the conceptual design stage. It would enable expeditious identification of the best engine design amongst several candidates.

Michael T Tong↗

Time series methods for the analysis of soundscapes and other cyclical ecological data

Biodiversity monitoring has entered an era of ‘big data’, exemplified by a near-continuous collection of sounds, images, chemical and other signals from organisms in diverse ecosystems. Such data streams have the potential to help identify new threats, assess the effectiveness of conservation interventions, as well as generate new ecological insights. However, appropriate analytical methods are often still missing, particularly with respect to characterizing cyclical temporal patterns. Here, we present a framework for characterizing and analysing ecological responses that represent nonstationary, complex temporal patterns and demonstrate the value of using Fourier transforms to decorrelate continuous data points. In our example, we use a framework based on three approaches (spectral analysis, magnitude squared coherence, and principal component analysis) to characterize differences in tropical forest soundscapes within and across sites and seasons in Gabon. By reconstructing the underlying, cyclic behaviour of the soundscape for each site, we show how one can identify circadian patterns in acoustic activity. Soundscapes in the dry season had a complex diel cycle, requiring multiple harmonics to represent daily variation, while in the wet season there was less variance attributable to the daily cyclic patterns. Our framework can be applied to most continuous, or near-continuous ecological data collected at a fine temporal resolution, allowing ecologists to explore patterns of temporal autocorrelation at multiple levels for biologically meaningful trends. Such methods will become indispensable as biological big data are used to understand the impact of anthropogenic pressures on biodiversity and to inform efforts to mitigate them.

54 ENVIRONMENTAL SCIENCES↗

GOOML - Finding Optimization Opportunities for Geothermal Operations: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach. We have used this framework to develop digital twins that provide steamfield operators with an operational environment to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management for real world applications. The GOOML modeling software is built on a generic component-based systems framework that allows for both historical and forecast analysis. A GOOML model can perform historical data-assimilation using first-principal thermodynamics to create a meaningful data model. Historical production data can then be coupled with a forecast framework to train machine-learning models of steamfield components to predict future outputs. This modeling environment enables digital exploration of steamfield design configurations and operational scenarios. GOOML digital twins have been developed for steamfields in New Zealand and the United States representing differing power generation and field conditions. These digital twins have been validated by comparing hindcast predictions against historical production data. Reinforcement learning experiments were conducted to demonstrate the ability to programmatically explore the operations space using machine learning agents. Our initial results are compelling; two to five percent increases in annual energy production were demonstrated by the GOOML models with no additional infrastructure build required. GOOML offers a new approach to geothermal operations by applying state-of-the-art machine learning algorithms, comprehensive data analytics, and interaction with digital twins. Through application of these tools, operators will realize greater availability and higher net generation which will increase the cost effectiveness of geothermal energy projects.

access↗

Cloud Giovanni: Reining in Costs and Improving Performance with Analytical Data Stores Using Scalable Serverless Architecture

Giovanni is the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure developed at NASA GES DISC which provides a simple and intuitive way to visualize, analyze, and access vast amounts of Earth science data. It receives large number of user requests each day for a variety of analysis and visualization services, which leads to the big data challenge of serving gradually increasing large data volumes with diverse statistical algorithms. We hereby propose a multi-dimensional accumulation method which provides fast and cost-efficient cloud analysis for diverse services including both area averaging and time averaging. This method involves the weighted volume integration over multiple variable dimensions (time and space), and is implemented in AWS using Athena providing serverless and highly scalable data analysis. Compared to the standard method, this approach dramatically reduced the computational time by order of magnitude with a minimal AWS cost incurred. For example, for a benchmark of 10-year area averaging over the 1x1 degree daily variable, the computational time was reduced from minutes to seconds, and the Athena cost is only $5 for 100,000 requests.

Zhang, Hailiang↗

FY2020 Energy Efficient Mobility Systems Annual Progress Report

EEMS Program activities during FY 2020 focused on analytical research and large-scale modeling and simulation to understand the impacts that new mobility technologies and services will have at the vehicle-, traveler-, and overall transportation system-level. This research included the development of a multi-fidelity, end-to-end transportation system models and tools to evaluate the complex interactions among the various actors within the mobility landscape, analysis of empirical data to characterize which solutions may provide the largest benefits, and development of new control systems and algorithms that use vehicle connectivity and automation to improve the performance and efficiency of individual vehicles as well as the overall traffic system. This document presents a brief overview of the EEMS Program and documents progress and results from projects within each of the EEMS activity areas. The Computational Modeling and Simulation key activity area summarizes work within the sub-areas of (1) the SMART (Systems and Modeling for Accelerated Research in Transportation) Mobility Lab Consortium, (2) Artificial Intelligence, High-Performance Computing, and Data Analytics, and (3) Core Simulation and Evaluation Tools. Additionally, the program’s advanced R&D projects are summarized within (4) the Connectivity and Automation Technology key activity area. Each of the individual progress reports provide a project overview and highlights of the technical results.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Stellar nucleosynthesis and chemical evolution of the solar neighborhood

Current theoretical models of nucleosynthesis (N) in stars are reviewed, with an emphasis on their implications for Galactic chemical evolution. Topics addressed include the Galactic population II red giants and early N; N in the big bang; star formation, stellar evolution, and the ejection of thermonuclearly evolved debris; the chemical evolution of an idealized disk galaxy; analytical solutions for a closed-box model with continuous infall; and nuclear burning processes and yields. Consideration is given to shell N in massive stars, N related to degenerate cores, and the types of observational data used to constrain N models. Extensive diagrams, graphs, and tables of numerical data are provided.

Clayton, Donald D.↗