Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “large data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Data Mining and Analysis

The Data Mining project seeks to bring the capability of data visualization to NASA anomaly and problem reporting systems for the purpose of improving data trending, evaluations, and analyses. Currently NASA systems are tailored to meet the specific needs of its organizations. This tailoring has led to a variety of nomenclatures and levels of annotation for procedures, parts, and anomalies making difficult the realization of the common causes for anomalies. Making significant observations and realizing the connection between these causes without a common way to view large data sets is difficult to impossible. In the first phase of the Data Mining project a portal was created to present a common visualization of normalized sensitive data to customers with the appropriate security access. The tool of the visualization itself was also developed and fine-tuned. In the second phase of the project we took on the difficult task of searching and analyzing the target data set for common causes between anomalies. In the final part of the second phase we have learned more about how much of the analysis work will be the job of the Data Mining team, how to perform that work, and how that work may be used by different customers in different ways. In this paper I detail how our perspective has changed after gaining more insight into how the customers wish to interact with the output and how that has changed the product.

iss↗

Downscaling Satellite Precipitation with Emphasis on Extremes: A Variational 1-Norm Regularization in the Derivative Domain

The increasing availability of precipitation observations from space, e.g., from the Tropical Rainfall Measuring Mission (TRMM) and the forthcoming Global Precipitation Measuring (GPM) Mission, has fueled renewed interest in developing frameworks for downscaling and multi-sensor data fusion that can handle large data sets in computationally efficient ways while optimally reproducing desired properties of the underlying rainfall fields. Of special interest is the reproduction of extreme precipitation intensities and gradients, as these are directly relevant to hazard prediction. In this paper, we present a new formalism for downscaling satellite precipitation observations, which explicitly allows for the preservation of some key geometrical and statistical properties of spatial precipitation. These include sharp intensity gradients (due to high-intensity regions embedded within lower-intensity areas), coherent spatial structures (due to regions of slowly varying rainfall),and thicker-than-Gaussian tails of precipitation gradients and intensities. Specifically, we pose the downscaling problem as a discrete inverse problem and solve it via a regularized variational approach (variational downscaling) where the regularization term is selected to impose the desired smoothness in the solution while allowing for some steep gradients(called 1-norm or total variation regularization). We demonstrate the duality between this geometrically inspired solution and its Bayesian statistical interpretation, which is equivalent to assuming a Laplace prior distribution for the precipitation intensities in the derivative (wavelet) space. When the observation operator is not known, we discuss the effect of its misspecification and explore a previously proposed dictionary-based sparse inverse downscaling methodology to indirectly learn the observation operator from a database of coincidental high- and low-resolution observations. The proposed method and ideas are illustrated in case studies featuring the downscaling of a hurricane precipitation field.

Hurricanes↗

Hadoop for High-Performance Climate Analytics: Use Cases and Lessons Learned

Scientific data services are a critical aspect of the NASA Center for Climate Simulations mission (NCCS). Hadoop, via MapReduce, provides an approach to high-performance analytics that is proving to be useful to data intensive problems in climate research. It offers an analysis paradigm that uses clusters of computers and combines distributed storage of large data sets with parallel computation. The NCCS is particularly interested in the potential of Hadoop to speed up basic operations common to a wide range of analyses. In order to evaluate this potential, we prototyped a series of canonical MapReduce operations over a test suite of observational and climate simulation datasets. The initial focus was on averaging operations over arbitrary spatial and temporal extents within Modern Era Retrospective- Analysis for Research and Applications (MERRA) data. After preliminary results suggested that this approach improves efficiencies within data intensive analytic workflows, we invested in building a cyber infrastructure resource for developing a new generation of climate data analysis capabilities using Hadoop. This resource is focused on reducing the time spent in the preparation of reanalysis data used in data-model inter-comparison, a long sought goal of the climate community. This paper summarizes the related use cases and lessons learned.

analytics↗

Application of Machine Learning to Rotorcraft Health Monitoring

Machine learning is a powerful tool for data exploration and model building with large data sets. This project aimed to use machine learning techniques to explore the inherent structure of data from rotorcraft gear tests, relationships between features and damage states, and to build a system for predicting gear health for future rotorcraft transmission applications. Classical machine learning techniques are difficult, if not irresponsible to apply to time series data because many make the assumption of independence between samples. To overcome this, Hidden Markov Models were used to create a binary classifier for identifying scuffing transitions and Recurrent Neural Networks were used to leverage long distance relationships in predicting discrete damage states. When combined in a workflow, where the binary classifier acted as a filter for the fatigue monitor, the system was able to demonstrate accuracy in damage state prediction and scuffing identification. The time dependent nature of the data restricted data exploration to collecting and analyzing data from the model selection process. The limited amount of available data was unable to give useful information, and the division of training and testing sets tended to heavily influence the scores of the models across combinations of features and hyper-parameters. This work built a framework for tracking scuffing and fatigue on streaming data and demonstrates that machine learning has much to offer rotorcraft health monitoring by using Bayesian learning and deep learning methods to capture the time dependent nature of the data. Suggested future work is to implement the framework developed in this project using a larger variety of data sets to test the generalization capabilities of the models and allow for data exploration.

machine learning↗

Ask-the-expert: Active Learning Based Knowledge Discovery Using the Expert

Often the manual review of large data sets, either for purposes of labeling unlabeled instances or for classifying meaningful results from uninteresting (but statistically significant) ones is extremely resource intensive, especially in terms of subject matter expert (SME) time. Use of active learning has been shown to diminish this review time significantly. However, since active learning is an iterative process of learning a classifier based on a small number of SME-provided labels at each iteration, the lack of an enabling tool can hinder the process of adoption of these technologies in real-life, in spite of their labor-saving potential. In this demo we present ASK-the-Expert, an interactive tool that allows SMEs to review instances from a data set and provide labels within a single framework. ASK-the-Expert is powered by an active learning algorithm for training a classifier in the backend. We demonstrate this system in the context of an aviation safety application, but the tool can be adopted to work as a simple review and labeling tool as well, without the use of active learning.

software↗

Ask-the-Expert: Active Learning Based Knowledge Discovery Using the Expert

Often the manual review of large data sets, either for purposes of labeling unlabeled instances or for classifying meaningful results from uninteresting (but statistically significant) ones is extremely resource intensive, especially in terms of subject matter expert (SME) time. Use of active learning has been shown to diminish this review time significantly. However, since active learning is an iterative process of learning a classifier based on a small number of SME-provided labels at each iteration, the lack of an enabling tool can hinder the process of adoption of these technologies in real-life, in spite of their labor-saving potential. In this demo we present ASK-the-Expert, an interactive tool that allows SMEs to review instances from a data set and provide labels within a single framework. ASK-the-Expert is powered by an active learning algorithm for training a classifier in the back end. We demonstrate this system in the context of an aviation safety application, but the tool can be adopted to work as a simple review and labeling tool as well, without the use of active learning.

GUI↗

Using ADOPT Algorithm and Operational Data to Discover Precursors to Aviation Adverse Events

The US National Airspace System (NAS) is making its transition to the NextGen system and assuring safety is one of the top priorities in NextGen. At present, safety is managed reactively (correct after occurrence of an unsafe event). While this strategy works for current operations, it may soon become ineffective for future airspace designs and high density operations. There is a need for proactive management of safety risks by identifying hidden and "unknown" risks and evaluating the impacts on future operations. To this end, NASA Ames has developed data mining algorithms that finds anomalies and precursors (high-risk states) to safety issues in the NAS. In this paper, we describe a recently developed algorithm called ADOPT that analyzes large volumes of data and automatically identifies precursors from real world data. Precursors help in detecting safety risks early so that the operator can mitigate the risk in time. In addition, precursors also help identify causal factors and help predict the safety incident. The ADOPT algorithm scales well to large data sets and to multidimensional time series, reduce analyst time significantly, quantify multiple safety risks giving a holistic view of safety among other benefits. This paper details the algorithm and includes several case studies to demonstrate its application to discover the "known" and "unknown" safety precursors in aviation operation.

aviation safet↗

Ask-The-Expert: Minimizing Human Review for Big Data Analytics Through Active Learning

In this CIF project, we worked toward semi-automating knowledge discovery from anomaly detection algorithms through the use of active learning. Active learning is an area of research within machine learning that uses an "expert in the loop" to learn from large data sets that have very few annotations or labels available, and where providing such labels is expensive. In our case, the task can be defined as the identification of safety events from flight operational data. Since traditional anomaly detection algorithms cannot differentiate between operationally relevant and irrelevant statistical anomalies, Subject Matter Experts (SMEs) have a lengthy and expensive burden of investigating every example identified by the detection algorithm, classifying and labeling them as relevant or irrelevant. Active learningidentifies the unlabeled example for which a label would most improve the classifier, asks the domain expert for a label, and repeats this process until there are no more resources (time, budget) available for labeling or a minimum required performance is reached. A positive label indicates an operationally significant safety event whereas a negative label indicates otherwise. Based on these few labels we propose to build an active learning system that utilizes the SME's time in the most effective manner by iteratively asking for labels for as few informative instances as possible. Our work was proposed to be a stepping stone toward implementation and deployment of the system with user interface to be pursued by the Aviation Operations and Safety Program (AOSP) given its interest in safety monitoring and discovery of safety incidents.

aviation safety↗

Detecting And Characterizing Archetypes of Unintended Consequences in Engineered Systems

When designing engineered systems, the potential for unintended consequences of design policies or design decisions exists despite best intentions. Conditions that might cause the formation of unintended consequences are often known only in hindsight. However, since these conditions are associated with a single event, it is difficult to uncover the general patterns of conditions leading to unintended consequences. In this research, patterns of conditions associated with unintended consequences are learned from historical data and represented in the form of archetypes. While previous work using systems theoretic modeling has identified high-level archetypes, this work leverages a self-organizing map to learn archetypes of unintended consequences from human-tagged risk factors in a large data set of lessons learned from adverse events at NASA. The sixty-six identified archetypes contain patterns of conditions such as complexity and human-machine interaction associated with the formation of unintended consequences. To validate the archetypes, a sample of the archetypes is represented using system dynamics in order to illustrate that the identified archetypes are specialized versions of known high-level archetypes of unintended consequences. While the research is based upon a specific dataset, the archetypes apply to any engineered system and the pattern of leading indicators open a new path to manage unintended consequences and mitigate the magnitude of potentially adverse outcomes.

Hannah S Walsh↗

The Ejectable Data Recorder: A Lean, Risk-Informed Approach for Hardware Development

NASA is developing the Orion spacecraft to transport crew from the Earth to the Moon as part of the Artemis series of missions. To provide a crew escape capability from pre-launch through ascent, the Orion vehicle is equipped with a Launch Abort System (LAS), built by Lockheed Martin, which pulls the capsule away from the launch vehicle in the event of an abort scenario. The Ascent Abort 2 (AA-2) test flight occurred on July 2, 2019,and tested a production version of the LAS to ensure that it can operate as intended, and to collect a large data set from hundreds of sensors on the vehicle to support Orion flight certification. In the original AA-2 architecture, a single-string set of communications antennas on the LAS would downlink all of the in-flight test data to ground stations. However, that communications architecture was predicted to have data dropouts during abort and jettison of the LAS, and would not support data transmission at all after LAS jettison. As a result, a comprehensive trade study was completed, yielding the addition of antennas on the crew module (CM), a buffer/rebroadcast capability for key portions of the flight, and an ejectable data recorder (EDR) subsystem. This EDR subsystem would serve as a backup to the radio frequency (RF) communications system, and would be non-flight critical, providing a unique capability that enabled management to take a different approach with the hardware and software development. The Crew Module and Separation Ring were developed as “Class 1”Flight Hardware, albeit with some tailoring approaches to enable efficiencies. The Class 1 designation requires full rigor for flight hardware and software, documenting everything that happens to a piece of hardware from procurement through disposal, requiring a full spectrum of acceptance tests, and the highest rigor of quality assurance processes. At the other end of the spectrum, Class 3hardware is controlled, but not intended for flight, and leaves the level of rigor up to the project manager. This classification is often used for research and development projects. Similarly,Class-1E has been recently defined at NASA for ISS payloads and technology development projects that are not flight critical and do not need the full rigor of Class 1 to be successful. The EDR subsystem was challenged at commencement to adopt a skunkworks and agile-like approach to hardware development, allowing for a different risk posture than the rest of the AA-2 hardware. After initially pursuing Class 1 processes, the EDR subsystem design evolved to incorporating numerous commercial components, leading to re-designation as a Class-1E subsystem. The resulting EDR subsystem was fully successful in meeting all flight system requirements, and achieved 100% retrieval of flight test data. This paper will discuss the risk posture of the EDR subsystem and the subsequent tailoring that was enacted as part of its Class-1E status.

EDR↗

Machine Learning for the Zwicky Transient Facility

The Zwicky Transient Facility is a large optical survey in multiple filters producing hundreds of thousands of transient alerts per night. We describe here various machine learning (ML) implementations and plans to make the maximal use of the large data set by taking advantage of the temporal nature of the data, and further combining it with other data sets. We start with the initial steps of separating bogus candidates from real ones, separating stars and galaxies, and go on to the classification of real objects into various classes. Besides the usual methods (e.g., based on features extracted from light curves) we also describe early plans for alternate methods including the use of domain adaptation, and deep learning. In a similar fashion we describe efforts to detect fast moving asteroids. We also describe the use of the Zooniverse platform for helping with classifications through the creation of training samples, and active learning. Finally we mention the synergistic aspects of ZTF and LSST from the ML perspective.

Ashish Mahabal↗

A Robust Schema for Storing and Managing Machine Learning Data and Models

- Machine Learning (ML) has enabled models that can improve efficiency and decrease computational cost - ML models are crucial in enabling Integrated Computational Materials Engineering (ICME) - Large data sets require robust means of storing ML data and models

Brandon L. Hearley↗

Ocean and Earth System Modelling

Petascale supercomputing infrastructure + modelling and analysis capabilities + interdisciplinary upper-ocean expertise Multiscale ocean turbulence simulation Physical-biogeochemical interactions Analysis of large data sets from remote sensing and Earth system model ensembles

Ocean↗

Easy, Scalable Subsetting of GEDI Point Clouds

The GEDI Subsetter, a Python tool developed for NASA’s Multi-mission Algorithm and Analysis Platform (MAAP), optimizes the accessibility and visualization of GEDI point clouds by enabling users to efficiently subset data in a convenient, scalable manner. Complex science data often requires users to learn new software skills and handle many large files. Handling and cleaning large data sets is tedious and error-prone. These challenges significantly impede analysis. One of the goals of NASA's MAAP is to provide a platform that lowers the barrier to conducting research and analysis at scale. When a group of MAAP users wanted to conduct above-ground biomass estimation using GEDI data, we found that their existing workflow for leveraging GEDI data suffered from the barriers mentioned above. Furthermore, their workflow did not scale easily beyond a small number of granules. We found that existing tools related to GEDI data retrieval and subsetting were too limiting, so the GEDI Subsetter was written to support MAAP users’ needs. Being able to run many subsetting jobs simultaneously in the MAAP, and parallelizing the code itself, has led to significant speed improvements in obtaining relevant data, reducing subsetting time from hours to minutes. MAAP users can now more quickly and easily obtain only the data relevant to their research, by choosing which GEDI collection they want to work with (L1A, L2A, L2B, or L4A), and how they want to subset it, by specifying an area of interest, a temporal range, and relevant attributes. This has significantly reduced the feedback loop for users, allowing them to much more quickly subset GEDI data and begin their analysis. Although the GEDI Subsetter originally targeted users of the MAAP, it is generalized such that it can also be used outside of the MAAP and includes a command-line interface for convenience. Furthermore, with minor modifications, it should be possible to use it with non-GEDI data as the general pattern should be applicable to other sparse/track-based sensors.

Charles Daniels↗

Improved Access and Analysis of Data Provides New Opportunities for Smart Cities

Making smart, informed, data-driven energy decisions for cities requires large amounts of high-value data and analysis. Cities’ energy data has traditionally been difficult to access and costly to store and manage for city administrators. When data is available, many cities lack the expertise to properly analyze and model large data sets. NREL’s transparency and knowledge about data availability, integration, and interpretation is critical to decision-making for cities.

cities' data↗

PV Reliability Lessons from 100,000 Systems

Despite the importance of reliability to the cost competitiveness of PV, large data sets enabling high-level investigation of the technology’s performance in the field are relatively scarce. Dirk C. Jordan, Chris Deline, Bill Marion and Teresa Barnes of the National Renewable Energy Laboratory, and Mark Bolinger of the Lawrence Berkeley National Laboratory study a unique data set of 100,000 PV systems in the US, drawing out tips for better reliability that have relevance to other parts of the world.

41 EE - Solar Energy Technologies Office (EE-4S)↗

On the Use of Smart Meter Data to Estimate the Voltage Magnitude on the Primary Side of Distribution Service Transformers: Preprint

This paper develops a novel method to estimate the voltage magnitude on the primary side of distribution service transformers. The proposed method relies exclusively on smart meters, and therefore it is fully data-driven. This is an important feature because electric utilities have detailed models of only the primary network--that is, the network between the distribution substation and the primary side of service transformers that are installed closer to end-customer sites. The network that connects the secondary side of service transformers to end-customer sites, referred to as the secondary network, is simply represented by a lumped load. For each secondary network, the proposed method uses data acquired from only 2 smart meters: the closest and the farthest--in the sense of electrical distance--from the service transformer. As a reference to this feature, the proposed method is named SM2Vp. To our knowledge, this is the first time a method is shown to provide actionable information for real-time operation and control of power distribution grids using only two smart meters per secondary network. This is important because utilities have experienced barriers in managing and using large data sets for real-time operation and control. SM2Vp is primarily intended to provide pseudo-measurements for distribution system state estimation, but it can also be used directly for voltage control schemes. The performance of SM2Vp is demonstrated by numerical simulations carried out on three secondary network synthetic models and by using field data provided by a utility partner serving customers in southwestern California. A maximum relative error of approximately 3.9% or less is observed for the primary voltage magnitude estimates in all numerical experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Advancing the Theory of Nuclear Data Evaluations [Abstract]

We present recent advances in the R-matrix formalism as well as the Bayesian evaluation framework for improved nuclear data evaluations. The advances in the R matrix formalism include: 1) direct processes, 2) doorway, as well as multistep, processes, and 3) various forms of the Reich-Moore approximation for eliminated capture channels. Furthermore, to address unreasonably small posterior uncertainties often encountered in nuclear data evaluations of large data sets using the conventional form of the Bayes’ theorem, we introduce imperfections (of the data or the model) as a formal evaluation tool for taming the evaluated uncertainties in harmony with Bayes’ theorem. These theoretical advances were motivated by the nuclear data evaluations of differential resolved resonance cross section data using the code SAMMY, as well as the integral benchmark experiments using the SCALE code system, being performed at Oak Ridge National Laboratory for the Nuclear Criticality Safety Program. Some pedagogical applications of the new formalism, as well as a snapshot of the SAMMY modernization efforts, will be presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗