SEARCH · Engineering Papers
Results for “Data exploration”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Tracking and data acquisition for space exploration
Tracking and data acquisition system between spacecraft and ground station
Tracking and data acquisition for space exploration.
Tracking and data acquisition systems for spacecraft and ground link through command, telemetry and tracking, noting data relay satellites, etc
The Value of Data and Metadata Standardization for Interoperability in Giovanni Or: Why Your Product's Metadata Causes Us Headaches!
Giovanni is a data exploration and visualization tool at the NASA Goddard Earth Sciences Data Information Services Center (GES DISC). It has been around in one form or another for more than 15 years. Giovanni calculates simple statistics and produces 22 different visualizations for more than 1600 geophysical parameters from more than 90 satellite and model products. Giovanni relies on external data format standards to ensure interoperability, including the NetCDF CF Metadata Conventions. Unfortunately, these standards were insufficient to make Giovanni's internal data representation truly simple to use. Finding and working with dimensions can be convoluted with the CF Conventions. Furthermore, the CF Conventions are silent on machine-friendly descriptive metadata such as the parameter's source product and product version. In order to simplify analyzing disparate earth science data parameters in a unified way, we developed Giovanni's internal standard. First, the format standardizes parameter dimensions and variables so they can be easily found. Second, the format adds all the machine-friendly metadata Giovanni needs to present our parameters to users in a consistent and clear manner. At a glance, users can grasp all the pertinent information about parameters both during parameter selection and after visualization.
Automated Vehicle Detection in a Nuclear Facility Using Low-Frequency Acoustic Sensors
This article presents an analysis of the method of construction and results for a classifier intended to identify vehicles using low-frequency acoustic data collected by a dis-tributed sensor network. This data is collected as part of a venture intended to explore data analytics and multisensor fusion techniques for the monitoring of activities at a test bed nuclear facility located at Oak Ridge National Laboratory in Oak Ridge, Tennessee. We describe the associated target signature and design a classifier based on a multilayer perceptron, followed by an analysis of its results. We discuss how overall accuracy is not the only consideration in constructing this classifier, and how for this application, it is actually desirable to operate at a lower level of accuracy in exchange for a reduction in the false alarm rate, as well as how this relates to the actual deployment of the classifier in practical use.
Navigating Uncertainty: Challenges in Visualizing Ensemble Data and Surrogate Models for Decision Systems
Uncertainty visualization plays a critical role in transforming ensemble simulation data into actionable insights by effectively communicating various dimensions of uncertainty within a system. The emergence of artificial intelligence-driven surrogate models trained on multirun ensemble data offers a transformative opportunity to replace computationally intensive simulations with fast estimates, enabling users to explore data spaces with unprecedented depth and interactivity. However, integrating ensemble data and surrogate models into decision-making workflows and tools introduces novel challenges for uncertainty visualization. These include reconciling and clearly communicating the unique uncertainties associated with ensembles and their surrogate model estimates, and leveraging these approximations to inform actionable decisions. This work explores these challenges in the context of high-dimensional data visualization, bridging discrete datasets with their continuous representations and addressing the complexities of systems that support iterative navigation between input and output spaces. We evaluate the role of uncertainty visualization in fostering intuitive, actionable interactions and identify critical hurdles in advancing this frontier of computational simulation.
Scenario adaptive disruption prediction study for next generation burning-plasma tokamaks
Next generation High Performance (HP) tokamaks risk damage from unmitigated disruptions at high current and power. Achieving reliable disruption prediction for a device's HP operation based on its Low Performance (LP) data is a key to its success. In this letter, through explorative data analysis and dedicated numerical experiments on multiple existing tokamaks, we demonstrate how the operational regimes of tokamaks can affect the power of a trained disruption predictor. First, our results suggest data-driven disruption predictors trained on abundant LP discharges work poorly on the HP regime of the same tokamak, which is a consequence of the distinct distributions of the tightly correlated signals related to disruptions in these two regimes. Second, we find that matching operational parameters among tokamaks strongly improves cross-machine accuracy which implies our model learns from the underlying scalings of dimensionless physics parameters like q 95 , β p and confirms the importance of these parameters in disruption physics and cross machine domain matching from the data-driven perspective. Finally, our results show in the absence of HP data from the target devices, the best predictivity of the HP regime for the target machine can be achieved by combining LP data from the target with HP data from other machines. Furthermore, these results provide a possible disruption predictor development strategy for next generation tokamaks, such as ITER and SPARC, and highlight the importance of developing baseline scenario discharges of future tokamaks on existing machines to collect more relevant disruptive data.
Approach to Managing MeaSURES Data at the GSFC Earth Science Data and Information Services Center (GES DISC)
A major need stated by the NASA Earth science research strategy is to develop long-term, consistent, and calibrated data and products that are valid across multiple missions and satellite sensors. (NASA Solicitation for Making Earth System data records for Use in Research Environments (MEaSUREs) 2006-2010) Selected projects create long term records of a given parameter, called Earth Science Data Records (ESDRs), based on mature algorithms that bring together continuous multi-sensor data. ESDRs, associated algorithms, vetted by the appropriate community, are archived at a NASA affiliated data center for archive, stewardship, and distribution. See http://measures-projects.gsfc.nasa.gov/ for more details. This presentation describes the NASA GSFC Earth Science Data and Information Services Center (GES DISC) approach to managing the MEaSUREs ESDR datasets assigned to GES DISC. (Energy/water cycle related and atmospheric composition ESDRs) GES DISC will utilize its experience to integrate existing and proven reusable data management components to accommodate the new ESDRs. Components include a data archive system (S4PA), a data discovery and access system (Mirador), and various web services for data access. In addition, if determined to be useful to the user community, the Giovanni data exploration tool will be made available to ESDRs. The GES DISC data integration methodology to be used for the MEaSUREs datasets is presented. The goals of this presentation are to share an approach to ESDR integration, and initiate discussions amongst the data centers, data managers and data providers for the purpose of gaining efficiencies in data management for MEaSUREs projects.
Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud
Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.
Revisit NGC 5466 tidal stream with Gaia , SDSS/SEGUE, and LAMOST
ABSTRACT By mining the data from Gaia Early Data Release 3, Sloan Digital Sky Survey/Sloan Extension for Galactic Understanding and Exploration Data Release 16, and Large Sky Area Multi-Object Fiber Spectroscopic Telescope Data Release 8, 11 member stars of the NGC 5466 tidal stream are detected and 7 of them are newly identified. To reject contaminators, a variety of cuts are applied in sky position, colour–magnitude diagram, metallicity, proper motion, and radial velocity. We compare our data to a mock stream generated by modelling the cluster’s disruption under a smooth Galactic potential plus the Large Magellanic Cloud (LMC). The concordant trends in phase space between the model and observations imply that the stream might have been perturbed by the LMC. The two most distant stars among the 11 detected members trace the stream’s length to 60° of sky, supporting and extending the previous length of 45°. Given that NGC 5466 is so distant and potentially has a longer tail than previously thought, we expect that the NGC 5466 tidal stream could be a useful tool in constraining the Milky Way gravitational field.
Blending Machine Learning and Interaction Design in Audio Explorer
The results of machine learning models can often be difficult to interpret, especially for domain experts. Audio Explorer, the winning entry of the 2018 VAST Challenge, is an interactive data exploration tool that effectively communicates machine learning results, using coordinated geospatial, temporal, and auditory visualizations to promote information discovery.
High-Fidelity Solar Irradiance Data: Simple Access to State-of-the-Art Information Accelerates Southeast Asia's Clean Energy Economic Transformation
High quality, robust, and reliable renewable energy resource data is foundational to climate-smart decision making, evidence-based policy planning, and clean energy investment mobilization. USAID and NREL, through the Advanced Energy Partnership for Asia, are expanding access to this critical resource data by providing free, high-fidelity solar resource data for Southeast Asia through the RE Data Explorer platform. This brief highlights several ways the Southeast Asia solar resource data has been used for power system planning and project development in the region.
Identifying Pathways for Enhanced Collaboration Between the Mining and Geothermal Industries
The locatable mineral industry is shifting toward improving environmental performance and becoming more sustainable, with numerous mining companies shifting to renewable energy technologies to power mine operations and at least one company pledging net-zero emissions by 2050. One potential electricity source to help achieve improved environmental performance and decarbonization within the mining industry is geothermal energy. As part of a study into potential collaboration between the geothermal and locatable mineral industries (focused on the portion of the Basin and Range Province within Nevada), the National Renewable Energy Laboratory (NREL) with support from the U.S. Department of Energy (DOE) Geothermal Technologies Office (GTO), investigated data, economic, and regulatory factors that may contribute to or inhibit synergies between the two industries. The objectives of the study included analyzing: The type and quality of data collected by the locatable mineral industry to determine feasibility for geothermal resource exploration; The regulatory pathways and potential barriers that could prevent development of geothermal resources discovered via a mining claim (and vice versa); The historical development of geothermal resources discovered via mineral exploration data; The value propositions for both the locatable mineral and the geothermal industries to collaborate.
High-Resolution Southeast Asia Wind Resource Data Set
Well informed decision-making is a key part of integrating variable renewable energy into the global energy marketplace. USAID and NREL, through the Advanced Energy Partnership for Asia, are expanding access to critical resource data by providing free, high-fidelity time-series wind resource data for Southeast Asia through the RE Data Explorer platform. This brief highlights the development of the Southeast Asia wind resource data set and discusses the impacts of this data.
MindSynchro
This report presents the developments and results of MindSynchro project as part of DOE OE FOA 1861. DOE and Pacific Northwest National Laboratory (PNNL) have made available to FOA awardees datasets containing years of real historical data recorded from various phasor measurement units (PMUs) which are installed in three large US interconnections: Texas (IC A), Western (IC B), and Eastern (IC C). The main goal of the project, which was successfully achieved, was to develop methods for detection and identification of events which are relevant for power grid operation. Tasks performed for achieving the project goals included data exploration and pre-processing, the development and application of physics-based features, data analysis and labeling based on unsupervised learning approaches, training and testing of DSSL models for classification of events which are relevant for power grid operation, and deployment of solutions to cloud environments. The methods developed in the project can potentially provide relevant benefits to power grid asset owners/operators in general in terms of situational awareness. Two main types of outcomes can be provided by these tools: Identification of specific relevant power grid event types: Semi-supervised ML methods developed in the project can adequately employ not only the relatively scarce labeled data but also the large amount of available unlabeled data to train models for detection of specific event types. Such methods enable the application of trained models for the detection of events in a population of PMUs much larger than that associated to the labeled events. Support in data labeling / label validation: Labels are critical for training of models for identification of specific types of events. However, labeling large amounts of data is a manual and tedious process. This means that such process is error prone and is not scalable. Methods developed in the project, based on ensembles of clustering models, have been successfully employed for turning manual labeling into a scalable process. Accurate identification of specific relevant events can provide the operators with immediate situational awareness that could otherwise require hours or days of analysis from domain experts. We envision that such methods could be initially employed in support of post-mortem analysis of events and, as confidence is gained, they could be employed for online/real-time support, providing, among other benefits, insights for avoiding major events which could happen due to a combination of smaller ones. On the longer term, related methods could potentially be employed to improve protection and control.
Multi-Sensor Data from A-Train Instruments Brought Together for Atmospheric Research
The A-Train is comprised of a series of instruments, developed independently, that measure highly related atmospheric components along the same flight path. In order to intercompare data from this multitude of sensors, researchers must access, subset, visualize, analyze and correlate distributed atmosphere measurements from the various A-Train instruments. The A-Train Data Depot (ATDD) has been operational for over a year, successfully performing the aforementioned functions on behalf of researchers, thus providing co-registered data from the Cloudsat, CALIOP, AIRS, and MODIS instruments for further intercomparisons. Of late, significant data from OM1 and POLDER are now included in the 'depot'. By specifying the desired spatial and temporal range, the researcher can subset, visualize, co-register, and access multi-sensor A-Train data related to: Cloud, aerosol, atmospheric temperature, and water vapor parameters (vertical profile visualizations); Cloud Pressure, cloud top temperature, water vapor, cloud optical thickness, and aerosol products (horizontal strips subsetted +/- 100km from the profile visualizations), and; Cloud pressure parameters (2-D line plots overlayed on the vertical profiles). All data is plotted using the GIOVANNI data exploration tool. A new feature of GIOVANNI is its ability to have collocated and subsetted data sets as well as PNG image files downloaded to the researcher's computing facility. By providing a convenient way to visualize and acquire multi-sensor data, ATDD affords users more time and effort to further their research.
NASA's Pilot Land Data System development program
The NASA Pilot Land Data System (PLDS) project is intended to enhance the effectiveness of data processing capabilities used by researchers applying remote sensing data in land science research. Two sites in the centerminous U.S. have been selected as study areas scanned by Landsat, Nimbus and GOES instruments. The data will be analyzed by teams of researchers representing different fields of expertise. The PLDS program will explore data management, networking and communications, system access capabilities, land analysis software, special processes and overall systems engineerng. The data will be processed by researchers working interactively through remote supermicrocomputer workstations using a variety of operating systems and on-site software capabilities.
The SAMPEX Data Processing Unit
The paper discusses salient features of the SAMPEX Data Processing Unit (DPU), the primary function of which is to collect sensor data to create telemetry packets for transmission to the solid-state recorder located within the Small Explorer Data System. Particular attention is given to the sensor interface electronics, the space command interface, the spacecraft telemetry interface, and the memory mapper of the DPU system; the task scheduling concept; and system reconfiguring. A block diagram of the DPU system is included.