Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Hydraulic fracture characterization by integrating multidisciplinary data from the Hydraulic Fracturing Test Site 2 (HFTS-2)

Various technologies have traditionally been used to monitor and describe hydraulic fractures from different perspectives. This work demonstrates the value of data integration for hydraulic fracture characterization when multiple data resources are available. The Hydraulic Fracturing Test Site 2 (HFTS 2) is a hydraulic fracturing research project in the Delaware Basin with multiple surveillance techniques including fiber optics sensing, microseismic, pressure/temperature gauges, etc. We integrated the multidisciplinary data from the HFTS-2 to characterize hydraulic fractures. The integrated data revealed interesting fracture propagation features including layering, vertical propagation affected by pore pressure gradient, and different microseismic activities due to difference in-situ conditions. Furthermore, these findings can be insightful for understanding hydraulic fracture propagation. The comparison among multiple surveillance data also helps us to evaluate the roles of various surveillance technologies and provides us experience to make informative decisions depending on different monitoring objectives.

58 GEOSCIENCES↗

CORAL: A framework for rigorous self-validated data modeling and integrative, reproducible data analysis

Abstract Background Many organizations face challenges in managing and analyzing data, especially when relevant datasets arise from multiple sources and methods. Analyzing heterogeneous datasets and additional derived data requires rigorous tracking of their interrelationships and provenance. This task has long been a Grand Challenge of data science and has more recently been formalized in the FAIR principles: that all data objects be Findable, Accessible, Interoperable, and Reusable, both for machines and for people. Adherence to these principles is necessary for proper stewardship of information, for testing regulatory compliance, for measuring the efficiency of processes, and for facilitating reuse of data-analytical frameworks. Findings We present the Contextual Ontology-based Repository Analysis Library (CORAL), a platform that greatly facilitates adherence to all 4 of the FAIR principles, including the especially difficult challenge of making heterogeneous datasets Interoperable and Reusable across all parts of a large, long-lasting organization. To achieve this, CORAL's data model requires that data generators extensively document the context for all data, and our tools maintain that context throughout the entire analysis pipeline. CORAL also features a web interface for data generators to upload and explore data, as well as a Jupyter notebook interface for data analysts, both backed by a common API. Conclusions CORAL enables organizations to build FAIR data types on the fly as they are needed, avoiding the expense of bespoke data modeling. CORAL provides a uniquely powerful platform to enable integrative cross-dataset analyses, generating deeper insights than are possible using traditional analysis tools.

97 MATHEMATICS AND COMPUTING↗

Snow Process Estimation Over the Extratropical Andes Using a Data Assimilation Framework Integrating MERRA Data and Landsat Imagery

A data assimilation framework was implemented with the objective of obtaining high resolution retrospective snow water equivalent (SWE) estimates over several Andean study basins. The framework integrates Landsat fractional snow covered area (fSCA) images, a land surface and snow depletion model, and the Modern Era Retrospective Analysis for Research and Applications (MERRA) reanalysis as a forcing data set. The outputs are SWE and fSCA fields (1985-2015) at a resolution of 90 m that are consistent with the observed depletion record. Verification using in-situ snow surveys showed significant improvements in the accuracy of the SWE estimates relative to forward model estimates, with increases in correlation (0.49-0.87) and reductions in root mean square error (0.316 m to 0.129 m) and mean error (-0.221 m to 0.009 m). A sensitivity analysis showed that the framework is robust to variations in physiography, fSCA data availability and a priori precipitation biases. Results from the application to the headwater basin of the Aconcagua River showed how the forward model versus the fSCA-conditioned estimate resulted in different quantifications of the relationship between runoff and SWE, and different correlation patterns between pixel-wise SWE and ENSO. The illustrative results confirm the influence that ENSO has on snow accumulation for Andean basins draining into the Pacific, with ENSO explaining approximately 25% of the variability in near-peak (1 September) SWE values. Our results show how the assimilation of fSCA data results in a significant improvement upon MERRA-forced modeled SWE estimates, further increasing the utility of the MERRA data for high-resolution snow modeling applications.

SWE↗

Demonstration and Evaluation of a Non-Invasive, Low-Cost, Strap-On Sensor for Natural Gas Meters

The U.S. General Services Administration (GSA) is interested in installing internet-connected, gas submeters to better understand gas consumption in its portfolio of buildings. The GSA in partnership with the National Renewable Energy Laboratory conducted a demonstration to assess a specific submeter technology. This technology was implemented at two GSA separate facilities located in Dallas, Texas. This demonstration evaluated hardware and software installations and integrations, data integrity and accuracy, and included an economic analysis. This demonstration evaluated a gas submeter technology provided by the vendor Vata Verks, who produces a non-invasive strap-on submeter. The company was founded with a mission to conserve water (and eventually gas) cheaply and simply, as explained on its website. The product intends to streamline submeter deployments for gas and water by eliminating most hardware costs while allowing for easy integration of submetered data into other systems such as the Building Automation System and removing tenant and building disruption This demonstration evaluated the product features when used on a gas utility meter.

03 NATURAL GAS↗

A-Train Data Depot: Integrating and Exploring Data Along the A-Train Tracks

The immense potential for new science findings as a result of inter-instrument data analysis has led to the development of a new data portal at GSFC: the A-train Data Depot. The power and utility of this new service to the general public is amplified immensely when the archived data are used in conjunction with online data analysis services like Giovanni. This presentation details some of the challenges of data usage from multiple distinct missions and how the tool sets we have developed can help to overcome these challenges, considerably cut down on analysis overhead and promote science exploration in an otherwise very challenging arena.

Leptoukh, G.↗

CO 2 Storage Site Screening Platform Development and CO 2 Storage Resource Analysis in SECARB Offshore Reservoirs Using SAS Viya

A major goal of the SECARB Offshore Partnership (DE-FE0031557) is to screen deep saline aquifers and hydrocarbon reservoirs in the central Gulf of Mexico for CO 2 sequestration and CO 2 -enhanced oil and gas recovery (EOR/EGR) and estimate the corresponding CO 2 storage resources for select reservoirs. CO 2 storage potential associated with offshore CO 2 -EOR is considerable and likely represents “low hanging fruit” for near-term CO 2 storage given the in-place infrastructure in the region. It is for these reasons that this assessment focuses on oil and gas fields. To this end, three major objectives have been completed and include (1) managing geological data derived from different sources, (2) building a reservoir screening platform for CO 2 storage, and (3) ranking the reservoirs based on the estimated CO 2 storage resources. The SAS ® Viya platform was used for data management and analytics. The Viya platform is a cloud service platform that provides data integration, data management, quick analytics, data visualization, machine learning functions, and application programming interfaces (APIs) for multi-programming languages. Different sources of data containing geologic information, reservoir properties, and EOR/EGR information were collected, cleaned, formatted, and loaded into the SAS ® Viya platform for evaluation. The major geological characteristics of both shelf and deep-water areas of the central Gulf were examined and compared to define the appropriate reservoir screening criteria. Next, a CO 2 storage site screening system was built in the SAS ® Viya platform with the pre-defined criteria. Finally, the CO 2 storage resources of the screened reservoirs were calculated and reported at the BOEM field level to identify fields with the highest estimated CO 2 storage resource. The fields with the largest total estimated CO 2 storage resource are located in the Mississippi Canyon protraction area. Due to proximity to the Mississippi Delta (indicative of less infrastructure) and large estimated CO 2 storage resources, future development activities may wish to focus efforts in the Mississippi Canyon protraction area.

02 PETROLEUM↗

CIF Report - Information Fusion and Data Analytics for Human Lunar Exploration

This project leverages the Concept Exploration Laboratory (CEL) to collect, warehouse, and augment data relevant to human lunar exploration as a platform for NA (S&MA) to develop operational data integration techniques. The project capitalizes on 16+ years of CEL experience applied to NASA, DoD, the City of Houston, the State of Texas, and private industry. The integrated data will be utilized in the two scenarios described in a definition of concept for development of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, and may be useful as pathfinders for the Digital Transformation Program.

information fusion↗

System for Secure Integration of Aviation Data

The Aviation Data Integration System (ADIS) of Ames Research Center has been established to promote analysis of aviation data by airlines and other interested users for purposes of enhancing the quality (especially safety) of flight operations. The ADIS is a system of computer hardware and software for collecting, integrating, and disseminating aviation data pertaining to flights and specified flight events that involve one or more airline(s). The ADIS is secure in the sense that care is taken to ensure the integrity of sources of collected data and to verify the authorizations of requesters to receive data. Most importantly, the ADIS removes a disincentive to collection and exchange of useful data by providing for automatic removal of information that could be used to identify specific flights and crewmembers. Such information, denoted sensitive information, includes flight data (here signifying data collected by sensors aboard an aircraft during flight), weather data for a specified route on a specified date, date and time, and any other information traceable to a specific flight. The removal of information that could be used to perform such tracing is called "deidentification." Airlines are often reluctant to keep flight data in identifiable form because of concerns about loss of anonymity. Hence, one of the things needed to promote retention and analysis of aviation data is an automated means of de-identification of archived flight data to enable integration of flight data with non-flight aviation data while preserving anonymity. Preferably, such an automated means would enable end users of the data to continue to use pre-existing data-analysis software to identify anomalies in flight data without identifying a specific anomalous flight. It would then also be possible to perform statistical analyses of integrated data. These needs are satisfied by the ADIS, which enables an end user to request aviation data associated with de-identified flight data. The ADIS includes client software integrated with other software running on flight-operations quality-assurance (FOQA) computers for purposes of analyzing data to study specified types of events or exceedences (departures of flight parameters from normal ranges). In addition to ADIS client software, ADIS includes server hardware and software that provide services to the ADIS clients via the Internet (see figure). The ADIS server receives and integrates flight and non-flight data pertaining to flights from multiple sources. The server accepts data updates from authorized sources only and responds to requests from authorized users only. In order to satisfy security requirements established by the airlines, (1) an ADIS client must not be accessible from the Internet by an unauthorized user and (2) non-flight data as airport terminal information system (ATIS) and weather data must be displayed without any identifying flight information. ADIS hardware and software architecture as well as encryption and data display scheme are designed to meet these requirements. When a user requests one or more selected aviation data characteristics associated with an event (e.g., a collision, near miss, equipment malfunction, or exceedence), the ADIS client augments the request with date and time information from encrypted files and submits the augmented request to the server. Once the user s authorization has been verified, the server returns the requested information in de-identified form.

Kulkarni, Deepak↗

Integrating multimodal data through interpretable heterogeneous ensembles

Motivation: Integrating multimodal data represents an effective approach to predicting biomedical characteristics, such as protein functions and disease outcomes. However, existing data integration approaches do not sufficiently address the heterogeneous semantics of multimodal data. In particular, early and intermediate approaches that rely on a uniform integrated representation reinforce the consensus among the modalities but may lose exclusive local information. The alternative late integration approach that can address this challenge has not been systematically studied for biomedical problems. Results: We propose Ensemble Integration (EI) as a novel systematic implementation of the late integration approach. EI infers local predictive models from the individual data modalities using appropriate algorithms and uses heterogeneous ensemble algorithms to integrate these local models into a global predictive model. We also propose a novel interpretation method for EI models. We tested EI on the problems of predicting protein function from multimodal STRING data and mortality due to coronavirus disease 2019 (COVID-19) from multimodal data in electronic health records. We found that EI accomplished its goal of producing significantly more accurate predictions than each individual modality. It also performed better than several established early integration methods for each of these problems. The interpretation of a representative EI model for COVID-19 mortality prediction identified several disease-relevant features, such as laboratory test (blood urea nitrogen and calcium) and vital sign measurements (minimum oxygen saturation) and demographics (age). These results demonstrated the effectiveness of the EI framework for biomedical data integration and predictive modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES↗

NASA's GeneLab: An Integrated Omics Data Commons and Workbench

GeneLab (http://genelab.nasa.gov) is a NASA initiative designed to accelerate “open science” biomedical research in support of the human exploration of space and the improvement of life on earth. The GeneLab Data Systems (GLDS) were developed to help investigators corroborate findings from “omics” (genomics, transcriptomics, proteomics, and metabolomics) assays and translate them into systems biology knowledge and, eventually, therapeutics, including countermeasures to support life in space. Phase I of the project (completed) emphasized developing key capabilities for submission, curation, storage, search, and retrieval of omics data from biomedical research in and of space environments. The development focus for Phase II (completed) was federated data search and retrieval of these kinds of data from other open-access repositories. The last phase of the project (in work) entails developing an omics analysis tool set, and a portal to visualize processed omics data, emphasizing integration with the data repository and search functions developed during the prior phases. The final product will be an open-access system where users can individually or collaboratively publish, search, integrate, analyze, and visualize omics data.

genome↗