Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Advancing Open Science in Atmospheric Research: Integrating Data Usability and Machine Learning

In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences." "In the dynamic realm of atmospheric sciences, the convergence of data science methodologies and open data marks a transformative era, driving research advancements and nurturing aspiring scientists. This abstract highlights two pivotal projects that epitomize open science principles, aligning seamlessly with the session's objective of interdisciplinary synergy and the cultivation of emerging talent. As a NASA-certified data center, our foremost endeavor focuses on enhancing the visibility and traceability of NASA datasets within atmospheric science research. This initiative not only elevates these datasets' prominence but also establishes a robust framework ensuring their credibility in scholarly discourse. By bridging the gap between data sources and research publications, this project serves as an educational catalyst, nurturing a new generation of scholars in open collaboration and dataset authenticity. Concurrently, our second project pioneers an early warning system for flooding events, utilizing machine learning algorithms to predict flooded fractions. Through multi-source data fusion and predictive modeling, this initiative goes beyond forecasting; it embodies the core of open science by enabling proactive risk mitigation strategies. This project not only advances atmospheric sciences but also fosters an environment where young scholars engage in practical, data-driven solutions. These intertwined projects exemplify the fusion of data science with open data solutions, ensuring both the usability of quality datasets and the cultivation of scientific knowledge among emerging scholars. By spotlighting these impactful use cases, our aim is to foster discussions emphasizing the importance of open collaboration, data integrity, and the nurturing of scientific talent in atmospheric sciences.

Jennifer Wei↗

Data Integration Support for Data Served in the OPeNDAP and OGC Environments

NASA is coordinating a technology development project to construct a gateway between system components built upon the Open-source Project for a Network Data AcceSs Protocol (OPeNDAP) and those made available made available via interfaces specified by the Open Geospatial Consortium (OGC). This project is funded though the Advanced Collaborative Connections for Earth-Sun System Science (ACCESS) Program and is a NASA contribution to the Committee on Earth Satellites (CEOS) Working Group on Information Systems and Services (WGISS). The motivation for the project is the set of data integration needs that have been expressed by the Coordinated Enhanced Observing Period (CEOP), an international program that is addressing the study of the global water cycle. CEOP is assembling a large collection in situ and satellite data and mode1 results from a wide variety of sources covering 35 sites around the globe. The data are provided by systems based on either the OPeNDAP or OGC protocols but the research community desires access to the full range of data and associated services from a single client. This presentation will discuss the current status of the OPeNDAP/OGC Gateway Project. The project is building upon an early prototype that illustrated the feasibility of such a gateway and which was demonstrated to the CEOP science community. In its first year as an ACCESS project, the effort has been has focused on the design of the catalog and data services that will be provided by the gateway and the mappings between the metadata and services provided in the two environments.

McDonald, Kenneth R.↗

CIRSS vertical data integration, San Bernardino study

The creation and use of a vertically integrated data base, including LANDSAT data, for local planning purposes in a portion of San Bernardino County, California are described. The project illustrates that a vertically integrated approach can benefit local users, can be used to identify and rectify discrepancies in various data sources, and that the LANDSAT component can be effectively used to identify change, perform initial capability/suitability modeling, update existing data, and refine existing data in a geographic information system. Local analyses were developed which produced data of value to planners in the San Bernardino County Planning Department and the San Bernardino National Forest staff.

Hodson, W.↗

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY↗

Prototype Local Data Integration System and Central Florida Data Deficiency

This report describes the Applied Meteorology Unit's (AMU) task on the Local Data Integration System (LDIS) and central Florida data deficiency. The objectives of the task are to identify all existing meteorological data sources within 250 km of the Kennedy Space Center (KSC) and the Eastern Range at Cape Canaveral Air Station (CCAS), identify and configure an appropriate LDIS to integrate these data, and implement a working prototype to be used for limited case studies and data non-incorporation (DNI) experiments. The ultimate goal for running LDIS is to generate products that may enhance weather nowcasts and short-range (less than 6 h) forecasts issued in support of the 45th Weather Squadron (45 WS), Spaceflight Meteorology Group (SMG), and the Melbourne National Weather Service (NWS MLB) operational requirements. The LDIS has the potential to provide added value for nowcasts and short term forecasts for two reasons. First, it incorporates all data operationally available in east central Florida. Second, it is run at finer spatial and temporal resolutions than current national-scale operational models. In combination with a suitable visualization tool, LDIS may provide users with a more complete and comprehensive understanding of evolving fine-scale weather features than could be developed by individually examining the disparate data sets over the same area and time. The utility of LDIS depends largely on the reliability and availability of observational data. Therefore, it is important to document all existing meteorological data sources around central Florida that can be incorporated by it. Several factors contribute to the data density and coverage over east central Florida including the level in the atmosphere, distance from KSC/CCAS, time, and prevailing weather. The central Florida mesonet consists of existing surface meteorological and hydrological data available from the Tampa NWS and data servers at Miami and Jacksonville. However the utility of these data for operational use is limited, mainly because there are relatively few additional meteorological observations within 50 km of KSC/CCAS to supplement existing METAR and KSC/CCAS tower reports.

Manobianco, John↗

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

Towards Integrating Data Quality Assessments and Radiometer Uncertainty for Determining the Expanded Uncertainty of Three-Component Solar Radiation Measurements

Accurate solar irradiance data are fundamental for determining the design and performance characteristics of photovoltaic systems. The uncertainty of solar irradiance measurements depends on many factors including radiometer design, calibration, installation, maintenance, and operational environment. The key contributors to this uncertainty can be classified as the measurement uncertainty of a particular radiometer and the operational uncertainty determined for the time of measurement. Radiometer measurement uncertainty estimates (U R ) can be based on well-established methods used as part of the radiometer calibration process. Estimates of operational uncertainties (U o ) require consideration of additional site-specific factors that affect data quality. A method is needed for establishing the accuracy of solar irradiance data by integrating an existing data quality process and measurement uncertainty estimates for specific radiometers. An algorithm has been developed to integrate data quality analyses and measurement uncertainty estimates for three-component solar irradiance data: global horizontal (total hemispheric) irradiance, direct normal (beam) irradiance, and diffuse horizontal (sky) irradiance collected at one- to 60-minute intervals. The algorithm has been tested using one-minute irradiance measurements. The goal of the project is to distribute a user-friendly software package based on the new algorithm.

data integrity↗

An integrated data management and informatics framework for continuous drug product manufacturing processes: A case study on two pilot plants

The pharmaceutical industry continuously looks for ways to improve its development and manufacturing efficiency. In recent years, such efforts have been driven by the transition from batch to continuous manufacturing and digitalization in process development. To facilitate this transition, integrated data management and informatics tools need to be developed and implemented within the framework of Industry 4.0 technology. Here, in this regard, the work aims to guide the data integration development of continuous pharmaceutical manufacturing processes under the Industry 4.0 framework, improving digital maturity and enabling the development of digital twins. This paper demonstrates two instances where a data integration framework has been successfully employed in academic continuous pharmaceutical manufacturing pilot plants. Details of the integration structure and information flows are comprehensively showcased. Approaches to mitigate concerns in incorporating complex data streams, including integrating multiple process analytical technology tools and legacy equipment, connecting cloud data and simulation models, and safeguarding cyber-physical security, are discussed. Critical challenges and opportunities for practical considerations are highlighted.

59 BASIC BIOLOGICAL SCIENCES↗

Aviation Data Integration System

During the analysis of flight data and safety reports done in ASAP and FOQA programs, airline personnel are not able to access relevant aviation data for a variety of reasons. We have developed the Aviation Data Integration System (ADIS), a software system that provides integrated heterogeneous data to support safety analysis. Types of data available in ADIS include weather, D-ATIS, RVR, radar data, and Jeppesen charts, and flight data. We developed three versions of ADIS to support airlines. The first version has been developed to support ASAP teams. A second version supports FOQA teams, and it integrates aviation data with flight data while keeping identification information inaccessible. Finally, we developed a prototype that demonstrates the integration of aviation data into flight data analysis programs. The initial feedback from airlines is that ADIS is very useful in FOQA and ASAP analysis.

Kulkarni, Deepak↗

Towards physics-inspired data-driven weather forecasting: integrating data assimilation with a deep spatial-transformer-based U-NET in a case study with ERA5

Abstract. There is growing interest in data-driven weather prediction (DDWP), e.g., using convolutional neural networks such as U-NET that are trained on data from models or reanalysis. Here, we propose three components, inspired by physics, to integrate with commonly used DDWP models in order to improve their forecast accuracy. These components are (1) a deep spatial transformer added to the latent space of U-NET to capture rotation and scaling transformation in the latent space for spatiotemporal data, (2) a data-assimilation (DA) algorithm to ingest noisy observations and improve the initial conditions for next forecasts, and (3) a multi-time-step algorithm, which combines forecasts from DDWP models with different time steps through DA, improving the accuracy of forecasts at short intervals. To show the benefit and feasibility of each component, we use geopotential height at 500 hPa (Z500) from ERA5 reanalysis and examine the short-term forecast accuracy of specific setups of the DDWP framework. Results show that the spatial-transformer-based U-NET (U-STN) clearly outperforms the U-NET, e.g., improving the forecast skill by 45 %. Using a sigma-point ensemble Kalman (SPEnKF) algorithm for DA and U-STN as the forward model, we show that stable, accurate DA cycles are achieved even with high observation noise. This DDWP+DA framework substantially benefits from large (O(1000)) ensembles that are inexpensively generated with the data-driven forward model in each DA cycle. The multi-time-step DDWP+DA framework also shows promise; for example, it reduces the average error by factors of 2–3. These results show the benefits and feasibility of these three components, which are flexible and can be used in a variety of DDWP setups. Furthermore, while here we focus on weather forecasting, the three components can be readily adopted for other parts of the Earth system, such as ocean and land, for which there is a rapid growth of data and need for forecast and assimilation.

54 ENVIRONMENTAL SCIENCES↗

DIP Architecture and Data Integration Services

This workshop will cover DIP architecture and data integration services. Participants will get a look at how the DIP architecture is set-up as well as how data integration services are planned to be hosted on the platform. The DIP architecture review is intended to cover how DIP was envisioned and how DIP is being developed to address data needs across the industry. Participants will have a chance to provide feedback on the DIP architecture and gain insight into how one might interface with the DIP to send or receive data. The data integration services portion is intended to cover DIP’s technical approach to data integration. As an example implementation, there will be a first look at possible data fusion on the platform, including utilizing NASA’s Fuser, and tailoring for industry data consumers. Descriptions, at a high-level, of input to and output of the Fuser will also be discussed.

ATM-X↗

Recap of DIP Workshop Series: #1 DIP Architecture and Data Integrations Services

This workshop will cover DIP architecture and data integration services to obtain feedback from America for Airlines (A4A) Air Traffic Management Council (ATMC). Participants will get a look at how the DIP architecture is set-up as well as how data integration services are planned to be hosted on the platform. The DIP architecture review is intended to cover how DIP was envisioned and how DIP is being developed to address data needs across the industry. Participants will have a chance to provide feedback on the DIP architecture and gain insight into how one might interface with the DIP to send or receive data. The data integration services portion is intended to cover DIP’s technical approach to data integration. As an example implementation, there will be a first look at possible data fusion on the platform, including utilizing NASA’s Fuser, and tailoring for industry data consumers. Descriptions, at a high-level, of input to and output of the Fuser will also be discussed.

ATM-X↗

A conceptual design for an integrated data base management system for remote sensing data

The requirements of potential users were considered in the design of an integrated data base management system, developed to be independent of any specific computer or operating system, and to be used to support investigations in weather and climate. Ultimately, the system would expand to include data from the agriculture, hydrology, and related Earth resources disciplines. An overview of the system and its capabilities is presented. Aspects discussed cover the proposed interactive command language; the application program command language; storage and tabular data maintained by the regional data base management system; the handling of data files and the use of system standard formats; various control structures required to support the internal architecture of the system; and the actual system architecture with the various modules needed to implement the system. The concepts on which the relational data model is based; data integrity, consistency, and quality; and provisions for supporting concurrent access to data within the system are covered in the appendices.

Maresca, P. A.↗

An emerging network storage management standard: Media error monitoring and reporting information (MEMRI) - to determine optical tape data integrity

Sophisticated network storage management applications are rapidly evolving to satisfy a market demand for highly reliable data storage systems with large data storage capacities and performance requirements. To preserve a high degree of data integrity, these applications must rely on intelligent data storage devices that can provide reliable indicators of data degradation. Error correction activity generally occurs within storage devices without notification to the host. Early indicators of degradation and media error monitoring 333 and reporting (MEMR) techniques implemented in data storage devices allow network storage management applications to notify system administrators of these events and to take appropriate corrective actions before catastrophic errors occur. Although MEMR techniques have been implemented in data storage devices for many years, until 1996 no MEMR standards existed. In 1996 the American National Standards Institute (ANSI) approved the only known (world-wide) industry standard specifying MEMR techniques to verify stored data on optical disks. This industry standard was developed under the auspices of the Association for Information and Image Management (AIIM). A recently formed AIIM Optical Tape Subcommittee initiated the development of another data integrity standard specifying a set of media error monitoring tools and media error monitoring information (MEMRI) to verify stored data on optical tape media. This paper discusses the need for intelligent storage devices that can provide data integrity metadata, the content of the existing data integrity standard for optical disks, and the content of the MEMRI standard being developed by the AIIM Optical Tape Subcommittee.

Podio, Fernando↗

Local Data Integration in East Central Florida Using the ARPS Data Analysis System

This paper describes the Applied Meteorology Unit's (AMU) efforts to configure, implement, and test a version of the Advanced Regional Prediction System (ARPS) Data Analysis System (ADAS; Brewster 1996) that assimilates all available data within 250 km of the Kennedy Space Center (KSC) and the Eastern Range at Cape Canaveral Air Station (CCAS). The objective for running a Local Data Integration System (LDIS) such as ADAS is to generate products which may enhance weather nowcasts and short-range (less than 6 h) forecasts issued in support of ground and aerospace operations at KSC/CCAS. A LDIS such as ADAS has the potential to provide added value because it combines observational data to produce gridded analyses of temperature, wind, and moisture (including clouds) and diagnostic quantities such as vorticity, divergence, etc. at specified temporal and spatial resolutions. In this regard, a LDTS along with suitable visualization tools may provide users with a ignore complete and comprehensive understanding of evolving weather than could be developed by individually examining the disparate data sets over the same area and time. The AMU implemented a working prototype of the ADAS which does not run in real-time. Instead, the AMU is evaluating ADAS through post-analyses of weather events for a warm and cool season case. The case studies were chosen to investigate the capabilities and limitations of a LDIS such as ADAS including the impact of non-incorporation of specific data sources on the utility of the subsequent analyses.

Case, Jonathan↗