Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A data integration framework of additive manufacturing based on FAIR principles

Abstract Laser-powder bed fusion (L-PBF) is a popular additive manufacturing (AM) process with rich data sets coming from both in situ and ex situ sources. Data derived from multiple measurement modalities in an AM process capture unique features but often have different encoding methods; the challenge of data registration is not directly intuitive. In this work, we address the challenge of data registration between multiple modalities. Large data spaces must be organized in a machine-compatible method to maximize scientific output. FAIR (findable, accessible, interoperable, and reusable) principles are required to overcome challenges associated with data at various scales. FAIRified data enables a standardized format allowing for opportunities to generate automated extraction methods and scalability. We establish a framework that captures and integrates data from a L-PBF study such as radiography and high-speed camera video, linking these data sets cohesively allowing for future exploration. Graphical abstract

36 MATERIALS SCIENCE↗

NASA and Advanced Air Mobility (AAM) National Campaign (NC) Tech Talk: Integrated Data Product

NASA Advanced Air Mobility (AAM) National Campaign has hosted multiple flight tests in a effort to promote public confidence and accelerate the realization of emerging aviation markets for passenger and cargo transportation in urban, suburban, rural, and regional environments. The goals and objectives for each flight test are intended to assess the performance of the vehicle’s capacity to operate safely, efficiently, and sustainably in urban areas. The Airspace Test and Infrastructure (ATI) team processes the flight data to evaluate the effectiveness of air vehicle and airspace management integration, while its analysis capabilities support the assessment of the efficacy of proposed vehicles and ground systems required to enable the vehicles’ operations. The ATI data services team creates many products from the data, first by transforming and merging the data into a single Integrated Data Product (IDP). This aligns data from many sources by time, so that, for example, the winds that impacted a flight can be aligned with its ground track and speed. A detailed IDP enables many types of statistical computation and visualizations that help researchers understand what transpired during the flight and how it can be explained. This Tech Talk will give an overview of the IDP process, including examples of the resulting products created by the data services team. It will also share an application that will allow participants to perform their own data analysis.

National Campaign, Advanced Air Mobility, Integrat↗

LSTM-Based Data Integration to Improve Snow Water Equivalent Prediction and Diagnose Error Sources

Accurate prediction of snow water equivalent (SWE) can be valuable for water resource managers. Recently, deep learning methods such as long short-term memory (LSTM) have exhibited high accuracy in simulating hydrologic variables and can integrate lagged observations to improve prediction, but their benefits were not clear for SWE simulations. Here we tested an LSTM network with data integration (DI) for SWE in the western United States to integrate 30-day-lagged or 7-day-lagged observations of either SWE or satellite-observed snow cover fraction (SCF) to improve future predictions. SCF proved beneficial only for shallow-snow sites during snowmelt, while lagged SWE integration significantly improved prediction accuracy for both shallow- and deep-snow sites. The median Nash–Sutcliffe model efficiency coefficient (NSE) in temporal testing improved from 0.92 to 0.97 with 30-day-lagged SWE integration, and root-mean-square error (RMSE) and the difference between estimated and observed peak SWE values d max were reduced by 41% and 57%, respectively. DI effectively mitigated accumulated model and forcing errors that would otherwise be persistent. Moreover, by applying DI to different observations (30-day-lagged, 7-day-lagged), we revealed the spatial distribution of errors with different persistent lengths. For example, integrating 30-day-lagged SWE was ineffective for ephemeral snow sites in the southwestern United States, but significantly reduced monthly-scale biases for regions with stable seasonal snowpack such as high-elevation sites in California. These biases are likely attributable to large interannual variability in snowfall or site-specific snow redistribution patterns that can accumulate to impactful levels over time for nonephemeral sites. These results set up benchmark levels and provide guidance for future model improvement strategies.

54 ENVIRONMENTAL SCIENCES↗

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES↗

The Role of Snowmelt and Subsurface Heterogeneity in Headwater Hydrology of a Mountainous Catchment in Colorado: A Model‐Data Integration Approach

Mountainous headwater streams are sustained by both snowmelt‐driven streamflow and groundwater discharge in the Upper Colorado River Basin. However, predicting headwater stream discharge magnitude and peak flow timing is challenging in mountainous terrains, where snowmelt rates vary with vegetation type and elevation, and heterogeneous subsurface physical properties influence groundwater storage and its release. We used a model‐data integration approach to investigate the roles of snowmelt and subsurface structure in stream discharge and groundwater level. We ran an ensemble of 100 integrated surface‐subsurface hydrologic models for a mountainous headwater catchment near Crested Butte, Colorado, USA. We also evaluated and calibrated these models against observed data sets, including snow depth measurements using distributed temperature probes, stream discharge, and groundwater levels. Calibration with multiple data sources using neural density estimators has further constrained uncertainty in subsurface properties and snowmelt rates. Results indicated that observed slower snowmelt rates in evergreen forests delayed the peak flow and baseflow onset. In upstream areas with lower subsurface permeability, water was stored within the subsurface but was not released as interflow or shallow groundwater flow, and thereby not contributing to downstream streamflow during recession limb periods. Double peaks in groundwater occurred in areas with spatial subsurface heterogeneity, in our case due to the contrast between granodiorite and Mancos shale. These process‐based insights into groundwater and snowmelt dynamics in mountainous headwaters will help improve predictions of headwater hydrology.

Wang, Lijing [University of Connecticut, Storrs, C↗

Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower (DIVERS-H)

U.S. hydropower plants face potential threats from shrinking water supply, rising demands, and warmer stream temperatures from various causes. Power plant owners, operators, and regulators require new tools to take advantage of and interpret the diverse range of scientific data being produced by both observational methods (for example, satellite, radar, stream gauges) and computer modeling methods that evaluate and predict how earth's dynamic systems (atmosphere, oceans, land surface, and sea ice) are changing and interacting. Combining datasets such as these with AI-based analyses introduces a novel decision support system to help users anticipate and address potential impacts on power generation stations. This new technology has been named DIVERS-H for "Data Integration and Visualization for Enhanced Resilience and Sustainability in Hydropower." In Phase I, technical feasibility was established with the development and demonstration of all the new technologies that are required. Most notably, DIVERS-H will use new artificial intelligence (AI) methods to capture the complex dynamics of water availability, demand, and environmental changes. In addition, new data management software was developed, and a prototype user interface was implemented as the precursor to a full scale decision support system. With technical research complete, the project focus now shifts to development of a commercial software product to provide users with actionable insight into water availability and the risk/resilience of critical systems at their locations of interest. Although DIVER-H was originally conceived as a tool for hydroelectric power applications, the same underlying technology can be readily applied to other water-consuming systems including coal, natural gas, oil, and nuclear power plants.

Chaudhary, Aashish [Kitware, Inc., Clifton Park, N↗

Open Data Integration (ODIN): A Concurrent, Distributed Message-Based Architecture and Framework for Disaster Response

The Runtime for Airspace Concept Evaluation (RACE) is an open-source software architecture and framework to build configurable, highly concurrent and distributed message-based systems that offer scalable, low-latency performance on commodity hardware. RACE was used in commercial aviation applications to rapidly build systems that span several machines (including synchronized displays), interface existing hardware simulators and other live data feeds, and incorporate sophisticated visualization components such as NASA WorldWind. These RACE applications validated elements of the FAA’s System Wide Information Management (SWIM) Program, handling up to 1000 messages/sec from diverse sources (SFDPS, TFM-DATA, TAIS, ASDE-X, ITWS and local ADS) for 4,500 simultaneous flights tracked in the next-generation air transportation system’s digital backbone. We have since generalized RACE to support Open Data Integration (ODIN) applications outside aviation. Systems built with RACE/ODIN can be deployed in the field, on commodity hardware, and operate with limited or intermittent connectivity to the outside world. Our primary use case is a web-server with local/persistent data storage that runs within and only serves the stakeholder network (e.g. an incident command post). We are tailoring the RACE/ODIN system to support wildland fire management for the upcoming NASA Wildland Fire Safety Demonstration Series. RACE-ODIN is under consideration for application in the Scalable Traffic Management for Emergency Response Operations project, or STEReO, which aims to create a system that can be deployed during emergencies, to coordinate multiple elements of disaster response. Such data sources predominantly come from existing services on the internet (e.g. weather and satellite data, imported from so called "edge servers") but can also include dynamic (real-time) data from computer simulations and within the stakeholder network (such as aircraft and personnel tracking information). We will present the architecture and ODIN system demonstration incorporating local data from instrumented power-line towers, interpolated weather data and geospatial data from space-based platforms.

Joseph C Coughlan↗

Open Data Integration (ODIN): A Concurrent, Distributed Message-Based Architecture and Framework for Disaster Response

The Runtime for Airspace Concept Evaluation (RACE) is an open-source software architecture and framework to build configurable, highly concurrent and distributed message-based systems that offer scalable, low-latency performance on commodity hardware. RACE was used in commercial aviation applications to rapidly build systems that span several machines (including synchronized displays), interface existing hardware simulators and other live data feeds, and incorporate sophisticated visualization components such as NASA WorldWind. These RACE applications validated elements of the FAA’s System Wide Information Management (SWIM) Program, handling up to 1000 messages/sec from diverse sources (SFDPS, TFM-DATA, TAIS, ASDE-X, ITWS and local ADS) for 4,500 simultaneous flights tracked in the next-generation air transportation system’s digital backbone. We have since generalized RACE to support Open Data Integration (ODIN) applications outside aviation. Systems built with RACE/ODIN can be deployed in the field, on commodity hardware, and operate with limited or intermittent connectivity to the outside world. Our primary use case is a web-server with local/persistent data storage that runs within and only serves the stakeholder network (e.g. an incident command post). We are tailoring the RACE/ODIN system to support wildland fire management for the upcoming NASA Wildland Fire Safety Demonstration Series. RACE-ODIN is under consideration for application in the Scalable Traffic Management for Emergency Response Operations project, or STEReO, which aims to create a system that can be deployed during emergencies, to coordinate multiple elements of disaster response. Such data sources predominantly come from existing services on the internet (e.g. weather and satellite data, imported from so called "edge servers") but can also include dynamic (real-time) data from computer simulations and within the stakeholder network (such as aircraft and personnel tracking information). We will present the architecture and ODIN system demonstration incorporating local data from instrumented power-line towers, interpolated weather data and geospatial data from space-based platforms.

Guillaume P Brat↗

Crew Health and Performance Integrated Data Architecture (CHP-IDA) TechPort May 2024

Future exploration missions to Mars will have increased need for crew autonomy. Crew Health & Performance (CHP) related data on the ISS is currently, manually downlinked and in disparate locations, which limits crew autonomy for future missions. The CHP-IDA project is developing a backend data system platform that grants the ability to seamlessly collect, store, process, and display CHP-related data for exploration missions. This platform allows for integration of data and advanced analytics that offer crew and ground teams better insight into the crew’s health and performance. It also enables applications that can improve the crew’s ability to provide more autonomous medical care during exploration missions. Data will be collected automatically to reduce crew and ground team time and effort and will synchronize across all in-mission vehicles, habitats, and ground as communication delay permits. The Human Research Program’s (HRP) Medical Data Architecture (MDA) project focused on this backend data architecture but for medical data only. The CHP-IDA project, a joint effort between HRP’s Exploration Medical Capability (ExMC) element and the Exploration Medical Integrated Product Team (XMIPT), expands this capability to all relevant CHP-related data. The additional inputs from nutrition, environment, exercise, radiation, and any other relevant sources will give more insight into crew’s health and performance. Currently, the Human Systems Engineering and Integration Division at Johnson Space Center (JSC) is designing the system. The team completed a system requirements review (SRR) in FY22 and now the focus is on core software development, testbed buildup, and use case scenario demonstration. An end-to-end demonstration with multiple data sources across CHP domains is schedule for the end of FY24 where all three focus areas will be displayed. Following this ground demo, the software will be completed, tested, and validated for flight.

Courtney M Schkurko↗

Accelerating nuclear-integrated data center pursuits in the USA: SWOT analysis, power-thermal management strategies and demonstration plan

Here, this study explores the increasing interest in leveraging nuclear power to meet the escalating energy demands of data centers in the United States (U.S.) by focusing on key factors that contribute to accelerated deployment. The study highlights the importance of N+1/N+2 power supplies (where N is the required number of units), outlines research and innovations in nuclear-integrated data center thermal management and demonstration plan. It also provides updates about status and costing of various reactor system designs. A summarized strengths, weaknesses, opportunities, and threats (SWOT) analysis shows the potential options for grid connectivity, reactors, and site selection. Suitable site discussions consider land and water availability, grid access, and optical fiber connectivity, and the study presents graded prospects for Department of Energy (DOE) sites with a specific example. Community engagement and partnerships are emphasized, particularly the roles of local government, federal agencies, utilities, and data center industry partners, which are crucial for accelerating deployment, business outreach, and approvals. The study provides actionable insights for stakeholders to accelerate the deployment of nuclear-powered data centers.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Machine learning-enabled model-data integration for predicting subsurface water storage

Subsurface water storage (SWS) is a key variable of the climate system and a storage component for precipitation and radiation anomalies, inducing persistence in the climate system. It plays a critical role in climate-change projections and can mitigate the impacts of climate change on ecosystems. However, because of the difficult accessibility of the underground, hydrologic properties and dynamics of SWS are poorly known. Direct observations of SWS are limited, and accurate incorporation of SWS dynamics into Earth system land models remains challenging. We propose a machine learning-enabled model-data integration framework to improve the SWS prediction at local to conus scales in a changing climate by leveraging all the available observation and simulation resources, as well as to inform the model development and guide the observation collection. The accurate prediction will enable an optimal decision of water management and land use and improve the ecosystem's resilience to the climate change.

Lu, Dan↗

Verifying Data Integrity of Electronically Scanned Pressure Systems at the NASA Glenn Research Center

The proper operation of the Electronically Scanned Pressure (ESP) System critical to accomplish the following goals: acquisition of highly accurate pressure data for the development of aerospace and commercial aviation systems and continuous confirmation of data quality to avoid costly, unplanned, repeat wind tunnel or turbine testing. Standard automated setup and checkout routines are necessary to accomplish these goals. Data verification and integrity checks occur at three distinct stages, pretest pressure tubing and system checkouts, daily system validation and in-test confirmation of critical system parameters. This paper will give an overview of the existing hardware, software and methods used to validate data integrity.

Panek, Joseph W.↗

Battery data integrity and usability: Navigating datasets and equipment limitations for efficient and accurate research into battery aging

A tremendous commitment of resources is needed to acquire, understand and apply battery data in terms of performance and aging behavior. There are many state of performance (SOP) and state of health (SOH) metrics that are useful to guide alignment of batteries to end-use, yet how these metrics are measured or extracted can make the difference between usable, valuable datasets versus data that lacks the necessary integrity to meet baseline confidence levels for SOP/SOH quantification. This work will speak to 1) types of data that support SOP and SOH evaluations on mechanistic terms, 2) measurement conditions needed to assure high data integrity, 3) equipment limitations that can compromise data high fidelity, and 4) the impact of cell polarization on data quality. A common goal in battery research and field use is to work from a data platform that supports economical paths of data capture while minimizing down-time for battery diagnostics. An ideal situation would be to utilize data obtained during normal daily use (“pulses or cycles of convenience”) without stopping the daily duty cycles to perform dedicated SOP/SOH diagnostic routines. However, difficulties arise in trying to make use of daily duty cycle data (denoted as cycle-by-cycle, CBC) that underscores the need for standardization of conditions: temperature and duty cycles can vary over the course of a day and throughout a week, month and year; polarization can develop within an immediate cycle and throughout successive cycles as a hysteresis. If CBC data is envisioned as a data source to determine performance and aging trends, it should be recognized that polarization is a frequent consequence of CBC and thus makes it difficult to separate reversible and irreversible components to metrics such as capacity loss and resistance increase over aging. Since CBC conditions can have a major impact on data usability, we will devote part of this paper to CBC data conditioning and management. Differential analyses will also be discussed as a means to detect changing trends in data quality. Our target cell chemistries will be lithium-ion types NMC/graphite and LMO/LTO.

25 ENERGY STORAGE↗

sciCAN: single-cell chromatin accessibility and gene expression data integration via cycle-consistent adversarial network

The boom in single-cell technologies has brought a surge of high dimensional data that come from different sources and represent cellular systems from different views. With advances in these single-cell technologies, integrating single-cell data across modalities arises as a new computational challenge. Here, we present an adversarial approach, sciCAN, to integrate single-cell chromatin accessibility and gene expression data in an unsupervised manner. We benchmarked sciCAN with 5 existing methods in 5 scATAC-seq/scRNA-seq datasets, and we demonstrated that our method dealt with data integration with consistent performance across datasets and better balance of mutual transferring between modalities than the other 5 existing methods. We further applied sciCAN to 10X Multiome data and confirmed that the integrated representation preserves biological relationships within the hematopoietic hierarchy. Finally, we investigated CRISPR-perturbed single-cell K562 ATAC-seq and RNA-seq data to identify cells with related responses to different perturbations in these different modalities.

59 BASIC BIOLOGICAL SCIENCES↗

CIRSS vertical data integration, San Bernardino County study phases 1-A, 1-B

User needs, data types, data automation, and preliminary applications are described for an effort to assemble a single data base for San Bernardino County from data bases which exist at several administrative levels. Each of the data bases used was registered and converted to a grid-based data file at a resolution of 4 acres and used to create a multivariable data base for the entire study area. To this data base were added classified LANDSAT data from 1976 and 1979. The resulting data base thus integrated in a uniform format all of the separately automated data within the study area. Several possible interactions between existing geocoded data bases and LANDSAT data were tested. The use of LANDSAT to update existing data base is to be tested.

Christenson, J.↗

Lunar Instrument Data Integration Into the Virtual Reality Mission Simulation System for Decision Communication and Situational Awareness

In situ resource utilization (ISRU) technologies are a key advancement required to make human habitation on the Moon and Mars viable. The upcoming Volatiles Investigating Polar Exploration Rover (VIPER) mission will provide crucial correlations between volatiles and lunar geology to understand the water content available for ISRU on the moon. The mission will require the coordination of multi-disciplinary teams across the country making real time decisions based on rover instrument data. The virtual reality Mission Simulation System (vMSS) is a virtual reality platform designed at MIT by the Resource Exploration and Science of our Cosmic Environment (RESOURCE) team to provide teams with a collaboration interface for planetary missions like VIPER. Herein we determine the integration pathway for analog based datasets that are examples of VIPER's two main instruments, the near-infrared volatile spectrometer subsystem (NIRVSS) and the neutron spectrometer subsystem (NSS), into vMSS to provide the most valuable visualization tools. Focusing on improving situational awareness, decision making, reducing task load and incorporating comments from scientists working previous analogs and on the current VIPER mission, we recommend critical elements to implement into vMSS and the best approaches for data visualization. We present a review of relevant analogs and state of the art mission software. We have developed a design concept and path to flight of analysed instrument data integrated with data maps that allow for virtual manipulation and annotation between non-co-located team members. We focus on pre-mission mapping of a priori data for improved situational awareness, layering of analysed instrument data, correlative mapping and interactive capabilities for in-mission decision making, as well as archiving and annotation tools for post-mission analysis. Finally, we lay out the roadmap for the future development of immersive sample site visualization capabilities and the use of integrated instrument data in vMSS with automated temporal and geospatial planning.

Cody Alison Paige↗

Integrating data types to estimate spatial patterns of avian migration across the Western Hemisphere

For many avian species, spatial migration patterns remain largely undescribed, especially across hemispheric extents. Recent advancements in tracking technologies and high-resolution species distribution models (i.e., eBird Status and Trends products) provide new insights into migratory bird movements and offer a promising opportunity for integrating independent data sources to describe avian migration. Here, we present a three-stage modeling framework for estimating spatial patterns of avian migration. First, we integrate tracking and band re-encounter data to quantify migratory connectivity, defined as the relative proportions of individuals migrating between breeding and nonbreeding regions. Next, we use estimated connectivity proportions along with eBird occurrence probabilities to produce probabilistic least-cost path (LCP) indices. In a final step, we use generalized additive mixed models (GAMMs) both to evaluate the ability of LCP indices to accurately predict (i.e., as a covariate) observed locations derived from tracking and band re-encounter data sets versus pseudo-absence locations during migratory periods and to create a fully integrated (i.e., eBird occurrence, LCP, and tracking/band re-encounter data) spatial prediction index for mapping species-specific seasonal migrations. To illustrate this approach, we apply this framework to describe seasonal migrations of 12 bird species across the Western Hemisphere during pre- and postbreeding migratory periods (i.e., spring and fall, respectively). We found that including LCP indices with eBird occurrence in GAMMs generally improved the ability to accurately predict observed migratory locations compared to models with eBird occurrence alone. Using three performance metrics, the eBird + LCP model demonstrated equivalent or superior fit relative to the eBird-only model for 22 of 24 species–season GAMMs. In particular, the integrated index filled in spatial gaps for species with over-water movements and those that migrated over land where there were few eBird sightings and, thus, low predictive ability of eBird occurrence probabilities (e.g., Amazonian rainforest in South America). This methodology of combining individual-based seasonal movement data with temporally dynamic species distribution models provides a comprehensive approach to integrating multiple data types to describe broad-scale spatial patterns of animal movement. Further development and customization of this approach will continue to advance knowledge about the full annual cycle and conservation of migratory birds.

59 BASIC BIOLOGICAL SCIENCES↗