Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “heterogeneous data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

LANDSAT-4 image data quality analysis

Seven heterogeneous areas within the Des Moines, Iowa area test site were selected to define candidate spectral training classes using a clustering algorithm. In addition to the 91 cluster (nonsupervised) classes, three supervised training classes were defined and subsequently included in the training statistics file. The identity of all 94 candidate classes were determined using available reference data. Through analysis of the interclass separabilities, the original 94 candidate training classes were reduced to 42 spectrally separable final classes. The minimum average transformed divergence values for the 42 spectral classes and for the best subsets of TM spectral bands are shown in a table.

Anuta, P. E.↗

Capturing Historic Reliability Performance Through Graph Databases: A Model Based System Engineering Approach

With the goal of improving the performance and reliability of high dependable technological systems such as nuclear power plants, advanced monitoring and health management systems are employed to inform system engineers on observed degradation processes and anomalous behaviors of assets and components. This information is captured in the form of large amount of data which can be heterogenous in nature (e.g., numeric, textual). Such large data availability poses challenges when system engineers are required to parse and analyze them in order to track historic reliability performance of assets and components. This paper tackles directly this challenge by providing means to organize data in the form of a graph: a knowledge graph. The presented approach distinguish itself from current knowledge graph-based methods by the fact that model-based system engineering (MBSE) models are used to “put data into context”. In particular, MBSE models are used as skeleton of a knowledge graph; numeric and textual data elements, once processed, are associated to MBSE model elements. Thus, a knowledge graph captures both system architecture (though MBSE models) and health/performance data. Such feature opens the door to new data analytics methods designed to identify causal relations between observed phenomena.

97 - MATHEMATICS AND COMPUTING↗

Enabling Interoperability in Earth System Digital Twins (ESDT): Integrating Observations, Models, and AI for Actionable Insights Through NASA'S Intelligent Systems Technology Program

NASA’s Intelligent Systems Technology Program (IST) is driving a paradigm shift in Earth science through the development of Earth System Digital Twins (ESDT). These integrated information systems create a dynamic "digital replica" of the Earth by harmonizing continuous, multi-source observations with high-fidelity models and state-of-the-art artificial intelligence (AI) that enable “What now?”, “What next?”, and “What if?” scenario building. These scenarios are reflected in NASA IST’s series of ESDTs, from the Coastal Zone Digital Twin that integrates complex data on the current state of the Chesapeake Bay to the Terrestrial Environmental Rapid-Replication and Assimilation Hydrometeorological (TerraHydro) AI-based ESDT that forecasts water movement across Earth’s surface, to the Agriculture Land Information System (AgLIS) which can be used to assess optimal planting dates and crop yield estimates. By bridging the gap between vast data archives and actionable insights, these projects enable a system-of-systems approach to understanding complex, interacting Earth processes. This poster will highlight recent innovations and future directions from NASA’s ESDT initiatives: Continuous Data Assimilation & Multi-Source Fusion. A core requirement of the ESDT work is the transition from static models to dynamic "living" replicas. This involves creating frameworks for the continual assimilation of near-real-time data from uncoordinated, heterogeneous sources, including satellite observations and airborne assets, and ground-based Internet of Things (IoT) sensors. These systems link design, operational status, and environmental data, ensuring the digital twin accurately reflects the current state of the physical Earth system. High-Fidelity Hybrid Modeling & Computational Acceleration to enable interactive "what-if" explorations, programs are moving beyond traditional, slow physical solvers by developing fast surrogate machine learning models and Deep Generative Models (DGMs). These hybrid approaches use neural networks to emulate complex physics, such as cloud feedback or ocean dynamics, at a fraction of the original computing cost, often leveraging advanced hardware like Graphics Processing Units (GPUs) to achieve the necessary scale. Federated Ecosystems & Interoperable Frameworks rather than building isolated tools, NASA IST is moving toward federated ESDTs and reusable analytic collaborative frameworks. This theme focuses on interoperability standards and common ontologies that allow specialized digital twins to interact and share data. This system-of-systems architecture supports multi-discipline investigations, such as analyzing how upstream watershed changes impact downstream urban flooding or how wildfire emissions affect regional air quality. By leveraging these advancements, ESDTs empower researchers and decision-makers to conduct real-time analysis and run complex hypothetical scenarios, ultimately improving our understanding of Earth’s evolving systems and informing critical real-world applications.

Earth System↗

Earth Science Data Analytics: Bridging Tools and Techniques with the Co-Analysis of Large, Heterogeneous Datasets

The continuum of ever-evolving data management systems affords great opportunities to the enhancement of knowledge and facilitation of science research. To take advantage of these opportunities, it is essential to understand and develop methods that enable data relationships to be examined and the information to be manipulated. This presentation describes the efforts of the Earth Science Information Partners (ESIP) Federation Earth Science Data Analytics (ESDA) Cluster to understand, define, and facilitate the implementation of ESDA to advance science research. As a result of the void of Earth science data analytics publication material, the cluster has defined ESDA along with 10 goals to set the framework for a common understanding of tools and techniques that are available and still needed to support ESDA.

science data analysis↗

Illustrating the Spatiotemporal Complexity of No2 Columns Using A Multi-Perspective Observing System: Moving Toward Geostationary Product Validation and Applications

As a precursor to secondary pollutants like ozone and PM2.5, nitrogen dioxide (NO2) is crucial to understand when addressing air quality issues. However, due to NO2’s short lifetime during the daytime and complexity of emission sources in urbanized regions, interpreting datasets from ground or satellite perspectives alone are challenged by variance in spatial and temporal resolutions. High resolution airborne mapping (< 1 km) of NO2 column densities across morning, midday, and afternoon add a unique perspective toward interpreting satellite data with respect to ground-measurements. This presentation focuses on the interpretation of spatiotemporal complexity of NO2 columns from the Synergistic TEMPO Air Quality Science Study (STAQS). The mission’s goal is to integrate geostationary observations from Tropospheric Emissions: Monitoring of Pollution (TEMPO) with traditional and enhanced air quality monitoring to improve the understanding of air quality science for increased societal benefit. We will demonstrate the interweaved perspective of NO2 columns from ground-based Pandora spectrometers and satellite-based observations (e.g., TROPOMI) as compared to high spatial resolution airborne observations from the GEOstationary Coastal and Air Pollution Events (GEO-CAPE) Airborne Simulator (GCAS). This includes the evaluation of each dataset through comparison to each other to identify potential biases in data products and the impact of heterogeneity on these comparisons. Airborne data will also be used as a proxy for geostationary observations with morning, midday, and afternoon raster maps collected over four cities (Los Angeles, Chicago, Toronto, and New York City). Finally, recent research outcomes will be presented to demonstrate how airborne and geostationary observations can be used to evaluate emission inventories and air quality models.

Laura Judd↗

A-Train Data Search and Visualization to Facilitate Multi-Instrument Cloud Studies

Now that the A-Train suite of datasets have become more mature, new and innovative science utilizing the various products has become more reliable and challenging. To perform multi-satellite research with A-Train data originating from heterogenous missions, scientists must access, subset visualize and analyze user specified datasets in ways unique to the dataset. Then hte datasets need to be co-registered and maybe merged. The A-Train Data Depot (ATDD) has been developed to save each scientist the effort and expense of developing these functions individually.

Kempler, Steven↗

Data Analysis Challenges for Multi-Messenger Astrophysics

Recent multi-messenger observations of gravitational-wave and high-energy neutrino sources together with electromagnetic signatures have opened new ways of observing the Universe. These promise a future in which physics and astronomy will be advanced by combining observations and data from across the electromagnetic spectrum with gravitational waves and neutrinos. We consider the challenges the field is facing in fully utilizing data for multi-messenger astrophysics. Such data come from heterogeneous detector networks and standards, and their analysis is often time-critical to guide further observations. In this area, science capabilities depend on the interplay among observation, theory and computational/modeling work. Advances in data science and computing present additional opportunities and considerations in analyzing such data. We invited ADASS participants to a Birds of a Feather session to engage in discussion on the challenges and opportunities in data analysis for multimessenger astrophysics.

Peter S Shawhan↗

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory↗

LANDSAT-4 image data quality analysis

Seven heterogeneous areas within the entire Des Moines, Iowa test site were selected to define candidate spectral training classes using a clustering algorithm. In addition to the 91 cluster nonsupervised classes, three supervised training classes were defined. The original candidate training classes were reduced to 42 spectrally separable training classes. The minimum and average transformed divergence values for the 42 spectral classes and for the best subsets of Y TM spectral bands are shown in a table. The best spectral band for any combination of 1 through 7 bands is the first middle IR band. The next best band is the near IR, followed by the red band and than the thermal IR. The best combination of four bands includes one from each of the four regions of the spectrum (visible, near IR, middle IR, and thermal IR).

Anuta, P. E.↗

Lunar impact basins and crustal heterogeneity - New western limb and far side data from Galileo

Multispectral images of the lunar western limb and far side obtained from Galileo reveal the compositional nature of several prominent lunar features and provide new information on lunar evolution. The data reveal that the ejecta from the Orientale impact basin (900 kilometers in diameter) lying outside the Cordillera Mountains was excavated from the crust, not the mantle, and covers pre-Orientale terrain that consisted of both highland materials and relatively large expanses of ancient mare basalts. The inside of the far side South Pole-Aitken basin (greater than 2000 kilometers in diameter) has low albedo, red color, and a relatively high abundance of iron- and magnesium-rich materials. These features suggest that the impact may have penetrated into the deep crust or lunar mantle or that the basin contains ancient mare basalts that were later covered by highlands ejecta.

Belton, Michael J. S.↗

Evolution of DUNE’s Production System

The DUNE experiment will start running in 2029 and record 30 PB/year of raw waveforms from Liquid Argon TPCs and photon detectors. The size of individual readouts can range from 100 MB to a typical 8 GB full readout of the detector, and even 100 TB for extended readouts from supernova candidates. These data then need to be cataloged, stored and distributed for processing worldwide. This massive amount of data and a heterogeneous computing environment necessitates a powerful and robust distributed computing infrastructure. In the process of building up that infrastructure, DUNE’s production system has recently undergone an overhaul, in which it has integrated 1) a new workflow management system (justIN) 2) a new data catalog (MetaCat) and 3) a state-of-the-art data management system (Rucio). Simulations of DUNE’s Far Detector and its prototypes ProtoDUNE Horizontal Drift (ProtoDUNE-HD) and ProtoDUNE Vertical Drift (ProtoDUNE-VD), as well as data from ProtoDUNE-HD serve as the first tests of this infrastructure.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ozone destruction through heterogeneous chemistry following the eruption of El Chichon

The results of ozone observations at northern midlatitudes in late 1982 through 1983, following the eruption of El Chichon are discussed, together with the observations of other trace gases which may be linked to possible variations in ozone chemistry. These results are related to the in situ aerosol observations following the El Chicon eruption, with particular attention given to data relevant to heterogeneous reactions, such as the aerosol surface area and weight percent H2SO4. It is shown that, at midlatitudes, the observed volcanic-particle surface area reached a maximum of about 50 sq microns/cu m (above a typical background value of about 0.75) at an altitude of 18-20 km in early 1983; this enhancement of surface area is about the same as that encountered in stratospheric clouds in the Antarctic, suggesting a possible basis for ozone depletion through heterogeneous chemistry. The fraction of ozone reduction that may have occurred as a result of heterogeneous chemicl effects is estimated.

Hofmann, David J.↗

GeneLab Analysis Working Group Kick-Off Meeting

Goals to achieve for GeneLab AWG - GL vision - Review of GeneLab AWG charter Timeline and milestones for 2018 Logistics - Monthly Meeting - Workshop - Internship - ASGSR Introduction of team leads and goals of each group Introduction of all members Q/A Three-tier Client Strategy to Democratize Data Physiological changes, pathway enrichment, differential expression, normalization, processing metadata, reproducibility, Data federation/integration with heterogeneous bioinformatics external databases The GLDS currently serves over 100 omics investigations to the biomedical community via open access. In order to expand the scope of metadata record searches via the GLDS, we designed a metadata warehouse that collects and updates metadata records from external systems housing similar data. To demonstrate the capabilities of federated search and retrieval of these data, we imported metadata records from three open-access data systems into the GLDS metadata warehouse: NCBI's Gene Expression Omnibus (GEO), EBI's PRoteomics IDEntifications (PRIDE) repository, and the Metagenomics Analysis server (MG-RAST). Each of these systems defines metadata for omics data sets differently. One solution to bridge such differences is to employ a common object model (COM) to which each systems' representation of metadata can be mapped. Warehoused metadata records are then transformed at ETL to this single, common representation. Queries generated via the GLDS are then executed against the warehouse, and matching records are shown in the COM representation (Fig. 1). While this approach is relatively straightforward to implement, the volume of the data in the omics domain presents challenges in dealing with latency and currency of records. Furthermore, the lack of a coordinated has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

GeneLab↗

Multi-omics data resource: Data package 23 (Pck023)

The data package consists of isolated pancreatic islets from adult male C57BL6/J mice treated with IL-1β + IFNγ, IL-1β + IFNγ + NMMA, or NMMA alone for 18 h and submitted for scRNA-seq. This study focused on the cell-type-specific effects of nitric oxide signaling in islets and characterized the heterogeneity of responses. Data contributors: Jennifer S Stancill & John A Corbett: Department of Biochemistry, Medical College of Wisconsin, Milwaukee, WI, USA Data repository: GSE183010 Publication: 10.1093/function/zqab063

Sarkar, Soumyadeep [Pacific Northwest National Lab↗

The Aggregate Description of Semi-Arid Vegetation with Precipitation-Generated Soil Moisture Heterogeneity

Meteorological measurements in the Walnut Gulch catchment in Arizona were used to synthesize a distributed, hourly-average time series of data across a 26.9 by 12.5 km area with a grid resolution of 480 m for a continuous 18-month period which included two seasons of monsoonal rainfall. Coupled surface-atmosphere model runs established the acceptability (for modelling purposes) of assuming uniformity in all meteorological variables other than rainfall. Rainfall was interpolated onto the grid from an array of 82 recording rain gauges. These meteorological data were used as forcing variables for an equivalent array of stand-alone Biosphere-Atmosphere Transfer Scheme (BATS) models to describe the evolution of soil moisture and surface energy fluxes in response to the prevalent, heterogeneous pattern of convective precipitation. The calculated area-average behaviour was compared with that given by a single aggregate BATS simulation forced with area-average meteorological data. Heterogeneous rainfall gives rise to significant but partly compensating differences in the transpiration and the intercepted rainfall components of total evaporation during rain storms. However, the calculated area-average surface energy fluxes given by the two simulations in rain-free conditions with strong heterogeneity in soil moisture were always close to identical, a result which is independent of whether default or site-specific vegetation and soil parameters were used. Because the spatial variability in soil moisture throughout the catchment has the same order of magnitude as the amount of rain failing in a typical convective storm (commonly 10% of the vegetation's root zone saturation) in a semi-arid environment, non-linearitv in the relationship between transpiration and the soil moisture available to the vegetation has limited influence on area-average surface fluxes.

White, Cary B.↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge and support human space missions. Through artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in space biosciences and engineered astronaut health systems, to enable Earth-independence and mission operations autonomy. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated mission biomonitoring, and 8) a Precision Space Health system. AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the space biology field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics to phenotypic data using an ensemble model to infer causality of rodent liver health disruption, 2) usage of explainable ML to interrogate muscular underpinnings of muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interactions, and 5) a suite of benchmarked open science datasets enabling programmers to identify best algorithms to answer space biology questions.

space biology↗

Biological Research and Space Health Enabled by Machine Learning to Support Deep Space Missions

A key science goal of the NASA “Moon to Mars” campaign is to understand how biology responds to the Lunar, Martian, and deep space environments in order to advance fundamental knowledge, reduce risk, and support safe, productive human space missions. Through the powerful emerging computer science approaches of artificial intelligence (AI) and machine learning (ML), a paradigm shift has begun in biomedical science and engineered astronaut health systems, to enable Earth-independence and autonomy of mission operations. We present a decadal view of AI/ML architecture to support deep space mission goals, developed in concert with leaders in the field. We describe current AI/ML methods to support 1) fundamental biology, 2) in situ analytics, 3) high performance computing hardware, 4) automated science, 5) self-driving labs, 6) remote data management, 7) integrated real-time mission biomonitoring, and 8) a Precision Space Health system. Cutting-edge AI/ML approaches that can be integrated to support these domains include active learning, explainable AI, adaptive learning, causal inference, knowledge graphs, federated learning, transfer learning, and large language models. Finally, we present results from several current ML projects that are underway in the field to address key challenges of small sample n, high feature count, heterogeneity, and sparse data. These include 1) connecting omics data to phenotypic data using an ensemble model to infer causality of spaceflight rodent liver health disruption, 2) usage of explainable ML to interrogate the muscular underpinnings of spaceflight muscle atrophy, 3) ML models analyzing and determining directed acyclic graphs of human space health risk leveraging rodent bone datasets, 4) usage of large pre-trained models connecting biomedical knowledgebases with small spaceflight datasets to understand gene-to-gene interaction networks, and 5) a suite of benchmarked open science datasets (spaceflight mouse liver; radiation DNA damage) enabling programmers to identify the best ML algorithms to answer space biological science questions.

space biology↗