Engineering PapersSearch

SEARCH · Engineering Papers

Results for “heterogeneous data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

New Ways of Facilitating Improved Data Discovery and Access for NASA's Suborbital Earth Science Observations

NASA conducts field research in various Earth Science disciplines utilizing airborne and other non-satellite platforms to acquire in situ and remotely sensed observations indicative of physical processes across a range of scales. Field efforts are key in the development and validation of instruments and satellite algorithm refinements. The heterogeneous data, with a range of file formats, scales, and acquisition methods, support research in several science areas. NASA’s archive process assigns data products to discipline-oriented Distributed Active Archive Centers (DAACs) for stewardship. Over time, individual DAACs have developed tools for data browsing and serving disparate user bases. As science becomes more interdisciplinary, researchers need to incorporate observations from multiple campaigns, and multiple DAACs, into their work. Motivated in part by this shifting paradigm of needs, the Catalog of Archived Suborbital Earth Science Investigations (CASEI) was created. CASEI provides a single starting point to browse, search, and discover airborne and field data. Contextual metadata are organized and inter-linked allowing intuitive, integrated exploration across all NASA DAACs. Campaign science objectives, platform and instrument configurations, geographical details, geophysical concepts, and more are tracked in CASEI’s database, facilitating multi-parameter search, browse, and discovery of relevant data products. Researchers are able to directly access associated data products, via DOI links, regardless of the DAAC where they reside. Significant events, key time periods of high science interest within the longer-duration campaign effort, are also indicated and allow for a more efficient identification of critical data subsets. This presentation describes CASEI’s development, intensive metadata curation process, and demonstrates the web interface experience. Initial content metrics and plans for continued maintenance will also be discussed.

metadata

An Analysis of Earth Science Data Analytics Use Cases

The increase in the number and volume, and sources, of globally available Earth science data measurements and datasets have afforded Earth scientists and applications researchers unprecedented opportunities to study our Earth in ever more sophisticated ways. In fact, the NASA Earth Observing System Data Information System (EOSDIS) archives have doubled from 2007 to 2014, to 9.1 PB (Ramapriyan, 2009; and https:earthdata.nasa.govaboutsystem-- performance). In addition, other US agency, international programs, field experiments, ground stations, and citizen scientists provide a plethora of additional sources for studying Earth. Co--analyzing huge amounts of heterogeneous data to glean out unobvious information is a daunting task. Earth science data analytics (ESDA) is the process of examining large amounts of data of a variety of types to uncover hidden patterns, unknown correlations and other useful information. It can include Data Preparation, Data Reduction, and Data Analysis. Through work associated with the Earth Science Information Partners (ESIP) Federation, a collection of Earth science data analytics use cases have been collected and analyzed for the purpose of extracting the types of Earth science data analytics employed, and requirements for data analytics tools and techniques yet to be implemented, based on use case needs. ESIP generated use case template, ESDA use cases, use case types, and preliminary use case analysis (this is a work in progress) will be presented.

data analytics

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING

ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability

Earth system predictability is challenged by the complexity of environmental dynamics and the multitude of variables involved. Current AI foundation models, although advanced by leveraging large and heterogeneous data, are often constrained by their size and data integration, limiting their effectiveness in addressing the full range of Earth system prediction challenges. To overcome these limitations, we introduce the Oak Ridge Base Foundation Model for Earth System Predictability (ORBIT), an advanced vision transformer model that scales up to 113 billion parameters using a novel hybrid tensor-data orthogonal parallelism technique. As the largest model of its kind, ORBIT surpasses the current climate AI foundation model size by a thousandfold. Performance scaling tests conducted on the Frontier supercomputer have demonstrated that ORBIT achieves 684 petaFLOPS to 1.6 exaFLOPS sustained throughput, with scaling efficiency maintained at 41% to 85% across 49,152 AMD GPUs. These breakthroughs establish new advances in AI-driven climate modeling and demonstrate promise to significantly improve the Earth system predictability.

Wang, Xiao

Scalable Federated Learning for Scientific Foundation Models on Leadership-Class Systems

Federated learning (FL) at leadership-class HPC systems remains largely unexplored, despite growing interest in deploying federated workflows on modern HPC systems. This paper provides the first system-level empirical characterization of federated fine-tuning of pretrained foundation models on an exascale supercomputer under a multi-node deployment. Using up to 96 concurrent FL clients deployed across Frontier nodes, we study the impact of client scale, model size, data heterogeneity, partial participation, and differential privacy on runtime, communication overhead, and convergence stability. Our results show that pretrained transformer models remain robust to heterogeneity, client dropout, and privacy noise, while system efficiency degrades rapidly with scale as synchronizat and orchestration dominate runtime. We further demonstrate that system-aware execution strategies, including intra-node aggregation and early aggregation, significantly reduce wall-clock time without degrading model quality. These findings establish a practical performance baseline and inform the design of communication-efficient FL systems on leadership-class HPC platforms.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)

Real-Twin

Real-Twin is a unified, model-agnostic scenario generation tool designed to streamline and standardize the evaluation of emerging mobility technologies. It provides an end-to-end framework that includes robust workflows, integrated tools, and comprehensive metrics to generate, calibrate, and benchmark microscopic traffic simulation scenarios across multiple platforms. Key Features of Real-Twin include: - Unified Scenario Generation: generate transferable, simulation-ready scenarios from heterogeneous data sources using a consistent workflow. - Automated Calibration Workflow: bridges simulation and real-world data, minimizing manual effort and making traffic simulation more accessible to researchers and engineers. - Model-Agnostic Compatibility: supports SUMO, VISSIM, and AIMSUN for cross-platform scenario generation and benchmarking. Enables reliable comparisons and reproducibility across different simulation tools. - Consistent Scenarios across Different Simulators: generate comparable simulation scenarios across different microscopic traffic simulators, providing users the ability to conduct benchmarking and cross-validation that are crucial for ensuring the reliability and reproducibility of simulation results. - Emerging Technology Support: includes a scenario database and pipeline for studying autonomous vehicles (AVs), with planned extensions to CAVs, EVs, and other advanced technologies.

Wang, Chieh (Ross) [Oak Ridge National Laboratory

Shaping the Future of Self-Driving Autonomous Laboratories Workshop

The "Shaping the Future of Self-Driving Autonomous Laboratories" workshop, held in Denver on November 7-8, 2024, brought together leading experts from materials science and computing to address the growing need to revolutionize scientific research through AI-driven autonomous laboratories. The workshop identified critical challenges, including the integration of heterogeneous data, development of AI systems that understand fundamental physical principles, and comprehensive safety protocols. Key recommendations emerged around developing universal laboratory equipment interfaces, implementing automated metadata collection systems, and creating hybrid AI approaches that combine data-driven learning with scientific principles. The workshop emphasized maintaining human oversight while leveraging automation, transforming scientific education to prepare the next generation of researchers, and establishing a national consortium leveraging DOE facilities as anchors for broader collaboration with academia and industry. Participants stressed the urgency of addressing the growing disconnect between human decision-making timescales and modern instrumentation capabilities, highlighting the need for strategic automation while preserving essential human insight and oversight in the research process.

36 MATERIALS SCIENCE

NERVA materials development

Materials development topics include: development of analysis techniques to adjust heterogeneous data; determination of thermal conductivity for AISI 347 stainless steel and elastic moduli and Poisson's ratio for Inconel 718 and Ti 5Al-2.5Sn; embrittlement effects of 1400 psi gaseous hydrogen for alloy 718 and Ti 5Al-2.5Sn; cryogenic radiation damage of Ti 5Al-2.5Sn; and evaluation of prepreg, impregnation, and fabric materials for optimum fibrous graphite properties. Component support topics include: tensile design allowable development of Ti 5Al-2.5Sn for turbopump applications; evaluation of fatigue, fracture toughness, and stress corrosion properties of AA 7039-T63 for pressure vessel applications; development of AISI 347 sheet tensile and creep properties for nozzle applications; evaluation of orbital weld techniques for aluminum line fabrication; material selection of shield materials; development of high load friction and wear properties of hard chrome/gold plate combinations; and evaluation of weld processes for NASS duct coolant channel fabrication.

Mandell, B.

Treatment of geodetic leveling in the integrated geodesy approach

Integrated geodesy is a method in which a wide variety of surveying measurements are modeled in terms of geometric positions and the earth's geopotential. Using heterogeneous data, both geometric and gravimetric quantities are simultaneously estimated by a least-squares procedure. Heretofore, geodetic leveling differences were reduced into pseudo-observables using assumed values of gravity prior to their inclusion into integrated geodesy least-squares adjustments. The errors in estimates of geometric and gravimetric quantities obtained from integrated geodesy adjustments of geodetic leveling difference, potential differences and Helmert height differences are compared.

Milbert, Dennis G.

DNET: A communications facility for distributed heterogeneous computing

This document describes DNET, a heterogeneous data communications networking facility. DNET allows programs operating on hosts on dissimilar networks to communicate with one another without concern for computer hardware, network protocol, or operating system differences. The overall DNET network is defined as the collection of host machines/networks on which the DNET software is operating. Each underlying network is considered a DNET 'domain'. Data communications service is provided between any two processes on any two hosts on any of the networks (domains) that may be reached via DNET. DNET provides protocol transparent, reliable, streaming data transmission between hosts (restricted, initially to DECnet and TCP/IP networks). DNET also provides variable length datagram service with optional return receipts.

Tole, John

SIMBAD quality-control

The astronomical database SIMBAD developed at the Centre de donnees astronomiques de Strasbourg presently contains 760,000 objects (stellar and non-stellar). It has the unique characteristic of being structured specifically for astronomical objects. All types of heterogeneous data (bibliographic references, measurements, and sets of identification) are connected with each object. The attributes that define quality of the database include the following. Reliability: cross-identification should not rely upon just exact values object coordinates. It also means that information attached to one simple object should be consistent. The existing data must be controlled in order to start with a reliable base and to cross-identify new data assuring the quality as data grows. Exhaustivity: delays between publication of new informations and their inclusion in the database should be as short as possible. The integrity of the database has to be maintained as data accumulates. Taking the amount of data into consideration and the rate of new data production, it is necessary to use automatic methods. One of the possibilities is to use multivariate data analysis. The factor-space is a n-dimensional relevancy space which is described by the n-axes representing a set of n subject matter headings; the words and phrases can be used to scale the axes and the documents are then a vector average of the terms within them. The application reported herein is based on the NASA-STI bibliographical database. The selected data concern astronomy, astrophysics, and space radiation (102,963 references from 1975 to 1991 included 8070 keywords). The F-space is built from this bibliographical data. By comparing the F-space position obtained from the NASA-STI keywords with the F-space position obtained from the SIMBAD references, the authors will be able to show whether it is possible to retrieve information with a restricted set of words only. If the comparison is valid, this will be a way to enter bibliographic information in the SIMBAD quality control process. Furthermore, it is possible to connect the physical measurements of stars from SIMBAD to literature concerning these stars from the NASA-STI abstracts. The physical properties of stars (e.g. UBV colors) are not randomly distributed. Stars are distributed among different clusters in a physical parameter space. The authors will show that there are some relations between this classification and the literature concerning these objects clusters in a factor space. They will investigate the nature of the relationship between the SIMBAD measurements and the bibliography. These would be new relationships that are not pre-established by an astronomer. In addition, the bibliography could be neutral information that can be used in combination with the measured parameters.

Lesteven, Soizick

Automated Classification of ROSAT Sources Using Heterogeneous Multiwavelength Source Catalogs

We describe an on-line system for automated classification of X-ray sources, ClassX, and present preliminary results of classification of the three major catalogs of ROSAT sources, RASS BSC, RASS FSC, and WGACAT, into six class categories: stars, white dwarfs, X-ray binaries, galaxies, AGNs, and clusters of galaxies. ClassX is based on a machine learning technology. It represents a system of classifiers, each classifier consisting of a considerable number of oblique decision trees. These trees are built as the classifier is 'trained' to recognize various classes of objects using a training sample of sources of known object types. Each source is characterized by a preselected set of parameters, or attributes; the same set is then used as the classifier conducts classification of sources of unknown identity. The ClassX pipeline features an automatic search for X-ray source counterparts among heterogeneous data sets in on-line data archives using Virtual Observatory protocols; it retrieves from those archives all the attributes required by the selected classifier and inputs them to the classifier. The user input to ClassX is typically a file with target coordinates, optionally complemented with target IDs. The output contains the class name, attributes, and class probabilities for all classified targets. We discuss ways to characterize and assess the classifier quality and performance and present the respective validation procedures. Based on both internal and external validation, we conclude that the ClassX classifiers yield reasonable and reliable classifications for ROSAT sources and have the potential to broaden class representation significantly for rare object types.

McGlynn, Thomas

An Ecological Forecasting Agent

The project goals are: Make data analysis faster and cheaper. Increase use of NASA data by removing barriers to data access. Cope with data heterogeneity. Support code reuse and rapid application development. Support multiple applications, users. Including fire and health domains. Improve QOS. Always provide an answer. Tell user how good it is, where it come from.

Golden, Keith

Adding Hierarchical Objects to Relational Database General-Purpose XML-Based Information Managements

NETMARK is a flexible, high-throughput software system for managing, storing, and rapid searching of unstructured and semi-structured documents. NETMARK transforms such documents from their original highly complex, constantly changing, heterogeneous data formats into well-structured, common data formats in using Hypertext Markup Language (HTML) and/or Extensible Markup Language (XML). The software implements an object-relational database system that combines the best practices of the relational model utilizing Structured Query Language (SQL) with those of the object-oriented, semantic database model for creating complex data. In particular, NETMARK takes advantage of the Oracle 8i object-relational database model using physical-address data types for very efficient keyword searches of records across both context and content. NETMARK also supports multiple international standards such as WEBDAV for drag-and-drop file management and SOAP for integrated information management using Web services. The document-organization and -searching capabilities afforded by NETMARK are likely to make this software attractive for use in disciplines as diverse as science, auditing, and law enforcement.

Lin, Shu-Chun

Bridging the Last Mile with Open-Source Advancements: Empowering Communities through Fusion of Aerosol Optical Depth (AOD) Products from Multi-Satellite Sensors

Aerosol Optical Depth (AOD) is a crucial parameter for understanding atmospheric aerosol distribution and their impact on climate and air quality. With the growing number of Earth observation satellites, there is an abundance of AOD products derived from various sensors onboard both geostationary and low-orbit satellites. The availability of multiple datasets provides an opportunity to harness the strengths of each sensor and create comprehensive and accurate AOD datasets for climate and air quality studies at different temporal and spatial scales. Our NASA aerosol MEaSURES project has made significant strides in recent years by undertaking the ambitious task of developing an open-source package tailored for fusing AOD products from different sources. The package is based on OOP (Object-Oriented Programming) design and is implemented in Python modules. Generic interfaces enable easy inclusion of large and heterogeneous data. The package may be utilized to produce harmonized AOD datasets with enhanced spatial and temporal coverage. The latest version of the package is able to process and integrate the dark-target AOD data from six different sensors: AHI Himawari-8, ABI GOES-West, ABI GOES-East, MODIS AQUA, MODIS TERRA, and VIIRS SNPP. Rigorous validation and intercomparison studies have been performed to assess the accuracy and reliability of the fused AOD product against ground-based measurements and reference datasets. The open-source nature of the developed package ensures transparency, reproducibility, and community engagement. The research community and stakeholders can access, contribute to, and further improve the fusion methodology, making it adaptable to other studies, or expanding it to include new satellite data as they become available. In this poster presentation, we will introduce the accomplishments and challenges faced during the development of the open-source package for AOD data fusion, and demonstrate the advantages of combining AOD products from the six aforementioned satellite sensors. The presentation aims to foster discussions, collaborations, and future directions in integrating Earth observation and remote sensing data, which may contribute to a better understanding of atmospheric aerosols and their impacts on our environment.

Zhaohui Zhang

Archi: Agentic Operations at the CMS Experiment

We present Archi, an open-source, end-to-end framework for scientific collaborations that combines the systematic ingestion and organization of heterogeneous data sources with the deployment of configurable, private, and extensible agents that retrieve and reason over them. An instance of Archi has been deployed for the Computing Operations team of the CMS experiment at CERN's LHC since February 2026 as a support agent for technical operators, offering retrieval and analysis capabilities by combining documentation, historical data, and live monitoring systems. We evaluate the system on operator feedback and a question set collected from production usage, graded by human and automated panels. The system proves effective at operational tasks, resolving real-world queries posed by CMS operators. We also observe that locally-hosted, open-weight models perform competitively, enabling fully private management of sensitive data.

Lugato, Pietro [MIT; CERN]

Issues and Solutions for Bringing Heterogeneous Water Cycle Data Sets Together

The water cycle research community has generated many regional to global scale products using data from individual NASA missions or sensors (e.g., TRMM, AMSR-E); multiple ground- and space-based data sources (e.g., Global Precipitation Climatology Project [GPCP] products); and sophisticated data assimilation systems (e.g., Land Data Assimilation Systems [LDAS]). However, it is often difficult to access, explore, merge, analyze, and inter-compare these data in a coherent manner due to issues of data resolution, format, and structure. These difficulties were substantiated at the recent Collaborative Energy and Water Cycle Information Services (CEWIS) Workshop, where members of the NASA Energy and Water cycle Study (NEWS) community gave presentations, provided feedback, and developed scenarios which illustrated the difficulties and techniques for bringing together heterogeneous datasets. This presentation reports on the findings of the workshop, thus defining the problems and challenges of multi-dataset research. In addition, the CEWIS prototype shown at the workshop will be presented to illustrate new technologies that can mitigate data access roadblocks encountered in multi-dataset research, including: (1) Quick and easy search and access of selected NEWS data sets. (2) Multi-parameter data subsetting, manipulation, analysis, and display tools. (3) Access to input and derived water cycle data (data lineage). It is hoped that this presentation will encourage community discussion and feedback on heterogeneous data analysis scenarios, issues, and remedies.

Acker, James

A Comparison of FIFE Observation with GEOS Assimilated Data Including a Heterogeneous LSM

Several recent studies have shown that much can be learned by comparing grid-point data from a data assimilation system with in-situ observations from field experiments. While the surface heterogeneity is acknowledged in these studies, they lack quantitative representations of the influence of heterogeneity on the near-surface meteorology and surface hydrologic and energy balance. Here, we use the Betts and Ball FIFE site-averaged data. Standard deviations of the site-average will provide an estimate of the FIFE site heterogeneity. Recently, the Mosaic Land-Surface Model (LSM) has been incorporated into the Goddard Earth Observing System (GEOS) Data Assimilation System (DAS). The Mosaic LSM computes the surface energy and hydrologic balance for nine distinct surface types at each grid-point. Each surface type is proportionally weighted to determine the mean grid point properties. Hence, we can compare modeled and observed grid-point variability in addition to the mean properties. Also, assimilated data sets created with and without the LSM are compared. The results indicate the importance of including quantitative estimates of heterogeneity in the analysis of the land surface hydrology and energy balances in assimilation systems.

Bosilovich, M.