Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

CODARcode/MGARD

MGARD is a software providing error-controlled lossy compression and data refactoring based on multi-grid theories. It transforms floating-point scientific data into a multilevel representation, followed by quantization and lossless encoding processes, resulting in a self-describing compressed buffer. It supports diverse data topologies, error control norms, and computing architectures.

Chen, Jieyang [University of Oregon]↗

Efficient Data Compression for 3D Sparse TPC via Bicephalous Convolutional Autoencoder

Real-time data collection and analysis in large experimental facilities present a great challenge across multiple domains, including high energy physics, nuclear physics, and cosmology. To address this, machine learning (ML)-based methods for real-time data compression have drawn significant attention. However, unlike natural image data, such as CIFAR and ImageNet that are relatively small-sized and continuous, scientific data often come in as three-dimensional 3D data volumes at high rates with high sparsity (many zeros) and non-Gaussian value distribution. This makes direct application of popular ML compression methods, as well as conventional data compression methods, suboptimal. To address these obstacles, this work introduces a dual-head autoencoder to resolve sparsity and regression simultaneously, called Bicephalous Convolutional AutoEncoder (BCAE). This method shows advantages both in compression fidelity and ratio compared to traditional data compression methods, such as MGARD, SZ, and ZFP. To achieve similar fidelity, the best performer among the traditional methods can reach only half the compression ratio of BCAE. Moreover, a thorough ablation study of the BCAE method shows that a dedicated segmentation decoder improves the reconstruction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Efficient and Flexible Hierarchical Data Layouts for a Unified Encoding of Scalar Field Precision and Resolution

To address the problem of ever-growing scientific data sizes making data movement a major hindrance to analysis, we introduce a novel encoding for scalar fields: a unified tree of resolution and precision, specifically constructed so that valid cuts correspond to sensible approximations of the original field in the precision-resolution space. Furthermore, we introduce a highly flexible encoding of such trees that forms a parameterized family of data hierarchies. We discuss how different parameter choices lead to different trade-offs in practice, and show how specific choices result in known data representation schemes such as zfp[52], idx[58], and jpeg2000 [76]. Lastly, we provide system-level details and empirical evidence on how such hierarchies facilitate common approximate queries with minimal data movement and time, using real-world data sets ranging from a few gigabytes to nearly a terabyte in size. Experiments suggest that our new strategy of combining reductions in resolution and precision is competitive with state-of-the-art compression techniques with respect to data quality, while being significantly more flexible and orders of magnitude faster, and requiring significantly reduced resources.

97 MATHEMATICS AND COMPUTING↗

Assessing data change in scientific datasets

Summary Scientific datasets are growing rapidly and becoming critical to next‐generation scientific discoveries. The validity of scientific results relies on the quality of data used and data are often subject to change, for example, due to observation additions, quality assessments, or processing software updates. The effects of data change are not well understood and difficult to predict. Datasets are often repeatedly updated and recomputing derived data products quickly becomes time consuming and resource intensive and may in some cases not even be necessary, thus delaying scientific advance. Despite its importance, there is a lack of systematic approaches for best comparing data versions to quantify the changes, and ad‐hoc or manual processes are commonly used. In this article, we propose a novel hierarchical approach for analyzing data changes, including real‐time (online) and offline analyses. We employ a variety of fast‐to‐compute numerical analyses, graphical data change representations, and more resource‐intensive recomputations of a subset of the data product. We illustrate the application of our approach using three scientific diverse use cases, namely, satellite, cosmological, and x‐ray data. The results show that a variety of data change metrics should be employed to enable a comprehensive representation and qualitative evaluation of data changes.

97 MATHEMATICS AND COMPUTING↗

Reply to Comment on ‘The advanced tokamak path to a compact net electric fusion pilot plant’

Abstract The Comment by Manheimer (submitted to Nucl. Fusion with this response) on our recent paper Buttery et al (2021 Nucl. Fusion 61 046028), has mischaracterized our paper with a series of misleading statements about its content, which it then seeks to refute. It also offers a series of assertions that are unsupported by refereed publication or scientific data—and sometimes in contradiction to the published record and well-established points in the scientific community. In this response we address the main thrusts and themes of the comment, while providing a point-by-point response in the appendix.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

ATLAS HL-LHC Demonstrators with Data Carousel: Dataon-Demand and Tape Smart Writing

The High Luminosity upgrade to the LHC (HL-LHC) is expected to deliver scientific data at the multi-exabyte scale. To tackle this unprecedented data storage challenge, the ATLAS experiment initiated the Data Carousel project in 2018. Data Carousel is a tape-driven workflow in which bulk production campaigns with input data resident on tape are executed by staging and promptly processing a sliding window to disk buffer such that only a small fraction of inputs are pinned on disk at any one time. Put in ATLAS production before Run3, Data Carousel continues to be our focus for seeking new opportunities in disk space savings, and enhancing tape usage throughout the ATLAS Distributed Computing (ADC) environment. These efforts are highlighted by two recent ATLAS HL-LHC demonstrator projects: data-on-demand and tape smart writing. In this paper, we will discuss the recent studies and outcomes from these projects. The research was conducted together with site experts at CERN and Tier-1 centers.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

5G Enabled Energy Innovation: Advanced Wireless Networks for Science (Workshop Report)

Rapidly expanding, new telecommunications infrastructure based on 5G technologies will disrupt and transform how we design, build, operate, and optimize scientific infrastructure and the experiments and services enabled by that infrastructure, from continental-scale sensor networks to centralized scientific user facilities, from intelligent Internet of Things devices to supercomputers. Concurrently, 5G will introduce, or exacerbate, challenges related to protecting infrastructure and associated scientific data as well as to fully leveraging opportunities related to expanded infrastructure scale and complexity. The U.S. Department of Energy (DOE) Office of Science operates scientific infrastructure, supporting some of the nation’s most advanced intellectual discoveries, spanning the country and including 30 world-class user facilities from supercomputers to accelerators. Along with field experiments and remote observatories, every aspect of DOE’s scientific enterprise will be affected by 5G, which amounts to a complete renovation of the underpinnings of the nation’s information infrastructure. In this report we explore the scientific opportunities and new research challenges associated with 5G, ranging from scalability to heterogeneity to cybersecurity. The rapid commercial deployment of 5G opens the opportunity to rethink and reinvent DOE’s scientific infrastructure and experimentation, from intelligent sensor networks at unprecedented scales to a digital continuum of cyberinfrastructure spanning low-power sensors, high-performance computing embedded within and at the edge of the network, and DOE’s large-scale user instrument and computing facilities. New programming paradigms, workflow and data frameworks, and AI-based system design, operation, and autonomous adaptation and optimization will be necessary in order to exploit these new opportunities. Field deployments and centralized scientific instruments can also be revolutionized, moving (without traditional performance penalties) from wired to wireless connectivity for data and control systems, improving flexibility, and opening new sensing modalities, including the use of the 5G electromagnetic spectrum itself as an environmental probe. For DOE science, in contrast to commercial 5G applications and settings, devices will be deployed in extreme environments such as cryogenically cooled instrument control systems and in remote settings with harsh conditions, requiring the design of new materials for RF communication and edge processing to operate in these regimes. Concurrently, 5G infrastructure comprises both hardware and sophisticated software systems - currently closed and proprietary. The cybersecurity challenges to 5G-empowered reinvention mirror the complexity and variety of new 5G features, from virtualization to private network slices to ubiquitous access. Research is also needed in order to accelerate the development of secure and open 5G software infrastructure, reducing reliance on hardware and software produced outside the United States and providing the transparency and rigorous evaluation and testing afforded through open software. Twelve broad research thrusts are laid out in four chapters, with a companion fifth chapter (and three additional research thrusts) underscoring the needs and opportunities for an aggressive testbed program co-designed by networking experts and scientists involved in the 15 research thrusts. The urgency of undertaking this research is fueled by a global, accelerating deployment of new telecommunications infrastructure that is designed for entertainment and commercial applications - barely scratching the surface of what 5G can do to extend U.S. leadership in scientific discovery.

42 ENGINEERING↗

FasTensor (FT) v0.0.1

The FasTensor implements the native array data programming mode namely SLOPE for modern data analysis tasks. It is designed to work in parallel on supercomputer and can handle terabyte data analysis tasks. It is thousand times faster than other data analysis systems such as Apache Spark and 38% faster than TensorFlow in dealing with large scale scientific data analysis. It has wide applications in image data analysis and mesh data analysis tasks in physical simulations and other fields.

Wu, Kesheng↗

Transitioning from File-Based HPC Workflows to Streaming Data Pipelines with openPMD and ADIOS2

This paper aims to create a transition path from file-based IO to streaming-based workflows for scientific applications in an HPC environment. By using the openPMP-api, traditional workflows limited by filesystem bottlenecks can be overcome and flexibly extended for in situ analysis. The openPMD-api is a library for the description of scientific data according to the Open Standard for Particle-Mesh Data (openPMD). Its approach towards recent challenges posed by hardware heterogeneity lies in the decoupling of data description in domain sciences, such as plasma physics simulations, from concrete implementations in hardware and IO. The streaming backend is provided by the ADIOS2 framework, developed at Oak Ridge National Laboratory. This paper surveys two openPMD-based loosely-coupled setups to demonstrate flexible applicability and to evaluate performance. In loose coupling, as opposed to tight coupling, two (or more) applications are executed separately, e.g. in individual MPI contexts, yet cooperate by exchanging data. This way, a streaming-based workflow allows for standalone codes instead of tightly-coupled plugins, using a unified streaming-aware API and leveraging high-speed communication infrastructure available in modern compute clusters for massive data exchange. We determine new challenges in resource allocation and in the need of strategies for a flexible data distribution, demonstrating their influence on efficiency and scaling on the Summit compute system. The presented setups show the potential for a more flexible use of compute resources brought by streaming IO as well as the ability to increase throughput by avoiding filesystem bottlenecks.

Poeschel, Franz↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Multi-resolution enhancement for full-spectrum neural representations

Scientific data acquisition continues to outpace storage and analysis capabilities, making voxel-basedrepresentations increasingly intractable. Implicit neural representations (INRs) offer a promising solutionby encoding signals through coordinate-based neural networks, serving as surrogates of data, withcomputational and storage requirements scaling with network complexity rather than data dimensionality.However, smaller INRs struggle to faithfully represent multiscale structures, high-frequency informationand fine textures that constitute a large proportion of scientific measurements. We propose WIEN-INR, atheoretically guided hierarchical INR framework that distributes modelling across resolution scales andenables improved representation capacity through a novel enhancement network to recover subtle details.This multiscale architecture allows smaller networks to retain the full spatial-frequency content of thesignal as well as preserve training efficiency and lower storage cost. Evaluated on distinct raw experimentalmeasurements across scales and complexities, WIEN-INR represents a practical step towards a broaderadoption of neural representations in scientific workflows, delivering compact, robust and high-fidelityrepresentations.

Ni, Yuan [SLAC National Accelerator Laboratory (SL↗

Deceptive Infusion of Data: A Novel Data Masking Paradigm for High-Valued Systems

This work addresses how analysts of a high-valued system (e.g., nuclear reactor, aircraft turbine designs) can extract findable, accessible, interoperable, and reusable scientific data for public dissemination to artificial intelligence and machine-learning (AI/ML) researchers in a manner that cannot be reverse-engineered, potentially compromising sensitive or proprietary information. State-of-the-art methods address this problem through data masking techniques, which allow access to a subset of the information while obfuscating private and potentially identifying information (e.g., personally identifying medical data). These methods are unsuitable for industrial engineering processes, where AI/ML tools need explicit access to all the data available to draw the best inference about the system to help optimize its performance and identify its vulnerabilities, etc. Our novel deceptive infusion of data paradigm provides a solution to this conundrum by developing a mathematical approach capable of concealing the identity of the system while providing full access to all the features employed by AI/ML tools to ensure their optimal performance.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Summary of Responses to the Request for Information (RFI) on Partnerships for Transformational Artificial Intelligence Models

The Department of Energy (DOE) issued a Request for Information (RFI) in December 2025 inviting public comments regarding partnerships for transformational Artificial Intelligence (AI) models for the Genesis Mission Consortium, a public-private partnership platform. This RFI solicited feedback from industry, nonprofit organizations, universities, independent research organizations and other stakeholders. Specifically, the RFI asked three questions on (1) mobilizing DOE National Laboratories to curate the scientific data in a responsible and privacy-preserving manner, (2) the extent to which existing general-purpose AI models can be leveraged and which scientific disciplines are priorities for such model development, and (3) mechanisms by which these AI models can be provided to scientific communities. This document summarizes the input from 194 unique nonproprietary responses from businesses, universities, nonprofit organizations, research institutes and laboratories as well as a variety of other contributors, including individual contributions.

97 MATHEMATICS AND COMPUTING↗

Predicting Execution Times for Disk-based and In-Situ Parallel Data Analytics (Final Technical Report)

In recent years, there has been a significant amount of interests in in-situ analytics on simulation programs. For a variety of reasons, it is desirable to be able to predict the execution time of an analytics program. At the same time, frameworks such as MapReduce have become popular for scientific data analytics. This paper focuses on developing performance models for predicting execution time of parallel data analytics, with a special emphasis on in-situ analytics. We take two distinct approach towards performance prediction. We first expand SKOPE (a SKeleton framewOrk for Performance Exploration) with performance models for disk data read, cache performance, and page fault penalty. Second, an analytical performance model is also developed. We have evaluated our performance prediction framework as well as the analytical model on three hardware setups with well-known data mining algorithms implemented in three programming paradigms, MapReduce, MATE (a MapReduce-like parallel system with an alternate API for multi-core environments) and Smart (a MapReduce-like framework for in-situ analytics). Results show that our performance prediction framework along with the incorporated performance models are capable of accurately predicting execution times for parallel scientific analytics on different hardware setups.

97 MATHEMATICS AND COMPUTING↗

A Strategy for NACS investment in Machine Learning

The Nuclear and Chemical Sciences (NACS) Division furnishes the expertise in the scientific areas of chemical, nuclear and isotopic sciences that are foundational in the Laboratory’s national security missions. This expertise is maintained and advanced through identification, development and application of state-of-the-art theoretical, computational and experimental methods and tools. Recent developments in artificial intelligence and machine learning (AI/ML) techniques enabled by advances in computing capabilities and widespread availability of powerful software implementations have made use of these techniques ubiquitous across both science and industry. While the scope of AI/ML applications is incredibly large and evolves very rapidly, the topics most relevant to NACS missions fall into the general category of detecting, categorizing or identifying features in large, complex datasets using either supervised or unsupervised learning. This covers both basic scientific data analysis and the development of efficient surrogate models of real-life technological systems, experimental detectors, or theoretical models. To remain at the forefront of its core scientific disciplines, NACS must both cultivate ML expertise as well as continuously explore applying this expertise to new problems or utilizing new methods. This document identifies the key areas where this support is critical and provides a strategy for investing in them.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

A roadmap toward scaling, reasoning and self-evolving foundation models for nuclear and particle physics

Foundation models have revolutionized artificial intelligence, with Large Language Models demonstrating unprecedented capabilities in multimodal understanding, reasoning and tool use. Nuclear and particle physics stands at a critical juncture where similar transformative potential awaits realization. The field generates exabytes of experimental data, exascale simulations, and decades of theoretical insights — yet these remain largely disconnected from modern Artifical Intelligence (AI) capabilities, with most physics AI applications confined to narrow, task-specific models that suffer from domain shifting when applied to real experimental data. We present a roadmap for FM4NPP (Foundation Model for Nuclear and Particle Physics), systematically scaling from current proof-of-concept models to trillion-parameter architectures capable of autonomous discovery. Our approach advances three critical frontiers: unified data infrastructure integrating detector data, scientific knowledge and computational tools across global facilities; multi-facility foundation models enabling cross-experiment knowledge transfer and accelerated discovery; and agentic AI capabilities for reasoning and autonomous tool use. The resulting self-evolving FM4NPP will transform physics research by converting time-intensive data analysis, theory derivation and computational bottlenecks into rapid AI–human collaborative discovery. This paradigm shift promises to fundamentally accelerate scientific progress in nuclear and particle physics, enabling researchers to focus on high-level insights while AI handles routine analysis and explores vast parameter spaces beyond human capacity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Accelerated, scalable and reproducible AI-driven gravitational wave detection

The development of reusable artificial intelligence (AI) models for wider use and rigorous validation by the community promises to unlock new opportunities in multi-messenger astrophysics. Here we develop a workflow that connects the Data and Learning Hub for Science, a repository for publishing AI models, with the Hardware-Accelerated Learning (HAL) cluster, using funcX as a universal distributed computing service. Using this workflow, an ensemble of four openly available AI models can be run on HAL to process an entire month's worth (August 2017) of advanced Laser Interferometer Gravitational-Wave Observatory data in just seven minutes, identifying all four binary black hole mergers previously identified in this dataset and reporting no misclassifications. This approach combines advances in AI, distributed computing and scientific data infrastructure to open new pathways to conduct reproducible, accelerated, data-driven discovery. By combining a repository for artificial intelligence models and a supercomputing cluster, an entire month's worth of advanced LIGO data is analysed in just 7 min, finding all binary black hole mergers previously identified in this dataset and reporting no misclassifications.

79 ASTRONOMY AND ASTROPHYSICS↗