Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data-intensive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Genesis Mission Data cards

As data-intensive research and artificial intelligence become central to DOE mission science, the need for machine-actionable dataset documentation has grown accordingly. However, many DOE-aligned communities, including the Office of Science, NNSA, and cross-laboratory collaborations, have developed independent metadata practices. This fragmentation creates friction for discovery, federation, and reuse across programs. To address these challenges, this talk introduces the Genesis Data Card: a shared metadata artifact developed in collaboration with a broad DOE community (Jefferson Lab and the National Lab of the Rockies, Oak Ridge, Sandia, Idaho, Berkeley, and Los Alamos). The Genesis Data Card aims to standardize dataset documentation across DOE-aligned initiatives while remaining extensible to discipline-specific needs. This talk will describe the data card template and the supporting code to validate completed data cards, using a companion LinkML schema. I'll walk through the design decisions behind the template, its alignment with existing standards, its treatment of sensitivity and governance metadata, and the phased roadmap toward lifecycle-integrated "xCards" that support autonomous discovery and reuse. The talk closes with current gaps, ongoing work, and how others can contribute datasets and feedback to the shared repository.

McSpadden, Helen [Thomas Jefferson National Accele↗

Report of the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

Scientific computing is undergoing rapid transformation as advances in artificial intelligence, heterogeneous computing, automation, and data-intensive research reshape not only computational tools but also the institutions, workforce models, and collaborative practices that support scientific discovery. This report synthesizes insights from the 2026 Workshop on Next-Generation Ecosystems for Scientific Computing, the second in a three-year series focused on strengthening scientific computing ecosystems through socio-technical co-design. Workshop discussions identified four interdependent strategic themes: software ecosystems for AI-enabled scientific discovery; trust, validation, and traceability; human-AI teaming and paradigm shifts; and workforce, pedagogy, and governance. The report translates these themes into eight priorities for community action spanning shared research infrastructure, trust and traceability, user experience, human-AI teaming, workforce development, cross-sector coordination, stewardship and sustainability, and evaluation of scientific value. Together, these priorities outline directions for building scientific computing ecosystems that remain trustworthy, sustainable, innovative, and resilient as AI assumes a growing role in scientific work.

AI↗

Leveraging History to Predict Infrequent Abnormal Transfers in Distributed Workflows

Scientific computing heavily relies on data shared by the community, especially in distributed data-intensive applications. This research focuses on predicting slow connections that create bottlenecks in distributed workflows. In this study, we analyze network traffic logs collected between January 2021 and August 2022 at the National Energy Research Scientific Computing Center (NERSC). Based on the observed patterns, we define a set of features primarily based on history for identifying low-performing data transfers. Typically, there are far fewer slow connections on well-maintained networks, which creates difficulty in learning to identify these abnormally slow connections from the normal ones. We devise several stratified sampling techniques to address the class-imbalance challenge and study how they affect the machine learning approaches. Our tests show that a relatively simple technique that undersamples the normal cases to balance the number of samples in two classes (normal and slow) is very effective for model training. This model predicts slow connections with an F1 score of 0.926.

97 MATHEMATICS AND COMPUTING↗

Challenges for Implementing FAIR Digital Objects with High Performance Workflows

New types of workflows are being used in science that couple traditional distributed and high-performance computing (HPC) with data-intensive approaches, and orchestrate ensembles of numerical simulations and artificial intelligence (AI) models. Such workflows may use AI models to supplement computation where numerical simulations may be too computationally expensive, to automate trivial yet time consuming operations, to perform preliminary selections among intractable numbers of combinations in domains as diverse as protein binding, fine-grid climate simulations, and drug discovery.

97 MATHEMATICS AND COMPUTING↗

Snowmass Topical Group Summary Report: IF04 -- Trigger and Data Acquisition Systems

A trend for future high energy physics experiments is an increase in the data bandwidth produced from the detectors. Datasets of the Petabyte scale have already become the norm, and the requirements of future experiments -- greater in size, exposure, and complexity -- will further push the limits of data acquisition technologies to data rates of exabytes per seconds. The challenge for these future data-intensive physics facilities lies in the reduction of the flow of data through a combination of sophisticated event selection in the form of high-performance triggers and improved data representation through compression and calculation of high-level quantities. These tasks must be performed with low-latency (i.e. in real-time) and often in extreme environments including high radiation, high magnetic fields, and cryogenic temperatures.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Snowmass Topical Group Summary Report: IF04 -- Trigger and Data Acquisition Systems

A trend for future high energy physics experiments is an increase in the data bandwidth produced from the detectors. Datasets of the Petabyte scale have already become the norm, and the requirements of future experiments -- greater in size, exposure, and complexity -- will further push the limits of data acquisition technologies to data rates of exabytes per seconds. The challenge for these future data-intensive physics facilities lies in the reduction of the flow of data through a combination of sophisticated event selection in the form of high-performance triggers and improved data representation through compression and calculation of high-level quantities. These tasks must be performed with low-latency (i.e. in real-time) and often in extreme environments including high radiation, high magnetic fields, and cryogenic temperatures.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Harnessing ML Privacy by Design Through Crossbar Array Non-idealities

Deep Neural Networks (DNNs), handling computeand data-intensive tasks, often utilize accelerators like Resistiveswitching Random-access Memory (RRAM) crossbar for energyefficient in-memory computation. Despite RRAM’s inherent nonidealities causing deviations in DNN output, this study transforms the weakness into strength. By leveraging RRAM non-idealities, the research enhances privacy protection against Membership Inference Attacks (MIAs), which reveal private information from training data. RRAM non-idealities disrupt MIA features, increasing model robustness and revealing a privacy-accuracy tradeoff. Empirical results with four MIAs and DNNs trained on different datasets demonstrate significant privacy leakage reduction with a minor accuracy drop (e.g., up to 2.8% for ResNet-18 with CIFAR-100).

artificial intelligence↗

photoD with Rubin ’s Data Preview 1: First stellar photometric distances and faint blue star deficits

Aims. We investigate the utility of Rubin’s Data Preview 1 (DP1) for estimating stellar number density profiles across the Milky Way halo. Methods. We used stellar broad-band near-UV to near-IR ugrizy photometry released in Rubin’s DP1 to estimate distance and metallicity for blue main sequence stars brighter than r = 24 in three ~1.1 sq. deg. fields at southern Galactic latitudes. Results. Compared to TRILEGAL simulations of the Galaxy’s stellar content, we found a likely deficit of blue main sequence turn-off stars with 22 < r < 24. We interpreted this discrepancy as a signature of a steeper halo number density profile at galactocentric distances 10–50 kpc than the canonical ~1/r 3 profile assumed in TRILEGAL simulations. Conclusions. This interpretation is consistent with earlier suggestions based on observations of more luminous, but much less numerous, evolved stellar populations, along with a few pencil beam surveys of blue main sequence stars in the northern sky. These results bode well for the future Galactic halo exploration with Rubin’s Legacy Survey of Space and Time (LSST).

Galaxy: fundamental parameters↗

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS↗

The Vera C. Rubin Observatory Data Preview 1

We present Rubin Data Preview 1 (DP1), the first data from the National Science Foundation–Department of Energy Vera C. Rubin Observatory, comprising raw and calibrated single-epoch images, coadds, difference images, detection catalogs, and ancillary data products. DP1 is based on 1792 optical–near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera (LSSTComCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile in late 2024. DP1 covers ∼15 deg 2 distributed across seven roughly equal-sized noncontiguous fields, each independently observed in six broad photometric bands, ugrizy. The median FWHM of the point-spread function across all bands is approximately 1"14, with the sharpest images reaching about 0." 58. The 5σ point-source depths for coadded images in the deepest field, the Extended Chandra Deep Field South, are u = 24.55, g = 26.18, r = 25.96, i = 25.71, z = 25.07, and y = 23.1. Other fields are no more than 2.2 mag shallower in any band, where they have nonzero coverage. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band in coadds, and 431 solar system objects, of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and is available to Vera C. Rubin Observatory data rights holders via the Rubin Science Platform, a cloud-based environment for the analysis of petascale astronomical data. While small compared to future LSST releases, its high quality and diversity of data support a broad range of early science investigations ahead of full operations in 2026.

Ground-based astronomy↗

RTN-095: The Vera C. Rubin Observatory Data Preview 1

We present Rubin Data Preview 1 (DP1), the first release of data from the NSF-DOE Vera C. Rubin Observatory, consisting of raw and calibrated single-epoch images, coadds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of ~15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. The median image quality across all bands, measured by the FWHM of the point-spread function, is approximately 1.13 arcseconds, with the sharpest images reaching about 0.65 arcseconds. DP1 contains approximately 2.3 million distinct astrophysical objects, of which 1.6 million are extended in at least one band, and 431 solar system objects, of which 93 are new discoveries. DP1 is approximately 3.5 TB in size and available to Rubin data rights holders via the Rubin Science Platform, a cloud-based environment for the analysis of petascale astronomical data. While small compared to future LSST releases, its high quality and diversity of data support a broad range of early science investigations across all four LSST themes, providing a valuable opportunity to engage with Rubin data ahead of the start of full operations in late 2025.

79 ASTRONOMY AND ASTROPHYSICS↗

SITCOMTN-154: Initial studies of photometric redshifts with LSSTComCam from DP1

This technote holds reports based on the first analyses of the Data Preview 1 (DP1) data by the Science Unit for photometric redshifts. Although photometric redshifts are not an official DP1 data product, the "Photo-z Science Unit" generated photo-z estimates for every galaxy in DP1 using the available multi-band imaging on a best-effort basis. This work included developing training and test datasets by matching DP1 data to high-quality reference redshifts obtained with spectroscopy, Grism data, and multi-band photometry. The Science Unit used the RAIL software package to make photometric redshift estimates using eight different algorithms, developed simple scientific performance metrics, used those metrics to explore how the performance of the algorithms varied with configuration changes, derived more optimized configurations of the algorithms and tested the performance of those configurations. This work, the resulting data products and expected data distribution mechanism are all described there.

79 ASTRONOMY AND ASTROPHYSICS↗

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

DMTN-277: The Monster: A reference catalog with synthetic ugrizy-band fluxes for the Vera C. Rubin observatory

In order to facilitate bootstrap photometric calibrations of early Rubin Observatory data we have created an all sky reference catalog called The Monster. This reference catalog uses a rank-ordered set of other reference catalogs to generate synthetic ugrizy-band fluxes that can be used calibrate images processed with the LSST science pipelines. This document describes the methodology used to create The Monster, documents the input external reference catalogs, and performs basic data validation of the first version of The Monster.

79 ASTRONOMY AND ASTROPHYSICS↗