Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “system metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

SPRUCE: Peat Core Sample Collection Metadata, Marcell Experimental Forest, Minnesota, August 2024

This data set contains metadata associated with peat core samples collected from the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment in August 2024. This sample metadata contains no analytical results and is a reference for analytical datasets. To ensure accessibility and discoverability, each sample was assigned an International Generic Sample Number (IGSN), a persistent identifier, using System for Earth and Extraterrestrial Sample Registration (SESAR). These samples were used for downstream analysis by multiple teams of researchers the results of which will be reported separately. This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. An aliquot of most samples is stored at Oak Ridge National Laboratory and may be available for further analysis. Access this collection event on SESAR https://doi.org/10.58052/IEJ9B00VQ. To inquire about obtaining archived samples for analysis, reach out using the Contact Sample Owner form located on the bottom of the landing page in SESAR. Note: Only dried and ground material from C Cores are available for new analysis.

Birkebak, Joshua [ORNL] (ORCID:0009000955611494)↗

Drifting Acoustic Measurements around C-Power's SeaRay WEC

The repository contains underwater noise measurements and associated metadata collected around C-Power's SeaRay wave energy converter on July 15, 2024 and July 16, 2024 while it was deployed at the U.S. Navy's Wave Energy Test Site (WETS) in Kaneohe, HI. Measurements were obtained using Drifting Acoustic Instrumentation SYstems (DAISYs). DAISYs consist of a surface expression connected to a hydrophone recording package by a tether. Both elements are instrumented to provide metadata (e.g., position, orientation, and depth). Information about how to build DAISYs is available at https://www.pmec.us/research-projects/daisy. The repository's primary content is a compressed archive (.zip format), containing multiple MATLAB binary data files (.mat format). The structure of each file is included in the repository as a Word document (Data Description MHK-DR.docx). Each file contains time series information for a single DAISY deployment (file naming convention: WETS_DAISY_[Drift #].mat) consisting of processed hydrophone data and associated metadata. During these measurements, C-Power's SeaRay was located at approximately 21.48112 N, 157.74451 W.

16 TIDAL AND WAVE POWER↗

Genomes OnLine Database (GOLD) v.8: overview and updates

The Genomes OnLine Database (GOLD) (https://gold.jgi.doe.gov/) is a manually curated, daily updated collection of genome projects and their metadata accumulated from around the world. The current version of the database includes over 1.17 million entries organized broadly into Studies (45 770), Organisms (387 382) or Biosamples (101 207), Sequencing Projects (355 364) and Analysis Projects (283 481). These four levels contain over 600 metadata fields, which includes 76 controlled vocabulary (CV) tables containing 3873 terms. GOLD provides an interactive web user interface for browsing and searching by a wide range of project and metadata fields. Users can enter details about their own projects in GOLD, which acts as a gatekeeper to ensure that metadata is accurately documented before submitting sequence information to the Integrated Microbial Genomes (IMG) system for analysis. In order to maintain a reference dataset for use by members of the scientific community, GOLD also imports projects from public repositories such as GenBank and SRA. Here, the current status of the database, along with recent updates and improvements are described in this manuscript.

59 BASIC BIOLOGICAL SCIENCES↗

EVSE Characterization

NextGen Profiles' EVSE characterization efforts explored performance variability in production EVSE through the use of EV emulation equipment and assessed how different operational conditions influence charging behavior. Data were collected at a frequency of 10 Hz from both the EV emulator and EVSE during each charge session and stored in a time-series database for further analysis. As part of the NextGen Profiles project, characterization of high-power EVSE was performed on both conductive and wireless charging infrastructure; however, only conductive charging data are currently included in this repository. This EVSE characterization was performed over a range of DC output currents and voltages, covering both nominal and off-nominal test conditions. This EVSE characterization dataset includes high-power charging data from two types of 350-kW-capable EVSE using liquid-cooled Combined Charging System-1 (CCS1, North American version) cables and connectors. To protect confidentiality, all EVSE metadata are anonymized, and the publicly released datasets are metered at 10-Hz frequency.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Optode Images, and Dissolved Oxygen Data from Sand and Sediment Flow-through Column Experiments associated with: "2-D Imaging of Dissolved Oxygen Concentration in Flow-Through Sediment Columns"

This data package is associated with the publication "2-D Imaging of Dissolved Oxygen Concentration in Flow-Through Sediment Columns" submitted to Frontiers in Water (Garayburu-Caruso et al., 2022). This work presents a novel application of planar optode technology to sediment columns that allows the study of well-constrained flow fields without the need to assume completely homogeneous flow. We introduce a flow-through column with the interior coated in an oxygen sensing film and a series of sampling ports as a unique tool capable of capturing changes in dissolved oxygen (DO) at high frequency (10 s) and spatial (< 1 mm) resolution across the length of a column. In addition, we provide custom scripts that allows image processing steps to be automated and consistently reproduced over a variety of experimental conditions. This dataset is comprised of four folders (1) 01_Riverbed_Sediment_Experiment, (2) 02_Sand_Experiment, (3) 01_Riverbed_Sediment_Raw_Images, and (4) 02_Sand_Experiment_Raw_Images. 01_Riverbed_Sediment_Experiment contains (1) a subfolder with csv files that contain statistical parameters from the linear regressions applied to each optode image, (2) a subfolder with csv files from DO data measured by in-line sensors, (3) a subfolder with csv files containing column and injection specific input parameters for the image processing script, and (4) a subfolder with scripts used to process the images and create associated figures. 02_Sand_Experiment contains (1) a subfolder with csv files containing column and injection-specific input parameters for the image processing script, and (2) a subfolder with scripts used to process the images and create associated figures. 01_Riverbed_Sediment_Raw_Images (spited in 3 parts due to size) contains a subfolder with raw red, and green images from the optode system during the MilliQ Injection, Vanillin Injection or Vanillin + Sampling Injection respectively. 02_Sand_Experiment_Raw_Images contains a subfolder with raw red, and green images from the optode system. Outside of the main folders there is a csv containing file-level metadata and a csv data dictionary defining column headers for all csv files contained in the data package.

54 ENVIRONMENTAL SCIENCES↗

History and Status of ALSEP and the Apollo Lunar Data Project

A suite of automated scientific instruments (the Apollo Lunar Surface Experiment Package, or ALSEP) was installed at each of the landing sites of Apollo 12, 14, 15, 16, and 17 from 1969 to 1972. They operated from deployment until decommissioning on 30 September 1977. These data were continuously transmitted to Earth and saved on the Range Tapes, which were recorded at the Manned Space Flight Network stations. These data were also broken out by experiment and sent to the experiment Principal Investigators on what were called the P.I. Tapes. Starting in April 1973 the Range Tape data were stored in digital format on 7-track magnetic tapes, the ARCSAV Tapes. In February 1976, the handling of the Range Tapes was transferred to UT Galveston. They produced 9-track tapes referred to as the Work Tapes. Following the Apollo program the Range and ARCSAV tapes, which were never archived, were lost. The Work Tapes were archived at the National Space Science Data Center (NSSDC). Some investigators archived their individual experiment data with NSSDC as well, but much of the data had minimal documentation, were not in digital form, or were stored in difficult to translate formats. Data from many experiments were never delivered to the NSSDC. The Lunar Data Project was started to address the problem of both missing and not readily usable data. Our effort has resulted in recovery of some of the ARCSAV tapes, recovery and digitization of a large volume of Apollo scientific and technical documentation, and restoration of many ALSEP and other Apollo data collections. Restoration involves deciphering formats, assembling necessary ancillary data (metadata), and packaging data in digital format to be archived with the Planetary Data System (PDS). Recovery of the data from the ARCSAV tapes involved having the tapes read on special equipment and extracting the individual experiment data out of the integrated data stream. We will report on the history and status of the various recovery efforts.

Work Tapes↗

Using Knowledge Analytics to Search and Characterize Mass Properties Aerospace Data

There is growing capability in the field of “Big Data” and “Data Analytics” which Mass Properties Engineers might like to take advantage of. This paper utilizes an implementation of the IBM Knowledge Analytics and Watson search capabilities to explore a corpus of material developed primarily with the interests of Mass Properties Engineers and vehicle concept developers at its forefront. The full collection of SAWE (Society of Allied Weight Engineers, Inc.) Technical Papers from 1939 through 2015 is a major portion of the knowledge content. Additional aerospace vehicle design information includes metadata from AIAA (American Institute for Aeronautics and Astronautics), and INCOSE (International Council on Systems Engineering) as well as author-provided personal search material. This data is processed with respect to certain expected content, data taxonomies and key words to become the core data in NASA Langley Research Center’s “Vehicle Analysis Analytics”, IBM Watson Content. Processed data becomes the corpus of information which is interrogated to provide examples of finding data for mass regression analysis, technology impacts on MPE (Mass Properties Engineering), mass properties control, standards, and other aspects of interest.

Cerro, Jeffrey A.↗

A Standard Reference Model for Data Archives

An implementable Data Archive Architecture is being developed for trusted digital repositories based on the Reference Model for an Open Archival Information System (OAIS) – ISO 14721. A set of interoperable protocols and interface specifications are planned that will offer capabilities for accessing, merging, and re-using data, both within and across the operational boundaries of trustworthy digital repositories. The model will also provide support for the fundamental scientific need to verify the reproducibility of results. This standards development task is being performed by the Data Archive Interoperability (DAI) working group within the Consultative Committee for Space Data Systems (CCSDS). The architecture integrates concepts from the OAIS Reference Model, the ISO/IEC 11179 Metadata Registry (MDR) standard, the CCSDS Reference Architecture for Space Information Management (RASIM), the proposed draft recommended practice document, Information Preparation to Enable Long Term Use (IPELTU), and three decades of digital repository development for science research.

Ambacher, Bruce↗

Water Observations of Flow/No-Flow for the East-Taylor Watershed, Colorado (June-July 2025 and 2026)

This dataset provides multi-year, ground-truth visual observations of surface water flow/no-flow conditions within the East-Taylor Watershed, Colorado, collected during June and July of 2025 and 2026. In June and July 2025, on-the-ground visual observations of flow/no-flow were collected as part of the Watershed Function Scientific Focus Area (SFA) and Rocky Mountain Biological Laboratory (RMBL) Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign (further details are provided within the CHESS Project Description). We obtained 377 water observations of flow/no-flow within the East-Taylor Watershed, Colorado. These ground-truth observations were collected to validate classification maps from remote sensing data and model results within the East-Taylor Watershed. In 2025, flow/no-flow measurements were collected using a field-based app for the CHESS Campaign (Zerion iForm). Within the field app, a water observation form was created to collect coordinates and metadata about the observation. Information collected for the water observation points included information about visually-assessed streamflow presence/absence (standard question obtained from Colorado State University’s StreamTracker project), flow estimate, stream or ponded area width, canopy cover, manganese films, iron seeps, and beaver activity. For 2025 water observations, this dataset contains: (1) a data file with the water observations and coordinates (2025_Water_Observations.csv); (2) a Keyhole Markup Language Zipped (KMZ) with the water observation locations and metadata (2025_Water_Observations_Locations.kmz); (3) photos (.jpg and .jpeg) of the water observation points, organized by location, contained within 2025_Water_Observations_FieldPhotographs.zip file; and (4) water observation protocols and figures (2025_Water_Observation_Protocols.pdf). In June and July 2026, on-the-ground visual observations of flow/no-flow were collected as part of the Watershed Function SFA project. We obtained 365 water observations of flow/no-flow within the East-Taylor Watershed, Colorado. The 2026 observations focused on collecting repeat measurements at the 2025 flow/no-flow observation locations conducted as part of the CHESS campaign. These ground-truth observations were collected to understand differences in flow/no-flow in 2026, given the unprecedented 2026 drought in Colorado. In 2026, flow/no-flow measurements were collected using ArcGIS (Geographic Information System) Survey123. Within the field app, a water observation form was created to collect coordinates and metadata about the observation. Information collected for the water observation points included repeat information from the 2025 water observation effort, including visually-assessed streamflow presence/absence (standard question obtained from Colorado State University’s StreamTracker project), flow estimate, stream or ponded area width, canopy cover, manganese films, iron seeps, beaver activity, and a new metadata component of estimated stream depth (for select locations). For 2026 water observations, this dataset contains: (1) a data file with the water observations and coordinates (2026_Water_Observations.csv); (2) a Keyhole Markup Language Zipped (KMZ) with the water observation locations and metadata (2026_Water_Observations_Locations.kmz); (3) photos (.jpg) of the water observation points, organized by location, contained within 2026_Water_Observations_FieldPhotographs.zip file; and (4) water observation protocols and figures (2026_Water_Observation_Protocols.pdf). For 2025 and 2026 water observations, this dataset contains: (1) a location metadata file (locations.csv); (6) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and (7) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. 2026-09-02: This dataset was updated to include 2026 water observation measurements. The 2025 observation files were also updated to ensure a consistent file naming convention across water observation years.

2018 NEON and 2025 CHESS Campaigns↗

CDDIS: NASA's Archive of Space Geodesy Data and Products Supporting GGOS

The Crustal Dynamics Data Information System (CDDIS) supports data archiving and distribution activities for the space geodesy and geodynamics community. The main objectives of the system are to store space geodesy and geodynamics related data and products in a central archive, to maintain information about the archival of these data,to disseminate these data and information in a timely manner to a global scientific research community, and provide user based tools for the exploration and use of the archive. The CDDIS data system and its archive is a key component in several of the geometric services within the International Association of Geodesy (IAG) and its observing systemthe Global Geodetic Observing System (GGOS), including the IGS, the International DORIS Service (IDS), the International Laser Ranging Service (ILRS), the International VLBI Service for Geodesy and Astrometry (IVS), and the International Earth Rotation and Reference Systems Service (IERS). The CDDIS provides on-line access to over 17 Tbytes of dataand derived products in support of the IAG services and GGOS. The systems archive continues to grow and improve as new activities are supported and enhancements are implemented. Recently, the CDDIS has established a real-time streaming capability for GNSS data and products. Furthermore, enhancements to metadata describing the contents ofthe archive have been developed to facilitate data discovery. This poster will provide a review of the improvements in the system infrastructure that CDDIS has made over the past year for the geodetic community and describe future plans for the system.

CNS↗

High Performance Access to Archival Data Stored in HDF4 and HDF5 on Cloud Object Stores Without Reformatting the Files

Cloud computing offers numerous advantages for users of extensive Earth science data collections. These benefits encompass direct online access to data files and granules from any location, scalable access supporting parallel computing workflows, and flexible computing tools enabling innovative experimentation with processing techniques. However, older archival file formats designed for distinct computing systems hinder efficient access to decade-long time-series data when compared to data stored in modern cloud-optimized formats like Web Object Stores (WOS), exemplified by Amazon Web Services’ Simple Storage Service (S3). We describe DMR++ (Dataset Metadata Response plus plus), a technology facilitating efficient access to HDF5 (Hierarchical Data Format, version 5) and HDF4 files stored on WOS systems without requiring data reformatting. DMR++ achieves performance comparable to technologies like Zarr while preserving the original file structure, a substantial benefit considering the vast quantity of archival files held by organizations such as NASA. Moreover, DMR++ typically outperforms cloud-optimized versions of HDF5. Essentially an XML (Extensible Markup Language) document usually stored alongside the described data, DMR++ can also be generated on-the-fly but is generally created during data staging to the WOS. Archival files that use HDF4/5 often store large arrays of numerical data. The data in these files is often compressed, typically reducing their size by a factor of four or more. To achieve efficient access to portions of those arrays, they are 'chunked' into smaller sub-arrays, each individually compressed. The chunk size is a compromise, where spinning disks can efficiently access data in smaller chunks while S3 favors larger chunks. A simple optimization of aggregating smaller chunks that are stored adjacently, transferring them in a single access and then individually decompressing them will improve performance. NASA data pose an additional challenge: special Application Programmer Interface (API) libraries are often needed to compute some variables. These libraries are incompatible with WOS environments. Our solution involves storing computed values in the DMR++ document or a companion file, making them accessible like other variables and eliminating the need for specialized APIs. We outline specific optimizations for both satellite grid and swath data stored in HDF4-EOS2 (Earth Observing System).

James Gallagher↗

Metadata Schemas and Ontologies for Building Energy Applications: A Critical Review and Use Case Analysis

With the increasing digitalization of processes throughout the lifecycle of buildings, data exchanged between stakeholders and between building systems has grown significantly. However, a lack of semantic interoperability between data in different systems is still prevalent, hindering the development of applications that can be reused across buildings and limiting the scalability of innovative solutions. Semantics refers to the description of the meaning of the data in a way that can be consistently understood by applications. Recently, several competing initiatives have been developing metadata schemas and ontologies to express this semantic information for different applications in the building domain. This paper systematically reviews these schemas and conducts an analysis of five of them to evaluate their applicability to three high-value use cases for building operations: energy audits, automated fault detection and diagnostics and optimal control. The survey finds 40 schemas published in the last 10 years but but their actual use in industry is difficult to estimate. Among the five selected ontologies, several gaps are highlighted in relation to the three use cases. Recommendations for the future include better harmonization of these initiatives, more centralized repositories and search engines for these schemas as well as better industry engagement to facilitate their adoption.

Smart Building, Sematic, Metadata, Ontology, Data ↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

Control systems and data management for high-power laser facilities

The next generation of high-power lasers enables repetition of experiments at orders of magnitude higher frequency than what was possible using the prior generation. Facilities requiring human intervention between laser repetitions need to adapt in order to keep pace with the new laser technology. A distributed networked control system can enable laboratory-wide automation and feedback control loops. These higher-repetition-rate experiments will create enormous quantities of data. A consistent approach to managing data can increase data accessibility, reduce repetitive data-software development and mitigate poorly organized metadata. An opportunity arises to share knowledge of improvements to control and data infrastructure currently being undertaken. We compare platforms and approaches to state-of-the-art control systems and data management at high-power laser facilities, and we illustrate these topics with case studies from our community

47 OTHER INSTRUMENTATION↗

Development of a Digital Twin for Hydrogen Dispersion and Safety Assessment in an Electrolyzer Based Hydrogen Production Facility

Digital twin models are virtual representations of physical systems that use real-time data to simulate and optimize performance. This study presents the development and initial implementation of a digital twin (DT) for the electrolyzer-based hydrogen production facility at NREL's Advanced Research on Integrated Energy Systems (ARIES), focused on enhancing safety and optimizing sensor placement through physics-based simulations and metadata integration. The DT incorporates detailed facility-specific information, including component layout, leak locations, and controlled release parameters, to model hydrogen dispersion under varying environmental conditions. Using steady-state computational fluid dynamics (CFD) simulations informed by real meteorological data, such as wind speed, direction, and vertical wind profiles, the DT enables visualization of hydrogen plume behavior and spatial concentration distributions. Comparative analysis between high and low wind speed scenarios illustrates the significant influence of wind dynamics on plume shape and extent, with horizontal momentum dominating dispersion at higher speeds, while buoyancy effects become more prominent under low wind conditions. These simulations generate a rich dataset embedded within the DT, allowing users to assess potential leak outcomes and identify optimal sensor locations based on concentration thresholds. The model supports scenario-based analysis to guide safety strategies and equipment deployment for open-area hydrogen infrastructure. The digital twin thus serves as a dynamic platform for virtual prototyping, providing predictive insight into hydrogen behavior and enhancing risk-informed decision-making. This initial phase establishes a validated foundation for future integration of transient, uncontrolled leak scenarios and real-time sensor feedback, positioning the DT as a critical tool for safety design, operational planning, and adaptive monitoring in hydrogen systems. Overall, the approach demonstrates the value of combining environmental data with digital simulations to inform safer and more efficient deployment of hydrogen technologies.

08 HYDROGEN↗

Development of a Digital Twin for Hydrogen Dispersion and Safety Assessment in an Electrolyzer-Based Hydrogen Production Facility: Preprint

Digital twin models are virtual representations of physical systems that use real-time data to simulate and optimize performance. This study presents the development and initial implementation of a digital twin (DT) for the electrolyzer-based hydrogen production facility at the National Renewable Energy Laboratory (NREL)'s Advanced Research on Integrated Energy Systems (ARIES), focused on enhancing safety and optimizing sensor placement through physics-based simulations and metadata integration. The DT incorporates detailed facility-specific information, including component layout, leak locations, and controlled release parameters, to model hydrogen dispersion under varying environmental conditions. Using steady-state computational fluid dynamics (CFD) simulations informed by real meteorological data, such as wind speed, direction, and vertical wind profiles, the DT enables visualization of hydrogen plume behavior and spatial concentration distributions. Comparative analysis between high and low wind speed scenarios illustrates the significant influence of wind dynamics on plume shape and extent, with horizontal momentum dominating dispersion at higher speeds, while buoyancy effects become more prominent under low wind conditions. These simulations generate a rich dataset embedded within the DT, allowing users to assess potential leak outcomes and identify optimal sensor locations based on concentration thresholds. The model supports scenario-based analysis to guide safety strategies and equipment deployment for open-area hydrogen infrastructure. The digital twin thus serves as a dynamic platform for virtual prototyping, providing predictive insight into hydrogen behavior and enhancing risk-informed decision-making. This initial phase establishes a validated foundation for future integration of transient, uncontrolled leak scenarios and real-time sensor feedback, positioning the DT as a critical tool for safety design, operational planning, and adaptive monitoring in hydrogen systems. Overall, the approach demonstrates the value of combining environmental data with digital simulations to inform safer and more efficient deployment of hydrogen technologies.

08 HYDROGEN↗

Frontier Job-Centric Telemetry Dataset

Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Software Implements a Space-Mission File-Transfer Protocol

CFDP is a computer program that implements the CCSDS (Consultative Committee for Space Data Systems) File Delivery Protocol, which is an international standard for automatic, reliable transfers of files of data between locations on Earth and in outer space. CFDP administers concurrent file transfers in both directions, delivery of data out of transmission order, reliable and unreliable transmission modes, and automatic retransmission of lost or corrupted data by use of one or more of several lost-segment-detection modes. The program also implements several data-integrity measures, including file checksums and optional cyclic redundancy checks for each protocol data unit. The metadata accompanying each file can include messages to users application programs and commands for operating on remote file systems.

Rundstrom, Kathleen↗