Documents and Records Management Plan
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
This standard presents a set of optional metadata fields for managed digital objects within the Nuclear Security Enterprise (NSE) and provides a deeper look at data representation in metadata by looking at the representation of 1) Records Management required metadata, and 2) common representations of technical/scientific data. Metadata standardization is a critical enabler for effectively sharing data, documents, and other digital objects between NSE sites, and for tracing the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a complementary field standard, recommending an optional set of fields that should be uniformly built for all managed digital objects within the NSE. This document specifically focuses on extending the shared discovery layer defined in the first white paper by introducing additional descriptive and data representation fields that improve cross-site search and interpretation.
Record of Technical Change for 30 CAUs in the Use Restriction Management Plan, Revision 2
This project focuses on the digitization and public accessibility to the Fusion Safety Archives at the Idaho National Laboratory. The first phase involves a thorough review of each document in the physical archives to determine its online availability. For documents that are available online, PDF copies and unique identifiers are collected for database integration. Documents not available online are delivered to Red Inc. for digitization. Additionally, defunct storage devices such as diskettes are sent to INL’s archival department for data retrieval where possible. The second phase of the project involves the creation of a comprehensive database to house the digital copies of the archives. The database will facilitate easy access and management of the digitized documents. Following the database creation, we plan to train a Retrieval-Augmented Generation (RAG) based AI on publicly available documents. The trained AI will be integrated into a front-facing application, allowing the public to easily access information from the Fusion Safety Archives. This project aims to preserve valuable historical data, improve accessibility, and promote transparency in fusion safety research.
Record of Technical Change for 17 CAUs in the Use Restriction Management Plan
The Advanced Materials and Manufacturing Technologies (AMMT) program continues to advance a data-driven approach to demonstrate the utility of additive manufacturing for fabricating components for nuclear applications. A key scientific goal is to leverage data to better understand manufacturing outcomes and thereby improve the performance, reliability, and lifespan of nuclear components. Ultimately, this effort supports the development of standards for certification and qualification of additively manufactured components, enabling broader industry adoption. In support of this objective, the AMMT program is building and deploying a data management platform to record, index, analyze, and make available the manufacturing data generated across the AMMT program. In FY 2023, the team conceptualized the architecture of the platform and, in FY 2024, deployed the first functional version at the Oak Ridge National Laboratory (ORNL) Manufacturing Demonstration Facility (MDF). In FY 2025, the platform was officially opened to all AMMT members. To enable this expansion, core modifications and enhancements were developed, including improvements to the user interface and workflows for data entry and retrieval. Most notably, robust security and access control mechanisms were implemented to protect data and manage information sharing. This effort featured a logging system, protected views, and controlled access mechanisms. This report documents these enhancements and the transition of the platform into program-wide use.
The Savannah River Site (SRS) Defense Waste Processing Facility (DWPF) processes a Monosodium Titanate/Sludge Solids (MST/SS) waste stream received from the Salt Waste Processing Facility (SWPF) via the Precipitate Reactor Feed Tank (PRFT). During processing, DWPF is required to provide evidence of compliance with the Waste Acceptance Product Specifications (WAPS) to ensure acceptance of their vitrified high-level waste (HLW) into the Civilian Radioactive Waste Management System. Production Records must document the constituents of the MST/SS material in the PRFT from each salt batch (StB) processed at SWPF. Savannah River Mission Completion (SRMC) has requested Savannah River National Laboratory (SRNL) to analyze PRFT samples representing each SWPF salt batch for thirty-two radionuclides. Additionally, elemental analysis of PRFT slurry and MST/SS solids was performed to aid SRMC in further refinement of the inputs and assumptions used in future frit development and Material Tracking Program calculations. The analyses of PRFT Batches 15, 22, and 26, which corresponds to material from the processing of StB4, StB5, and StB7, respectively, are reported herein. The unwashed dried solids of the PRFT batches were found to be 86-87% MST. Additionally, the total sulfur values are well below the assumed 982 mg of sulfate/kg of PRFT slurry used in Material Tracking Program calculations.
The US Department of Energy (DOE) Oak Ridge Reservation (ORR) is located in Anderson and Roane Counties, Tennessee. A portion of the ORR, known as Self-Sufficiency Parcel 2 (SSP2) is planned for transfer for private use. The SSP2 Site is approximately 670 acres (Figure 1-1), although the current plan is to only clear and develop a portion of this acreage. Any inquiries about the land transfer and future development should be directed to DOE Oak Ridge Environmental Management, as this is beyond the scope of the Natural Resources Management Team (NRMT). NRMT records bat data for the entire ORR, including acoustic monitoring, mist netting and cave surveys. A few surveys have previously been conducted for small land transfers adjacent to SSP2 (See Appendix A), but not for the entire SSP2 area. Since bats have a large range, it was decided that collecting data while SSP2 was still accessible would be beneficial for the NRMT dataset. Acoustic data was therefore collected within and near SSP2 during the summer of 2024. This write-up is not a Biological Assessment (BA). However, the data and information provided can be used during the creation of a BA and consultations with US Fish and Wildlife Service (USFWS) in order to comply with federal directives of the Endangered Species Act of 1973 (16 U. S. C. 153 et seq.). The SSP2 site was surveyed during summer roosting/maternity season of 2024 using ultrasonic acoustic monitors to record calls from all bat species whose home ranges include the ORR. Special note was taken for presence of Federally listed Endangered and Threatened (T&E) bat species, as well as bat species which are Proposed for Federal listing, Candidate for federal listing, and state listed. Summer roosting season, from May 15 to August 15, is crucial to forest-dwelling T&E bat species for rearing young and foraging. Results of these surveys indicate the presence of three Federally listed bat species: Gray bat (Myotis grisescens--Endangered), Indiana bat (Myotis sodalis--Endangered), and Northern long-eared bat (Myotis septentrionalis--Endangered). Two additional bat species were present on the SSP2 Site: Tricolored bat (Perimyotis subflavus--Proposed for Federal listing) and Little brown bat (Myotis lucifugus—Candidate for Federal listing).
A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Recalculated based on data sourced from WestDAAT - Changed using a Site ID column to identify unique records to using aa combination of Site ID and Allocation ID - Removed the Water Management Area (WMA) column from the harmonized records. The replacement is a separate file which stores the relationship between allocations and WMAs. This allows for allocations to contribute to water right amounts to multiple WMAs during the subsequent cumulative process. - Added a column describing a water rights legal status - Added "Unspecified" was a water source category - Added an acre-foot (AF) column - Added a column for the classification of the right's owner v1.02 - Added a .RData file to the dataset as a convenience for anyone exploring our code. This is an internal file, and the one referenced in analysis scripts as the data objects are already in R data objects. v1.01 - Updated the names of each file with an ID number less than 3 digits to include leading 0s v1.0 - Initial public release Description Here we present an updated database of Western U.S. water right records. This database provides consistent unique identifiers for each water right record, and a consistent categorization scheme that puts each water right record into one of seven broad use categories. These data were instrumental in conducting a study of the multi-sector dynamics of inter-sectoral water allocation changes though water markets (Grogan et al., *in review*). Specifically, the data were formatted for use as input to a process-based hydrologic model, Water Balance Model (WBM), with a water rights module (Grogan et al., *in review*). While this specific study motivated the development of the database presented here, water management in the U.S. West is a rich area of study (e.g., Anderson and Woosly, 2005; Tidwell, 2014; Null and Prudencio, 2016; Carney et al., 2021) so releasing this database publicly with documentation and usage notes will enable other researchers to do further work on water management in the U.S. West. We produced the water rights database presented here in four main steps: (1) data collection, (2) data quality control, (3) data harmonization, and (4) generation of cumulative water rights curves. Each of steps (1)-(3) had to be completed in order to produce (4), the final product that was used in the modeling exercise in Grogan et al. (*in review*). All data in each step is associated with a spatial unit called a Water Management Area (WMA), which is the unit of water right administration utilized by the state in which the right came from. Steps (2) and (3) required use to make assumptions and interpretation, and to remove records from the raw data collection. We describe each of these assumptions and interpretations below so that other researchers can choose to implement alternative assumptions an interpretation as fits their research aims. Motivation for Changing Data Sources The most significant change has been a switch from collecting the raw water rights directly from each state to using the water rights records presented in WestDAAT, a product of the Water Data Exchange (WaDE) Program under the Western States Water Council (WSWC). One of the main reasons for this is that each state of interest is a member of the WSWC, meaning that WaDE is partially funded by these states, as well as many universities. As WestDAAT is also a database with consistent categorization, it has allowed us to spend less time on data collection and quality control and more time on answering research questions. This has included records from water right sources we had previously not known about when creating v1.0 of this database. The only major downside to utilizing the WestDAAT records as our raw data is that further updates are tied to when WestDAAT is updated, as some states update their public water right records daily. However, as our focus is on cumulative water amounts at the regional scale, it is unlikely most records updates would have a significant effect on our results. The structure of WestDAAT led to several important changes to how HarWR is formatted. The most significant change is that WaDE has calculated a field known as `SiteUUID`, which is a unique identifier for the Point of Diversion (POD), or where the water is drawn from. This separate from `AllocationNativeID`, which is the identifier for the allocation of water, or the amount of water associated with the water right. It should be noted that it is possible for a single site to have multiple allocations associated with it and for an allocation to be able to be extracted from multiple sites. The site-allocation structure has allowed us to adapt a more consistent, and hopefully more realistic, approach in organizing the water right records than we had with HarDWR v1.0. This was incredibly helpful as the raw data from many states had multiple water uses within a single field within a single row of their raw data, and it was not always clear if the first water use was the most important, or simply first alphabetically. WestDAAT has already addressed this data quality issue. Furthermore, with v1.0, when there were multiple records with the same water right ID, we selected the largest volume or flow amount and disregarded the rest. As WestDAAT was already a common structure for disparate data formats, we were better able to identify sites with multiple allocations and, perhaps more importantly, allocations with multiple sites. This is particularly helpful when an allocation has sites which cross WMA boundaries, instead of just assigning the full water amount to a single WMA we are now able to divide the amount of water between the number of relevant WMAs. As it is now possible to identify allocations with water used in multiple WMAs, it is no longer practical to store this information within a single column. Instead the stAllocationToWMATab.csv file was created, which is an allocation by WMA matrix containing the percent Place of Use area overlap with each WMA. We then use this percentage to divide the allocation's flow amount between the given WMAs during the cumulation process to hopefully provide more realistic totals of water use in each area. However, not every state provides areas of water use, so like HarDWR v1.0, a hierarchical decision tree was used to assign each allocation to a WMA. First, if a WMA could be identified based on the allocation ID, then that WMA was used; typically, when available, this applied to the entire state and no further steps were needed. Second was the spatial analysis of Place of Use to WMAs. Third was a spatial analysis of the POD locations to WMAs, with the assumption that allocation's POD is within the WMA it should belong to; if an allocation still had multiple WMAs based on its POD locations, then the allocation's flow amount would be divided equally between all WMAs. The fourth, and final, process was to include water allocations which spatially fell outside of the state WMA boundaries. This could be due to several reasons, such as coordinate errors / imprecision in the POD location, imprecision in the WMA boundaries, or rights attached with features, such as a reservoir, which crosses state boundaries. To include these records, we decided for any POD which was within one kilometer of the state's edge would be assigned to the nearest WMA. Other Changes WestDAAT has Allowed In addition to a more nuanced and consistent method of assigning water right's data to WMAs, there are other benefits gained from using the WestDAAT dataset. Among those is a consistent categorization of a water right's legal status. In HarDWR v1.0, legal status was effectively ignored, which led to many valid concerns about the quality of the database related to the amounts of water the rights allowed to be claimed. The main issue was that rights with legal status' such as "application withdrawn", "non-active", or "cancelled" were included within HarDWR v1.0. These, and other water rights status' which were deemed to not be in use have been removed from this version of the database. Another major change has been the addition of the "unspecified water source category. This is water that can come from either surface water or groundwater, or the source of which is unknown. The addition of this source category brings the total number of categories to three. Due to reviewer feedback, we decided to add the acre-foot (AF) column so that the data may be more applicable to a wider audience. We added the ownerClassification column so that the data may be more applicable to a wider audience. File Descriptions The dataset is a series of various files organized by state sub-directories. In addition, each file begins with the state's name, in case the file is separate from its sub-directory for some reason. After the state name is the text which describes the contents of the file. Here is each file described in detail. Note that st is a placeholder for the state's name. stFullRecords_HarmonizedRights.csv: A file of the complete water records for each state. The column headers for each of this type of file are: state - The name of the state to which the allocations belong to. FIPS - The two digit numeric state ID code. siteID - The site location ID for POD locations. A site may have multiple allocations, which are the actual amount of water which can be drawn. In a simplified hypothetical, a farm stead may have an allocation for "irrigation" and an allocation for "domestic" water use, but the water is drawn from the same pumping equipment. It should be noted that many of the site ID appear to have been added by WaDE, and therefore may not be recognized by a given state's water rights database. allocationID - The allocation ID for the water right. For most states this is the water right ID, and what is recommended to use should a right be looked up on a given state's water rights database. The water amounts associated with these IDs tend to be finer scaled than those associated with siteID. It should be noted that some allocations may be extracted from multiple sites, particularly for larger Places of Use. ownerClassification - A classification of the types of owners for water rights. The most common is `Private` which incorporates a wide range of entities. Several classifications would be grouped into a government category, most of which are for the U.S. Federal Government. These allocations could be listed as "Federal", "United States of America", or as the names of any number of federal agencies. The last major grouping of entities is for "Native American"s. priorityDate - The date we use as the water right priority date for our modeling analysis. This is the legal priority date when it is available. However, for some rights, specifically from California and New Mexico, we used a pseudo priority date (e.g. well completion date or start of well drilling date) when a legal priority date was not available. The most questionable dates come from New Mexico, where the only date associated with certain water right records was the date the allocation was recorded in the database. As the allocation record creation tended to be within a few months of the filing of the application of the water right, from manually double checking the water rights, and our analysis focuses on aggregating water rights on the timescale of years, we determined it was acceptable to use such dates to include as many records as possible. primaryBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories WestDAAT. This column is the original WaDE category for the primary water use at the PoD site. allocationBeneficialUse - From the numerous state water use categories, WaDE categorized them into 21 categories for WestDAAT. This column is the original WaDE category
Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. These experiments generate petabytes of data, with volumes soon expected to reach exabytes. Consequently, there is a growing need for computation, including structured data processing from raw data to consumer-ready derived data, extensive Monte Carlo simulation campaigns, and a wide range of end-user analysis. To manage these computational and storage demands, centralized workflow and data management systems are implemented. However, decisions regarding data placement and payload allocation are often made disjointly and via heuristic means. A significant obstacle in adopting more effective heuristic or AI-driven solutions is the absence of a quick and reliable introspective dynamic model to evaluate and refine alternative approaches. In this study, we aim to develop such an interactive system using real-world data. By examining job execution records from the PanDA workflow management system, we have pinpointed key performance indicators such as queuing time, error rate, and the extent of remote data access. The dataset includes five months of activity. Additionally, we are creating a generative AI model to simulate time series of payloads, which incorporate visible features like category, event count, and submitting group, as well as hidden features like the total computational load—derived from existing PanDA records and computing site capabilities. These hidden features, which are not visible to job allocators, whether heuristic or AI-driven, influence factors such as queuing times and data movement.
Explore the source record for details and available documents.
Overview This dataset compiles dam sediment management and monitoring information from surveys, case studies, and journal articles. Additionally, features described by the National Inventory of Dams (i.e., presence of sluice gates) are included to indicate known infrastructure features that may address sediment releases. The location and description of records from downstream monitoring gages are catalogued in order to help with tracking conditions over time (e.g., before and after management actions, as operations change, etc.). The data help address national scale understanding of challenges and solutions related to the accumulation of sediment behind a dam as well as downstream passage. Sediment trapping causes problems as it reduces storage capacity, disrupts dam and reservoir function, impedes access for recreation, alters water quality/habitat conditions, and contributes to riverbank and coastal erosion within the reservoir. Data compilation from a variety of sources is a first step towards assessing system-wide efficacy of management solutions. This dataset was developed under the Water Power Technologies Office funded effort which began as a Seedling on Reservoir Sedimentation Data, and was supported by the Reservoir Sedimentation Modeling Framework and Data Analysis project. These projects have addressed challenges in describing sediment transport, trapping, and management at dams throughout the US. Methodology An outer join on dams/reservoirs with surveys and survey reports (documented in the RESSED database, USBR or USACE databases, project websites, etc.) with the National Inventory of Dams, based on the NIDID to determine dams with documented management and/or sluice gates. Additional dams with documented management activity were identified through review of technical articles from the past 25 years in Journal of Hydrology, Journal of Water Resources Planning and Management, Geomorphology, Journal of Hydraulic Engineering, Water, Journal of Cleaner Production, International Journal of Sediment Research, Nature Scientific Reports, Earth Surface Processes and Landforms, and Environmental Science and Pollution Research. Individual records were created for each survey or management activity documented. To evaluate downstream sediment monitoring records, the nhdPlusTools and dataRetrieval packages in R were used to find gages within 10km of each dam in the management database. Length of record and location of matched gages were retrieved for those parameters relevant to sediment concentration or total sediment discharge.
Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.
The Glen Canyon Dam (GCD) is a Colorado River Storage Project (CRSP) power resource that is a component of the Salt Lake City Area Integrated Projects (SLCA/IP). The 2016 record of decision (ROD) for the GCD long-term experimental and management plan (LTEMP) final Environmental Impact Statement (EIS) specified criteria for GCD monthly water releases, daily and hourly operating limits, and experimental releases. This report examines the financial implications of the high flow experiment (HFE) conducted at GCD during the spring of Water Year (WY) 2023 as required by the LTEMP HFE Protocol. This report is part of a series of reports that describe the financial costs of LTEMP experimental releases since the 2016 ROD was adopted in January 2017. Previous reports analyzed the impact of several past HFEs and Bug Flow Experiments. This report focuses on the HFE conducted in April 2023. For this experimental release, financial costs of approximately $1.33 million were incurred because the HFE required sustained water releases exceeding the power plant’s maximum turbine flow rate. In addition, during the experiment, operators were not allowed to shape GCD power production, either to follow Firm Electric Service (FES) customer day-ahead energy deliveries or to respond to market prices. This study identifies the main factors contributing to the HFE costs and examines the interdependencies among these factors. It applies an integrated set of tools to estimate Western Area Power Administration (WAPA) financial impacts by simulating GCD under two types of cases; namely, (1) a “With Experiment” case that mimics the operations that actually occurred and (2) a “Without Experiment” case that simulates operations under the assumption that the HFE did not occur. The “With Experiment” case mimics operations during the HFE and the entire month the HFE occurred. It complies with LTEMP hourly and daily operating criteria. The “Without Experiment” case assumes that the HFE did not occur. The monthly water release volume is assumed to be identical under both cases. The Colorado River Storage Project Python-based model (CRiSPPy) model was the main modeling tool used to simulate the dispatch of the GCD hydropower plant and associated water releases from Lake Powell. In the modeling process, the research team used extensive data sets and historical information on SLCA/IP power plant characteristics, hydrologic conditions, and WAPA’s power purchases and sales prices. In addition to estimating the financial impact of the HFE, the team used the CRiSPPy model to gain insights into the interplay among ROD operating criteria, exceptions made to criteria to accommodate the HFE, and WAPA operating practices.
The Glen Canyon Dam (GCD) is a Colorado River Storage Project (CRSP) power resource that is a component of the Salt Lake City Area Integrated Projects (SLCA/IP). The 2016 record of decision (ROD) for the GCD long-term experimental and management plan (LTEMP) final Environmental Impact Statement (EIS) specifies criteria for GCD monthly water releases, daily and hourly operating limits, and experimental releases. This report presents a financial analysis of the Smallmouth bass (Micropterus dolomieu) (SMB) flows implemented at GCD during Water Year (WY) 2024. These bypass flows were introduced by the U.S. Bureau of Reclamation (USBR) as an emergency response to the growing threat posed by invasive SMB in the Colorado River ecosystem downstream of the dam. SMB are a non-native predatory species that pose a significant threat to native fish populations, including the endangered humpback chub (Gila cypha). The thermal regime below GCD, typically cold due to hypolimnetic releases from Lake Powell, has historically served as a thermal barrier limiting SMB establishment. However, persistently low reservoir levels in recent years have reduced stratification in Lake Powell, allowing warmer water to be released downstream. This has enabled SMB to spawn successfully below the dam, prompting urgent ecological concerns. To mitigate the risk of SMB proliferation, the USBR implemented a series of bypass flows in WY 2024. Drawn from a lower elevation than the penstocks, the bypass structures released cooler water downstream. These short-duration bypass flows aimed to keep temperatures cool enough to prevent SMB from spawning, thereby reducing the ecological threat posed by this invasive species. Although motivated by ecological objectives, these bypass flows came with financial tradeoffs. Releasing water through the bypass structures instead of the turbines at GCD reduced hydropower generation, resulting in a significantly lower financial position for Western Area Power Administration (WAPA), which is responsible for marketing the electricity produced by the GCD Powerplant. This report analyzes the financial impact of the SMB flows implemented from July to November 2024. These experimental releases led to an estimated financial cost of approximately $18.9 million, primarily driven by the substantial volume of water diverted through the bypass structures. This study applies an integrated set of tools to estimate WAPA financial impacts by simulating GCD under two types of cases; namely, (1) a “With Experiment” case that mimics the water operations that actually occurred, including the SMB bypass flows, and (2) a “Without Experiment” case that simulates operations under the assumption that the SMB flows did not occur. Both cases comply with LTEMP hourly and daily operating criteria, and the monthly water release volumes are assumed to be identical under both cases. The Colorado River Storage Project Python-based model (CRiSPPy) model was the main modeling tool used to simulate the dispatch of the GCD hydropower plant and associated water releases from Lake Powell. In the modeling process, the research team used extensive data sets and historical information on SLCA/IP power plant characteristics, hydrologic conditions, and WAPA’s power purchases and sales prices.
The Glen Canyon Dam (GCD) is a Colorado River Storage Project (CRSP) power resource that is a component of the Salt Lake City Area Integrated Projects (SLCA/IP). The 2016 record of decision (ROD) for the GCD long-term experimental and management plan (LTEMP) final Environmental Impact Statement (EIS) specifies criteria for GCD monthly water releases, daily and hourly operating limits, and experimental releases.
Managing IT equipment by hand is prone to errors and delays, severely impacting operational continuity and productivity. Manual inventory systems often result in time delays, inconsistent record-keeping, equipment shortages, and increased workloads for IT staff. At Savannah River National Laboratory (SRNL), my internship focused on creating an automated inventory management solution using Microsoft Power Automate and SharePoint Lists. This solution seamlessly integrates with the existing Microsoft 365 infrastructure, thus eliminating the need for additional software purchases or dedicated server space. By providing real-time updates and reducing manual data entry, the new system ensures a more reliable and maintainable approach to IT asset management.
Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.