Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data versions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 2 Sensor Data v2-1

This is the version v2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Please see v2-1 TEMPEST L2 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods. * Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021 * TEMPEST 1: June 22, 2022 * TEMPEST 2: June 6-7, 2023 * TEMPEST 3: June 11-13, 2024

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

Validation of the Community Land Model Version 5 over the Contiguous United States (CONUS) using in situ and remote sensing data sets

The Community Land Model (CLM) is an effective tool to simulate the biophysical and biogeochemical processes and their interactions with the atmosphere. Although CLM Version 5 (CLM5) constitutes various updates in these processes, its performance in simulating energy, water and carbon cycles over the Contiguous United States (CONUS) at scales which land surface changes and hydrometeorological and hydroclimatological applications are more locally relevant is yet to be assessed. In this study, we conducted three simulations at 0.125? during 1979-2018 over the CONUS using different configurations of CLM, namely CLM5-biogeochemistry (CLM5BGC), CLM4.5BGC, and CLM5-satellite phenology (CLM5SP). We validated and compared their simulations against multiple remote-sensed and in-situ datasets. Overall, the parametric and structural updates (e.g., carbon cost for nitrogen uptake, variable soil thickness, dry surface layer) in CLM5 improve its ability in capturing terrestrial biogeochemical dynamics. The low evapotranspiration in CLM5BGC is associated with biases in simulating vegetation phenological characteristics rather than soil water limitations. The mismatch between CLM5BGC-simulated peak leaf area index and reference data can be attributed to CLM5BGC's inability in simulating phenology of trees and grasses. The differences between CLM-simulated irrigation and reference estimates can be attributed to differences between processes represented in models and in reality, and uncertainties in input and validation datasets. Evaluation against observations at small catchments suggest that hydrologic parameters needed to be calibrated to improve simulations of runoff, especially subsurface runoff. Additional efforts are needed to incorporate spatially-distributed plant phenology and physiology parameters and regional-specific agricultural management practices (e.g., planting, harvest).

Cheng, Yanyan↗

Second-generation downscaled earth system model data using generative machine learning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. As with the first Sup3rCC version, the data still include temperature, wind speed and direction at multiple heights, pressure, three components of downwelling solar radiation, and relative humidity—all at 4-kilometer (km) hourly resolution over the contiguous United States. This is a 25x spatial enhancement and 24x temporal enhancement of the source 100-km daily-average ESM data. This extension of the Sup3rCC dataset includes data from six ESMs from two shared socioeconomic pathways (SSPs) totaling 400 years of data with multiple future projections of changing meteorological conditions. The scenario selection was based on a structured evaluation of historical ESM skill and comprehensive representation of possible trajectories of future climate change in temperature, humidity, precipitation, solar irradiance, and near-surface wind speeds. The inclusion of multiple future projections is intended to enable users to assess key drivers of un 36 certainty and variability. All data are double-bias corrected, resulting in a product that can be used out-of-the-box for energy system analysis with minimal historical bias. The potential applications of Sup3rCC data extend to various topics in renewable energy resource assessment, energy systems modeling, and grid resilience studies. High-resolution future meteorological projections are critical for evaluating the effects of changing meteorological conditions on renewable energy generation, energy demand, and for optimizing energy storage and grid infrastructure. The 4-km hourly resolution of the downscaled data enables understanding of spatial and temporal variability at the scales necessary for energy system operational planning. In addition, the dataset can support risk assessments by providing detailed information on possible future extreme weather events and long-term meteorological variability at scales relevant to energy infrastructure. By offering an enhanced representation of possible future meteorological conditions, the second-generation Sup3rCC dataset enables more precise modeling of energy resilience and adaptation strategies in response to changing meteorological conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Fission Product Yield Data Adjustment in a Prototype Version of TSURFER

The TSURFER (Tool for Sensitivity/Uncertainty analysis of Response Functionals using Experimental Results) module of Oak Ridge National Laboratory’s (ORNL’s) SCALE code system has been updated to perform nuclear data adjustments for fixed-source irradiation/depletion problems. TSURFER uses a generalized linear least squares (GLLS) approach to consolidate a prior set of measured responses and corresponding calculated values to create the most self-consistent set of nuclear data. Traditionally, TSURFER adjustments have been performed for multigroup nuclear data such as reaction cross sections. In this work, TSURFER is expanded to perform adjustments to independent fission product yields and branching ratios that need equality constraints. To preserve equality constraints after the data adjustment procedure, an updated GLLS formulation includes a new Lagrange multiplier that forces data adjustment to sum to 0 for a given fission yield/branching ratio parent. A test problem illustrates that the newly updated TSURFER module satisfies the required constraint that adjustments for fission yield data sum to 0.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART): Data Format and Preparation Guidance, Version 2.0

The Joint Office of Energy and Transportation maintains the Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART), which provides a centralized hub for submitting electric vehicle (EV) charging infrastructure data directed by the Federal Highway Administration (23 CFR 680.112) EV-ChART will provide a streamlined data submission process and an integrated set of analytic tools, connect to other data sources, and empower data sharing and access across stakeholders, including the public. Any data shared publicly will be aggregated and anonymized to stay in accordance with 23 CFR 680. This EV-ChART Data Format and Preparation Guidance provides a comprehensive overview of the data reporting requirements as authorized under 23 CFR 680.112. The guidance is intended to be used alongside the EV-ChART Data Input Template, which defines the tabular data structure that these data submissions must follow.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

1000 Soils Pilot Dataset, version 8, May 2025

This record hosts data generated by the 1000 Soils Pilot. Data will be updated as more become available. Please see the most recent data upload for current data. A beta visualization tool is available for some data types at https://shinyproxy.emsl.pnnl.gov/app/1000soils. Please submit any suggestions or comments through the 'contact' tab. We are actively working to improve visualizations and value all feedback. Data completed include: Geochemistry, texture, respiration, and enzyme activities FTICR-MS organic matter chemistry Microbial biomass C and N TOC/TDN of water-extractable OM X-ray computed tomography (derived metrics available here, raw data available upon request) Metagenomes; a variety of data formats are available upon request Soil hydraulic properties Data in progress: LC-MS/MS in development, timeline TBD, inquire for status 1000S_processed_BGC_summary.csv contains all available biogeochemical data; microbial biomass C and N; and TOC/TDN of water-extractable OM; and 1000S_Tomography.xslx contains a summary of data generated via X-ray computed tomography. icr_v2_corems2.csv contains FTICR-MS data processed by CoreMS version 2. These data are merged by formula across instrument runs to enable cross-sample comparisons. Technical replicates are merged by retaining peaks present in 2 out of 3 replicates. 1000Soils_Metadata_Site_Mastersheet_v1.csv contains site information. Soil Hydraulics_corrected_02042025.xlsx contains soil hydraulics information. Readme File_v4.xlsx is the readme file. Please contact the MONet project (monet.emsl@pnnl.gov) or Emily Graham (emily.graham@pnnl.gov) with questions. The following file and all raw data are available upon request: icr_by_mass_for_single_sample_analysis_only.csv contains FTICR-MS data processed by CoreMS and is intended for usage in the calculation of biochemical transformations within samples only. These data are not acceptable for cross-sample comparison of masses because they are from multiple instrument runs. For more information, please see: https://www.emsl.pnnl.gov/monet and https://sc-data.emsl.pnnl.gov/monet Acknowledgment: Soil data were provided by the Molecular Observation Network (MONet) at the Environmental Molecular Sciences Laboratory (https://ror.org/04rc0xn13), a DOE Office of Science user facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830. The work (proposal: 10.46936/10.25585/60008970) conducted by the U.S. Department of Energy, Joint Genome Institute (https://ror.org/04xm1d337), a DOE Office of Science user facility, is supported by the Office of Science of the U.S. Department of Energy operated under Contract No. DE-AC02-05CH11231. The Molecular Observation Network (MONet) database is an open, FAIR, and publicly available compilation of the molecular and microstructural properties of soil. Data in the MONet open science database can be found at https://sc-data.emsl.pnnl.gov/.

biogeochemistry↗

BuildingSync® v.2.7.0 (released 9.11.2025) [SWR-18-28]

BuildingSync® is a building data exchange schema to better enable integration between software tools and building data workflows. The schema's original use case was focused on commercial building energy audits; however, several additional use cases have been realized including building energy modeling and more high-level generic building data exchange. Version 2.7.0 adds new elements for file attachment feature and FederalBuilding, and generalizes usage of Optional Elements (e.g. EquipmentCondition, EquipmentID) to all assets/systems. BuildingSync helps streamline the data exchange process, improving the value of the data, minimizing duplication of effort for subsequent building data collection efforts (including audits), and facilitating the achievement of greater energy efficiency. This in done in part by standardizing on (a) reporting audits in an electronic format, (b) tracking proposed, implemented, and discarded energy conservation measures, and (c) storing building characteristics (at multiple levels) for audits, benchmarking, and building energy analysis. BuildingSync has several documents and tools available to help users understand how to best leverage BuildingSync. The list below are only a subset of the resources available. If new resources are discovered, then feel free to create a new pull request with the additions. Generic BuildingSync information is available on the DOE website and the project website. BuildingSync Examples - These examples are kept up to date and show a wide range of implementations. Any new update to BuildingSync is required to pass validation on these example files. BuildingSync Use Case Validator allows for users to determine if their instance complies with a specific use case for BuildingSync by checking if the required elements are implemented in an uploaded instance. An API is also provided for automated integration into other tools. Also, the website contains an easy way to view the entirety of the schema and how elements relate to the Building Exchange Data Exchange Specification. The Validator is open sourced here Use Case TestSuite provides a Python package for easier generation of BuildingSync use cases. BuildingSync use cases depend on the generation of schematron documents, which is time-consuming and difficult to implement well. The TestSuite allows users to define a use case using a more palatable CSV template, which it then turns into a Schematron document. The source code is available here. BuildingSync to OpenStudio/EnergyPlus. The translator is open sourced here. This project will translate a Level 1 (and partial Level 2) ASHRAE Energy Audit to a fully defined OpenStudio and EnergyPlus model. This project is in early Beta testing and any feedback is welcome!

Long, Nicholas [National Renewable Energy Lab. (NR↗

Tidal Resource Data from Sequim Bay Inlet, WA, August 2020

Data from a Nortek Signature1000 deployed on a lander for 14 days in Aug 2020 in the entrance to Sequim Bay, WA. Raw data were processed using the DOLfYN python package and standardized using the ME Data Pipeline python package, tsdat version 0.2.12. Processed data were partitioned into 24 hour increments and saved in the NETCDF file format.

16 TIDAL AND WAVE POWER↗

Data for Grogan et al. "Bringing Hydrologic Realism to Water Markets"

This data set provides model output and post-processing files required to reproduce the results, tables, and figures in the paper "Bringing Hydrologic Realism to Water Markets" by Grogan et al. (in review). Other input data used in this study includes: Lisk, M., Grogan, D., Zuidema, S., Caccese, R., Peklak, D., Zheng, J., Fisher-Vanden, K., Lammers, R., Olmstead, S., & Fowler, L. (2023). Harmonized Database of Western U.S. Water Rights (HarDWR) (Version v1) [Data set]. MSD-LIVE Data Repository. https://doi.org/10.57931/2205619 Two models were used in this study: (1) The University of New Hampshire Water Balance Model WBM, and (2) a Water Market Model. Market model code and model output post-processing code that make use of these data can be found here Model output files are: 1. WBM output files: scenario[x]_wbm_output.zip Where [x] is one of 1, 2, 2a, 3, and 3a Each zipped directory contains 7 gridded NetCDF files, each reporting the 10-year annual average value of a given variable, in units of average mm/day: File Name: wbm_indUseGross_yc.nc; Description: Water withdrawals by industry (part of the urban sector) File Name: wbm_domUseGross_yc.nc; Description: Water withdrawals by the domestic sector (part of the urban sector) File Name: wbm_irrigationGross_yc.nc; Description: Water withdrawals for agriculture File Name: wbm_irrigationExtra_yc.nc; Description: Water withdrawals from unsustainable groundwater for agriculture File Name: wbm_indUseEvap_yc.nc; Description: Consumptive water use by industry File Name: wbm_domUseEvap_yc.nc; Description: Consumptive water use by the domestic sector File Name: wbm_irrigationNet_yc.nc; Description: Consumptive water use by agriculture The file full_cell_area.nc gives the area of each grid cell in km2, which is used for converting water depth to water volume. 2. Water market model output & post processing output Folder: marketTrdSummaries/ Description: Files in this folder are used as input to code 1_WelfareCalculation_actual_trades.R. They summarize historical water right trade transactions in each state. File Name: welfare_gain_by_state_sector.csv; Description: Welfare gains by state and sector, as shown in Figure 3F. Used in code Figure3.R and produced (as a .xlsx file) by code 2_DemandCurves_simulated_trades.R File Name: welfare_data_actual.rdata; Description: welfare gains by WMA from actual historical trades, as shown in Figure 3A. This data is the output of code 1_WelfareCalculation_actual_trades.R File Name: welfare_summary_simulated.xlsx; Description: Welfare gains by state as simulated by the market model in Scenario 1. Produced by code 2_DemandCurves_simulated_trades.R, and used in code 4_WelfareCalculation.R. File Name: welfare_summary_cutoffs.xlsx; Description: Welfare gains by state as simulated by the market model in Scenario 2. Produced by code 3_DemandCurves_simulated_trades_cutoffs.R, and used in code 4_WelfareCalculation.R. File Name: welfare_summary_cutoffs_SGMS.xlsx; Description: Welfare gains by state as simulated by the market model in Scenario 2a. Produced by code 3_DemandCurves_simulated_trades_cutoffs.R, and used in code 4_WelfareCalculation.R. File Name: welfare_data_actual.rdata; Description: Spatial data, actual historical welfare gains by WMA as shown in Figure 3A. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. File Name: welfare_data_simulated.rdata; Description: Spatial data, simulated Scenario 1 welfare gains by WMA as shown in Figure 3B. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. File Name: welfare_data_simulated_cutoffs.rdata; Description: Spatial data, simulated Scenario 2 welfare gains by WMA. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. File Name: welfare_data_simulated_cutoffs_SGMA.rdata; Description: Spatial data, simulated Scenario 2a welfare gains by WMA. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. Additional files are provided for efficient reproduction of tables and figures. These include: File Name: wma_thresold_dates_Scenario2(a).csv; Description: Wet vs. paper right threshold dates for each WMA. Shown in Figure 2A,B. Produced and used by code calculate_thresolds_Figure2.R File Name: WWRTradeBounds (directory); Description: Trade boundary shapefile required to reproduce Figure 3A-D. Used in code Figure3.R File Name: welfare_region_totals.csv; Description: Welfare gains for the entire study region, as shown in Figure 3E. Used in code Figure3.R File Name: Welfare_gain_by_state_sector.csv; Description: Welfare gains by state and sector, as shown in Figure 3F. Used in code Figure3.R and produced (as a .xlsx file) by code 2_DemandCurves_simulated_trades.R File Name: WECC_MERIT_5min_v3b_mask.nc; Description: Gridded file that identified which land grid cells are in the WBM model domain, used for processing in code Figure4.py File Name: Table_1.csv; Description: All data in Table 1, reproducible from WBM output files using code table_1.R

Economics↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

CHESS 2025: Spectrometer orthorectified at-sensor radiance from NEON AOP imaging spectroscopy surveys

This dataset provides Level 1 (L1) orthorectified at-sensor radiance derived from measurements collected by the Imaging Spectrometer-1 (NIS-1) onboard the NEON (National Ecological Observatory Network) Airborne Observation Platform (AOP) for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). NIS-1 captures light reflected from the Earth’s surface in 426 discrete wavelength bands as raw digital numbers (DNs; Level 0). These data are then calibrated to physical units (uW/cm²·sr·nm) following the processing steps described in the NEON Imaging Spectrometer Level 1B Calibrated Radiance Algorithm Theoretical Basis Document (ATBD; Gallery 2022). The data delivered here are the primary inputs for the surface reflectance product in “Custom surface reflectance, shade masks, and equivalent water thickness maps for the Colorado Headwaters Ecological Spectroscopy Study” (Carroll et al. 2026). For intertemporal comparison, the radiance data here are most directly relatable to the v2 radiance data in “NEON AOP Imaging Spectroscopy Survey of Upper East River Colorado Watersheds: Raw-Space Radiance and Observational Variable Dataset” (Goulden et al. 2018), to which the same processing methodology was applied. Together, the radiance and reflectance data enable users to exploit the unique reflection signatures of different surface objects for land cover classification, foliar trait mapping, plant vigor assessment, water content estimation, trace-element identification, and other scientific applications. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. Within each domain, data are delivered by flightline as orthorectified and calibrated hyperspectral rasters in Hierarchical Data Format version 5 (HDF5) format, with radiance values provided in uW/cm²·sr·nm on a fixed, uniform Universal Transverse Mercator (UTM) grid at 1 meter spatial resolution. The radiance rasters include all 426 NIS-1 spectral bands, along with associated quality-assurance (QA) and diagnostic and ancillary layers needed for atmospheric correction workflows. Orthorectified radiance is produced from pushbroom spectrometer observations by applying NEON’s radiometric calibration (including bad pixel masking, dark subtract, dark pedestal shift correction, electronic panel ghost correction, grating ghost correction, deblur correction and flat-fielding) and spectral calibration (using spectral response function band centers and full-width at half-maximum intensity), followed by geolocation and regridding to the fixed grid. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

DFAT: A web-based toolkit for estimating demand flexibility in building-to-grid integration

Demand Flexibility Assessment Tool (DFAT) is an open source web-based tool that estimates the demand flexibility potential of common control strategies in commercial buildings. The toolkit features a demand flexibility estimation tool that contains two calculators, basic and advanced, based on the level of input of customer data. The basic version calculates demand shed metrics for the control strategy “global temperature adjustment” and “cycle on/off compressors” using customer building information, local weather data, and electrical meter data. The advanced version, which uses detailed HVAC equipment data, calculates demand flexibility metrics for control strategies such as static pressure reset, global temperature adjustment, and cycle on/off compressors. In addition to the demand flexibility estimation tool, this toolkit offers a benchmarking tool that helps facility operators, aggregators, and utility resource managers assess demand flexibility opportunities, quantify/verify performance, and compare their performance against that of their peers.

Leong, Michael↗

Westcott g factors extended to arbitrary neutron energy spectra

Westcott 𝑔 factors are used in Neutron Activation Analysis (NAA) and Prompt Gamma-ray Activation Analysis (PGAA) to evaluate the impact of non-1∕𝑣 behavior in the neutron-capture cross sections of certain nuclei on activation product yields. This non-1∕𝑣 behavior arises from the presence of neutron resonances in the neutron- capture cross sections that overlap with the source neutron spectrum at low (< 5 eV) energies. Historically, Westcott 𝑔 factors that have been cataloged for NAA and PGAA applications are the result of calculations that assume a Maxwellian neutron flux distribution with a given temperature. In this work, we use this approach with updated neutron-capture cross sections from the Evaluated Nuclear Data File, version VIII.1 (ENDF/B-VIII.1) to tabulate Westcott 𝑔 factor values for a broad range of Maxwellian distribution temperatures, comparing the results against currently-available 𝑔 factors from International Atomic Energy Agency tables and other sources. Here, it was discovered during this analysis that the use of guided thermal and cold-neutron beams at certain facilities necessitates an approach for evaluating Westcott 𝑔 factors based on arbitrary non-Maxwellian spectra. In this paper, we present an approach for calculating 𝑔 factors with user-specified neutron spectra, and we demonstrate these methods to obtain Westcott 𝑔-factors for guided- and cold-neutron beams at the Budapest Research Reactor and the Forschungsreaktor München II reactor. As part of this work, open-source software has been developed that can be used to perform these calculations for applications in PGAA and NAA experiments.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

KG-Hub—building and exchanging biological knowledge graphs

Knowledge graphs (KGs) are a powerful approach for integrating heterogeneous data and making inferences in biology and many other domains, but a coherent solution for constructing, exchanging, and facilitating the downstream use of KGs is lacking. Here we present KG-Hub, a platform that enables standardized construction, exchange, and reuse of KGs. Features include a simple, modular extract–transform–load pattern for producing graphs compliant with Biolink Model (a high-level data model for standardizing biological data), easy integration of any OBO (Open Biological and Biomedical Ontologies) ontology, cached downloads of upstream data sources, versioned and automatically updated builds with stable URLs, web-browsable storage of KG artifacts on cloud infrastructure, and easy reuse of transformed subgraphs across projects. Current KG-Hub projects span use cases including COVID-19 research, drug repurposing, microbial–environmental interactions, and rare disease research. KG-Hub is equipped with tooling to easily analyze and manipulate KGs. KG-Hub is also tightly integrated with graph machine learning (ML) tools which allow automated graph ML, including node embeddings and training of models for link prediction and node classification.

59 BASIC BIOLOGICAL SCIENCES↗

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗

Local models for scatter estimation and descattering in polyenergetic X-ray tomography

We propose a new modeling approach for scatter estimation and descattering in polyenergetic X-ray computed tomography (CT) based on fitting models to local neighborhoods of a training set. X-ray CT is widely used in medical and industrial applications. X-ray scatter, if not accounted for during reconstruction, creates a loss of contrast in CT reconstructions and introduces severe artifacts including cupping, shading, and streaks. Even when these qualitative artifacts are not apparent, scatter can pose a major obstacle in obtaining quantitatively accurate reconstructions. Our approach to estimating scatter is, first, to generate a training set of 2D radiographs with and without scatter using particle transport simulation software. To estimate scatter for a new radiograph, we adaptively fit a scatter model to a small subset of the training data containing the radiographs most similar to it. We compared local and global (fit on full data sets) versions of several X-ray scatter models, including two from the recent literature, as well as a recent deep learning-based scatter model, in the context of descattering and quantitative density reconstruction of simulated, spherically symmetrical, single-material objects comprising shells of various densities. Our results show that, when applied locally, even simple models provide state-of-the-art descattering, reducing the error in density reconstruction due to scatter by more than half.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Pride of Satellites in the Constellation Leo? Discovery of the Leo VI Milky Way Satellite Ultra-faint Dwarf Galaxy with DELVE Early Data Release 3

Abstract We report the discovery and spectroscopic confirmation of an ultra-faint Milky Way satellite in the constellation of Leo. This system was discovered as a spatial overdensity of resolved stars observed with Dark Energy Camera (DECam) data from an early version of the third data release of the DECam Local Volume Exploration (or DELVE) survey. The low luminosity ( M V = − 3.5 6 − 0.37 + 0.47 ; L V = 230 0 − 700 + 1200 L ⊙ ), large size ( R 1 / 2 = 9 0 − 30 + 30 pc), and large heliocentric distance ( D = 11 1 − 6 + 9 kpc) are all consistent with the population of ultra-faint dwarf galaxies (UFDs). Using Keck/DEIMOS observations of the system, we were able to spectroscopically confirm nine member stars, while measuring a tentative mass-to-light ratio of 70 0 − 500 + 1400 M ⊙ / L ⊙ and a nonzero metallicity dispersion of σ [ Fe / H ] = 0.1 9 − 0.11 + 0.14 , further confirming Leo VI’s identity as a UFD. While the system has a highly elliptical shape, ϵ = 0.5 4 − 0.29 + 0.19 , we do not find any conclusive evidence that it is tidally disrupting. Moreover, despite the apparent on-sky proximity of Leo VI to members of the proposed Crater-Leo infall group, its smaller heliocentric distance and inconsistent position in energy–angular momentum space make it unlikely that Leo VI is part of the proposed infall group.

79 ASTRONOMY AND ASTROPHYSICS↗