Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data repository”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Blockchain for Fault-Tolerant Grid Operations Version 2.0

This report explores the potential of distributed ledger technology (DLT) as a transformative tool to enhance fault-tolerant operations in electrical distribution systems. Leveraging DLT's core attributes, including an immutable decentralized ledger, distributed consensus mechanisms, and state replication capabilities, this study focuses on three critical use cases. A central aspect of this research centers on the utilization of a consensus-driven ledger, providing actors within the system, such as distributed resources, with access to a reliable data repository. This empowers these actors to collaborate effectively and make informed decisions, all securely recorded on the blockchain. The first use case concentrates on data configuration, utilizing mathematical criteria---particularly, the chi-squared test for gross error detection---to identify trustworthy sensors for advanced decision-making. Building upon this foundation of trust, the second use case, topology identification, accurately determines circuit breaker states, unveiling the distribution network's topology. Ultimately, the third use case leverages this trust to execute switching actions, reconfiguring feeders and restoring power to disconnected customers after fault events. The concept of trust serves as a cornerstone in this approach, marking a departure from traditional fault location, isolation, and service restoration (FLISR) methods. Additionally, the blockchain-based architecture introduces decentralization, empowering disconnected areas to make autonomous decisions, even when communication with a central control center is disrupted. The primary contributions of this report are twofold: (1) a novel approach for evaluating distribution system voltage areas while preserving data ownership and (2) the implementation of interactions between distribution network areas using the actor model. Unlike the previous sequential approach for evaluating the area connection voltages, which required a radial network topology, this study's area model reduction enables a more versatile approach. The area model reduction addresses issues of prolonged data waiting times and multiple points of failure within the previous approach. Notably, the presented evaluation for the reduced network model area connection reveals a significant increase in the differences in voltage magnitudes. Simulation and evaluation of area agents across four distinct cases elucidate the area-level interaction behavior during a fault event. Simulations demonstrate that the proposed distributed FLISR (DFLISR) approach can successfully restore service to an affected area. Varying message delays and message loss probabilities in each simulation case underscore their impacts on restoration times, ranging from 3 min and 32 s to 6 min and 19 s. In contrast, power is not restored in an area in one of our simulation cases.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Model America - data and models of every U.S. building

The 5-year goal of the 'Model America' concept was to generate a model of every building in the United States. This data repository delivers on that goal. Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,714,640 buildings detected in the United States and this dataset contains 122,930,327 (97.8%) buildings which resulted in a successful simulation. Future, annual updates have been proposed that may include additional buildings, data improvements, or other algorithmic enhancements. This dataset of 122.9 million buildings includes: Models (state_county.zip) - OpenStudio (v3.1.0) and EnergyPlus (v9.4) building energy models. Please note that the download requires the free Globus Connect Personal (https://www.globus.org/globus-connect-personal); Each model has approximately 3,000 building input descriptors that can be extracted. Please see the EnergyPlus(v9.4) 2,784-page Input/Output Reference Guide (https://energyplus.net/sites/all/modules/custom/nrel_custom/pdfs/pdfs_v9.4.0/InputOutputReference.pdf) for everything that can be retrieved or simulated from these models. These models were derived from the following metadata, which is not included in this dataset: 1. ID - unique building ID 2. County - county name 3. State - state name 4. CZ - ASHRAE Climate Zone designation 5. Clim_Zone - text label of climate zone 6. est_year - estimated year of construction 7. est_commercial - estimated building type (0=residential, 1=commercial) 8. Centroid - building center location in latitude/longitude (from Footprint2D) 9. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 10. Height - building height (meters) 11. Area2D - footprint area (ft2) 12. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 13. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 14. NumFloors - number of floors (above-grade) 15. Area - estimate of total conditioned floor area (ft2) 16. Standard - building vintage. These models are made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy's (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). This research used resources of the Argonne Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC02-06CH11357. Please cite as: New, Joshua R., Adams, Mark, Bass, Brett, Berres, Anne, and Clinton, Nicholas (2021). 'Model America - data and models of every U.S. building. [Data set].' Constellation, doi.ccs.ornl.gov/ui/doi/339, April 14, 2021

24 POWER TRANSMISSION AND DISTRIBUTION↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗

TRAILS Output Files

Overview This data repository contains ZIP files that store compressed versions of the output of running the WaterPaths utility planning and management tool in the DU Re-Evaluation mode (to download the tool, please see this GitHub repository). The tool was used to simulate the six-utility North Carolina Research Triangle problem. Details on the contents of each ZIP file can be seen below. Data details Temporal range: Weekly data for 2,344 weeks from 2015 to 2060 (45 years). Spatial range: Six water utilities in the North Carolina Research Triangle region (0: Chapel Hil/OWASA, 1: Durham, 2: Cary, 3: Raleigh, 4: Pittsboro, and 5: Chatham) File types: CSV and OUT Different solutions available The solution numbers correspond to the different pathway strategies (henceforth referred to as "solutions") discussed in paper's main and supporting text (abstract and link to the paper here). They are as follows: Sol92: The Durham-focused pathway strategy Sol132: The Raleigh-focused pathway strategy Sol140: The regionally-robust pathway strategy Objectives files These files can be accessed by unzipping solXX_objectives_pathways.zip that contains 1,000 Objectives_RDMXX_solsXX_to_XX.csv files. Each CSV file will consist of a row representing all the objective values for that specific solution, while every six columns represents the reliability, restriction frequency, infrastructure net present value ($ mil), peak financial cost, worst-case cost, and unit cost ($ per MG; in that order) for each of the six utilities. There will be 1,000 such files, denoting the performance of the six utilities across the 1,000 deeply uncertain states of the world (DU SOWs). Pathway files These files can be accessed by unzipping solXX_objectives_pathways.zip that contains 1,000 Pathways_sXX_RDMXX.out file. Each OUT corresponds to the set of infrastructure being triggered in a specific DU SOW, and each file will have the name file will consist of four tab-delimited columns that are described as follows: Realization: The realization in which an infrastructure options being triggered utility: The utility currently triggering infrastructure week: The week in which a specific infrastructure option is being triggered infra.: The infrastructure option being triggered If the OUT file contains only the header line, no infrastructure was triggered for that specific DU SOW. Policies files These files can be obtained by unzipping Policies.zip. Each of the 1,000 CSV files within the unzipped folder will contain weekly water use restriction policies for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: 0rest_m: restriction multiplier for utility 0 (values between 0 and 1) 1rest_m: restriction multiplier for utility 1 (values between 0 and 1) 2rest_m: restriction multiplier for utility 2 (values between 0 and 1) 3rest_m: restriction multiplier for utility 3 (values between 0 and 1) 4rest_m: restriction multiplier for utility 4 (values between 0 and 1) 5rest_m: restriction multiplier for utility 5 (values between 0 and 1) 0transf: transfer volume for utility 0 (in MGD) 1transf: transfer volume for utility 1 (in MGD) 2transf: transfer volume for utility 2 (in MGD) 3transf: transfer volume for utility 3 (in MGD) 4transf: transfer volume for utility 4 (in MGD) 5transf: transfer volume for utility 5 (in MGD) Water Sources files These files can be obtained by unzipping WaterSources_subset.zip. Each of the 100 CSV files within the unzipped folder will contain weekly state variables at each water source for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: Xvolume: available water volume from source X (in MGD) Xs_area: surface area of source X (in ACF) Xdemand: demand drawn from a water source from source X (in MGD) Xup_spill: upstream spillage from source X (in MGD) Xww_inflow: wastewater inflow from source X (in MGD) Xcatch_inflow: upstream catchment inflow to source X (in MGD) Xevap: evaporation multiplier for source X (values between 0 and 1) Xds_spill: downstream spillage from source X (in MGD) X_Y_alloc_cap: the allocated capacity from source X to utility Y (values between 0 and 1) X_Y_alloc_dem: the allocated demand from source X to utility Y (values between 0 and 1) Xtrmt_alloc_Y: the allocated treatment capacity from source X to utility Y (values between 0 and 1) Utilities files These files can be obtained by unzipping Utilities_subset.zip. Each of the 100 CSV files within the unzipped folder will contain weekly state variables at each utility for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: Xst_vol: total available storage volume of utility X (in MG) Xcapacity: total storage capacity of utility X (in MG) Xnet_inf: : net inflow for all storage infrastructure for utility X (in MGD) Xst_rof: short term ROF for utility X (values between 0 and 1) Xst_stor_rof: short-term storage ROF for utility X (values between 0 and 1) Xst_trmt_rof: short-term treatment ROF for utility X (values between 0 and 1) Xlt_rof: long-term ROF for utility X (values between 0 and 1) Xlt_stor_rof: long-term storage ROF for utility X (values between 0 and 1) Xlt_trmt_rof: long-term treatment ROF for utility X (values between 0 and 1) Xrest_demand: restricted demand for utility X (in MGD) Xunrest_demand: unrestricted demand for utility X (in MGD) Xunfulf_demand: unfulfilled demand for utility X (in MGD) Xwastewater: wastewater return for utility X (in MGD) Xtreat_capacity: total treatment capacity for utility X (in MG) Xcont_fund: reserve (contingency) fund balance for utility X Xins_pout: insurance payout for utility X (% annual volumetric revenue) Xins_price: insurance price for utility X (% annual volumetric revenue) Xinfra_npv: infrastructure net present value for utility ($mil) Xst_vol: total available storage volume of utility X (in MG) Xdebt_serv: debt service for utility X (usually once per year if the infrastructure is triggered; % annual volumetric revenue) Xstor_vol: total stored volume (in MGD) Xobs_ann_dem: observed annual demand for utility X (in MGD) Xproj_dem: projected annual demand for utility X (in MGD) Xpv_debt_serv: present value of debt service payments for utility X (% annual volumetric revenue) Xgross_rev: gross revenue for utility X ($mil) Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

Artificial Intelligence↗

Creation of a prototype biomimetic fish to better understand impact trauma caused by hydropower turbine blade strikes

Biomimetic model organisms could be useful surrogates for live animals in many applications if the models have sufficient biofidelity. One such application is for use in field and laboratory tests of fish mortality associated with passage through hydropower turbines. Laboratory trials suggest that blade strikes are especially injurious and often causes mortality when fish are struck by thinner blades moving at higher velocities. Dose-response relationships have been created from these data, but the exact relationship between fish mortality and the actual forces enacted on fish during simulated blade strike testing remains unknown. Here, we describe the methods used to create a prototype biomimetic model fish composed of ballistic gelatin and covered with a surrogate skin to better approximate the biomechanical properties of a fish body. Frozen fish were scanned with high-fidelity laser scanners, and a 3D-printed, reusable mold was created from which to cast our gelatin model. Computed tomography scan data, imaged directly or taken from online data repositories, were also successfully used to create CAD models for use in additive manufacturing of molds. One 3-axis accelerometer was embedded into the gelatin to compare accelerometer data to dose-response data from previous laboratory research on live fish. The resulting model ( i.e. , Gelfish) had a statistically indistinguishable tissue durometer to that of real fish tissue and preliminary blade strike impact testing suggested its overall flexibility was similar to that of live fish. Gelfish was designed with biofidelity as its guiding principle and our results suggest initial experimentation was successful. Future research will include replication of initial Gelfish test results, quantitative measurement of model flexibility relative to real fish, and inclusion of surrogate skeletal structures to enhance biofidelity. Use of more sophisticated sensors would also better quantify the physical forces of blade strike impact and help determine how said forces correlate with rates of mortality observed during tests on live fish.

36 MATERIALS SCIENCE↗

Microbiome data management in action workshop: Atlanta, GA, USA, June 12–13, 2024

Microbiome research is revolutionizing human and environmental health, but the value and reuse of microbiome data are significantly hampered by the limited development and adoption of data standards. While several ongoing efforts are aimed at improving microbiome data management, significant gaps still remain in terms of defining and promoting adoption of consensus standards for these datasets. The Strengthening the Organization and Reporting of Microbiome Studies (STORMS) guidelines for human microbiome research have been endorsed and successfully utilized by many research organizations, publishers, and funding agencies, and have been recognized as a consensus community standard. No equivalent effort has occurred for environmental, synthetic, and non-human host-associated microbiomes. To address this growing need within the microbiome research community, we convened the Microbiome Data Management in Action Workshop (June 12–13, 2024, in Atlanta, GA, USA), to bring together key decision makers in microbiome science including researchers, publishers, funders, and data repositories. The 50 attendees, representing the diverse and interdisciplinary nature of microbiome research, discussed recent progress and challenges, and brainstormed actionable recommendations and paths forward for coordinated environmental microbiome data management and the modifications necessary for the STORMS guidelines to be applied to environmental, non-human host, and synthetic microbiomes. The outcomes of this workshop will form the basis of a formalized data management roadmap to be implemented across the field. These best practices will drive scientific innovation now and in years to come as these data continue to be used not only in targeted reanalyses but in large-scale models and machine learning efforts.

54 ENVIRONMENTAL SCIENCES↗

Model America: Data and Models for every U.S. Building

The 5-year goal of the “Model America” concept was to generate a model of every building in the United States. This data repository delivers on that goal with "Model America v1". Oak Ridge National Laboratory (ORNL) has developed the Automatic Building Energy Modeling (AutoBEM) software suite to process multiple types of data, extract building-specific descriptors, generate building energy models, and simulate them on High Performance Computing (HPC) resources. For more information, see AutoBEM-related publications (bit.ly/AutoBEM). There were 125,715,609 buildings detected in the United States. Of this number, 122,146,671 (97.2%) buildings resulted in a successful generation and simulation of a building energy model. This dataset includes the full 125 million buildings. Future updates may include additional buildings, data improvements, or other algorithmic model enhancements in "Model America v2". This dataset contains OSM and IDF zip files for every U.S. county. Each zip file contains the generated buildings from that county. The .csv input data contains the following data fields: 1. ID - the Unique Building Identifier (UBID), generated using the Pacific Northwest National Laboratory (PNNL) BuildingID framework 2. Centroid - building center location in latitude/longitude (from Footprint2D) 3. Footprint2D - building polygon of 2D footprint (lat1/lon1_lat2/lon2_...) 4. State_abbr - state name 5. Area - estimate of total conditioned floor area (ft2) 6. Area2D - footprint area (ft2) 7. Height - building height (ft) 8. NumFloors - number of floors (above-grade) 9. WWR_surfaces - percent of each facade (pair of points from Footprint2D) covered by fenestration/windows (average 14.5% for residential, 40% for commercial buildings) 10. CZ - ASHRAE Climate Zone designation 11. BuildingType - DOE prototype building designation (IECC=residential) as implemented by OpenStudio-standards 12. Standard - building vintage This data is made free and openly available in hopes of stimulating any simulation-informed use case. Data is provided as-is with no warranties, express or implied, regarding fitness for a particular purpose. We wish to thank our sponsors which include Oak Ridge National Laboratory (ORNL) Laboratory Directed Research and Development (LDRD), U.S. Dept. of Energy’s (DOE) Building Technologies Office (BTO), Office of Electricity (OE), Biological and Environmental Research (BER), and National Nuclear Security Administration (NNSA). Update (September 23, 2025): We corrected the ID field in all state-level.csv input files to ensure one-to-one consistency with the corresponding .osm and .idf output files. The schema and file structure are unchanged; only the values in the ID column were modified. No files were added or removed, and the .zip bundles (containing .osm / .idf) are unchanged. The corrected .csv inputs were re-extracted in March 2025 from the original data generated ~ 2021 (Theta supercomputer runs), and published here to align input IDs with model outputs. Update (September 6, 2026): The Model America dataset was updated to replace the previous building ID field with the Unique Building Identifier (UBID), using the Pacific Northwest National Laboratory (PNNL) BuildingID framework. UBIDs provide standardized, location-based identifiers for individual building footprints and improve interoperability with other building and geospatial datasets. The data files containing the previous building identifiers were updated to include UBIDs. This update standardizes building identification; the underlying Model America building characteristics and energy simulation results were not recomputed as part of this update.

54 ENVIRONMENTAL SCIENCES↗

RCSB Protein Data Bank tools for 3D structure-guided cancer research: human papillomavirus (HPV) case study

Abstract Atomic-level three-dimensional (3D) structure data for biological macromolecules often prove critical to dissecting and understanding the precise mechanisms of action of cancer-related proteins and their diverse roles in oncogenic transformation, proliferation, and metastasis. They are also used extensively to identify potentially druggable targets and facilitate discovery and development of both small-molecule and biologic drugs that are today benefiting individuals diagnosed with cancer around the world. 3D structures of biomolecules (including proteins, DNA, RNA, and their complexes with one another, drugs, and other small molecules) are freely distributed by the open-access Protein Data Bank (PDB). This global data repository is used by millions of scientists and educators working in the areas of drug discovery, vaccine design, and biomedical and biotechnology research. The US Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) provides an integrated portal to the PDB archive that streamlines access for millions of worldwide PDB data consumers worldwide. Herein, we review online resources made available free of charge by the RCSB PDB to basic and applied researchers, healthcare providers, educators and their students, patients and their families, and the curious public. We exemplify the value of understanding cancer-related proteins in 3D with a case study focused on human papillomavirus.

60 APPLIED LIFE SCIENCES↗

Automated Metadata Extraction: Challenges and Opportunities

Proper application of the FAIR data principles is what separates a vibrant data ecosystem, in which research data are frequently shared and reused, from a lifeless data graveyard. Automated metadata extraction systems have been proposed as a means of bolstering the findability, interoperability, and reusabil- ity of data repositories with little or no human intervention. These extraction systems mine metadata by crawling a repository and applying lightweight extractors that, for various types of file (e.g., image, CSV file), extract or synthesize relevant attributes. In practice, however, the automated creation of generally useful metadata is fraught with challenges. Data consumers may have different perspectives as to what metadata representations are useful, the standards for recording metadata tend to change over time, and the software model for processing updates can introduce unnecessary human and computational effort. Thus, generalizing extraction for a broad audience of data consumers is a difficult and relatively unsolved problem.In this work, we explore these challenges faced by extraction systems in the context of constructing our own extraction system for science data. We first define the metadata extraction problem and provide context to the issues faced in generalizing metadata. Additionally, we identify potential research directions to help alleviate many of these challenges for all automated extraction systems. Ultimately, this work represents a first step in designing ubiquitous metadata extraction systems that can maximize the value of research data while minimizing the human efforts required in doing so.

Skluzacek, Tyler↗

Identification and Characterization of ten Escherichia coli Strains Encoding Novel Shiga Toxin 2 Subtypes, Stx2n as Well as Stx2j, Stx2m, and Stx2o, in the United States

The sharing of genome sequences in online data repositories allows for large scale analyses of specific genes or gene families. This can result in the detection of novel gene subtypes as well as the development of improved detection methods. Here, we used publicly available WGS data to detect a novel Stx subtype, Stx2n in two clinical E. coli strains isolated in the USA. During this process, additional Stx2 subtypes were detected; six Stx2j, one Stx2m strain, and one Stx2o, were all analyzed for variability from the originally described subtypes. Complete genome sequences were assembled from short- or long-read sequencing and analyzed for serotype, and ST types. The WGS data from Stx2n- and Stx2o-producing STEC strains were further analyzed for virulence genes pro-phage analysis and phage insertion sites. Nucleotide and amino acid maximum parsimony trees showed expected clustering of the previously described subtypes and a clear separation of the novel Stx2n subtype. WGS data were used to design OMNI PCR primers for the detection of all known stx1 (283 bp amplicon), stx2 (400 bp amplicon), intimin encoded by eae (221 bp amplicon), and stx2f (438 bp amplicon) subtypes. These primers were tested in three different laboratories, using standard reference strains. An analysis of the complete genome sequence showed variability in serogroup, virulence genes, and ST type, and Stx2 pro-phages showed variability in size, gene composition, and phage insertion sites. The strains with Stx2j, Stx2m, Stx2n, and Stx2o showed toxicity to Vero cells. Stx2j carrying strain, 2012C-4221, was induced when grown with sub-inhibitory concentrations of ciprofloxacin, and toxicity was detected. Taken together, these data highlight the need to reinforce genomic surveillance to identify the emergence of potential new Stx2 or Stx1 variants. The importance of this surveillance has a paramount impact on public health. Per our description in this study, we suggest that 2017C-4317 be designated as the Stx2n type-strain.

59 BASIC BIOLOGICAL SCIENCES↗

Broadband Multi-wavelength Properties of M87 during the 2017 Event Horizon Telescope Campaign

In 2017, the Event Horizon Telescope (EHT) Collaboration succeeded in capturing the first direct image of the center of the M87 galaxy. The asymmetric ring morphology and size are consistent with theoretical expectations for a weakly accreting supermassive black hole of mass ∼6.5 × 109 M ⊙. The EHTC also partnered with several international facilities in space and on the ground, to arrange an extensive, quasi-simultaneous multi-wavelength campaign. This Letter presents the results and analysis of this campaign, as well as the multi-wavelength data as a legacy data repository. We captured M87 in a historically low state, and the core flux dominates over HST-1 at high energies, making it possible to combine core flux constraints with the more spatially precise very long baseline interferometry data. We present the most complete simultaneous multi-wavelength spectrum of the active nucleus to date, and discuss the complexity and caveats of combining data from different spatial scales into one broadband spectrum. We apply two heuristic, isotropic leptonic single-zone models to provide insight into the basic source properties, but conclude that a structured jet is necessary to explain M87’s spectrum. We can exclude that the simultaneous γ-ray emission is produced via inverse Compton emission in the same region producing the EHT mm-band emission, and further conclude that the γ-rays can only be produced in the inner jets (inward of HST-1) if there are strongly particle-dominated regions. Direct synchrotron emission from accelerated protons and secondaries cannot yet be excluded.

79 ASTRONOMY AND ASTROPHYSICS↗

Nuclear Physics Network Requirements Review Report

The Energy Sciences Network (ESnet) is the Office of Science’s high-performance network user facility, delivering highly reliable data transport capabilities optimized for the requirements of data-intensive science. In essence, ESnet is the circulatory system that enables the U.S. Department of Energy (DOE) science mission by connecting each and every DOE lab and its user facilities. ESnet is funded and stewarded by the Advanced Scientific Computing Research (ASCR) Program and managed and operated by the Scientific Networking Division at Lawrence Berkeley National Laboratory (LBNL). ESnet is widely regarded as a global leader in the research and education networking community. ESnet connects DOE national laboratories, user facilities, and major experiments so scientists can use remote instruments and computing resources as well as share data with collaborators, transfer large data sets, and access distributed data repositories. While ESnet provides network connectivity, it cannot be characterized as an internet service provider as it is specifically built to provide a range of network services that are tailored to meet the unique requirements of DOE’s data-intensive science.

97 MATHEMATICS AND COMPUTING↗

Visual Systems Mapping to Define and Compare Woody Biomass LCAs for Sustainable Systems

The challenge addressed in this research centres on the need to choose between several biomass sources and energy production processes, while supporting rural economies and resilience of forest systems. A key barrier to effective decision-making for strategies using biomass is the lack of standardized and transparent life cycle assessment (LCA) baselines. These baselines are critical for assessing the impacts of biomass strategies but often vary due to regional factors and chosen simplifying assumptions of the LCAs. However, omitting key variables can mean the LCA omits key feedback and balancing loops relevant to fully assessing impacts of the change or test scenario. To address these complexities, this project employs a systems engineering approach: visual systems mapping. This technique is used to define the boundaries and dynamic behaviours of LCA baselines, enhancing transparency. By examining five literature sources and their documented baseline scenarios, the systems mapping case-studies demonstrates an approach to documenting and archiving these baselines. Recommendations are that visual systems mapping should be used to document key assumptions, such as baselines, of LCAs. Further, where possible open data repositories should hold key information about LCA baselines and reproducible workflows (e.g., using open-source tools) should be used to improve transparency and comparability in LCAs. Given the consensus within the broader scientific community on the importance of replicable data practices, this research reinforces the need for standardized frameworks and systems engineering tools in LCAs. This research demonstrates a pathway to more transparent, standardized, and comparable LCAs, that may bolster decisions for biomass systems.

Davis, Maggie [ORNL] (ORCID:0000000181319328)↗

Data for Grogan et al. "Bringing Hydrologic Realism to Water Markets"

This data set provides model output and post-processing files required to reproduce the results, tables, and figures in the paper "Bringing Hydrologic Realism to Water Markets" by Grogan et al. (in review). Other input data used in this study includes: Lisk, M., Grogan, D., Zuidema, S., Caccese, R., Peklak, D., Zheng, J., Fisher-Vanden, K., Lammers, R., Olmstead, S., & Fowler, L. (2023). Harmonized Database of Western U.S. Water Rights (HarDWR) (Version v1) [Data set]. MSD-LIVE Data Repository. https://doi.org/10.57931/2205619 Two models were used in this study: (1) The University of New Hampshire Water Balance Model WBM, and (2) a Water Market Model. Market model code and model output post-processing code that make use of these data can be found here Model output files are: 1. WBM output files: scenario[x]_wbm_output.zip Where [x] is one of 1, 2, 2a, 3, and 3a Each zipped directory contains 7 gridded NetCDF files, each reporting the 10-year annual average value of a given variable, in units of average mm/day: File Name: wbm_indUseGross_yc.nc; Description: Water withdrawals by industry (part of the urban sector) File Name: wbm_domUseGross_yc.nc; Description: Water withdrawals by the domestic sector (part of the urban sector) File Name: wbm_irrigationGross_yc.nc; Description: Water withdrawals for agriculture File Name: wbm_irrigationExtra_yc.nc; Description: Water withdrawals from unsustainable groundwater for agriculture File Name: wbm_indUseEvap_yc.nc; Description: Consumptive water use by industry File Name: wbm_domUseEvap_yc.nc; Description: Consumptive water use by the domestic sector File Name: wbm_irrigationNet_yc.nc; Description: Consumptive water use by agriculture The file full_cell_area.nc gives the area of each grid cell in km2, which is used for converting water depth to water volume. 2. Water market model output & post processing output Folder: marketTrdSummaries/ Description: Files in this folder are used as input to code 1_WelfareCalculation_actual_trades.R. They summarize historical water right trade transactions in each state. File Name: welfare_gain_by_state_sector.csv; Description: Welfare gains by state and sector, as shown in Figure 3F. Used in code Figure3.R and produced (as a .xlsx file) by code 2_DemandCurves_simulated_trades.R File Name: welfare_data_actual.rdata; Description: welfare gains by WMA from actual historical trades, as shown in Figure 3A. This data is the output of code 1_WelfareCalculation_actual_trades.R File Name: welfare_summary_simulated.xlsx; Description: Welfare gains by state as simulated by the market model in Scenario 1. Produced by code 2_DemandCurves_simulated_trades.R, and used in code 4_WelfareCalculation.R. File Name: welfare_summary_cutoffs.xlsx; Description: Welfare gains by state as simulated by the market model in Scenario 2. Produced by code 3_DemandCurves_simulated_trades_cutoffs.R, and used in code 4_WelfareCalculation.R. File Name: welfare_summary_cutoffs_SGMS.xlsx; Description: Welfare gains by state as simulated by the market model in Scenario 2a. Produced by code 3_DemandCurves_simulated_trades_cutoffs.R, and used in code 4_WelfareCalculation.R. File Name: welfare_data_actual.rdata; Description: Spatial data, actual historical welfare gains by WMA as shown in Figure 3A. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. File Name: welfare_data_simulated.rdata; Description: Spatial data, simulated Scenario 1 welfare gains by WMA as shown in Figure 3B. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. File Name: welfare_data_simulated_cutoffs.rdata; Description: Spatial data, simulated Scenario 2 welfare gains by WMA. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. File Name: welfare_data_simulated_cutoffs_SGMA.rdata; Description: Spatial data, simulated Scenario 2a welfare gains by WMA. Produced by code 4_WelfareCalculation.R and used by code Figure3.R. Additional files are provided for efficient reproduction of tables and figures. These include: File Name: wma_thresold_dates_Scenario2(a).csv; Description: Wet vs. paper right threshold dates for each WMA. Shown in Figure 2A,B. Produced and used by code calculate_thresolds_Figure2.R File Name: WWRTradeBounds (directory); Description: Trade boundary shapefile required to reproduce Figure 3A-D. Used in code Figure3.R File Name: welfare_region_totals.csv; Description: Welfare gains for the entire study region, as shown in Figure 3E. Used in code Figure3.R File Name: Welfare_gain_by_state_sector.csv; Description: Welfare gains by state and sector, as shown in Figure 3F. Used in code Figure3.R and produced (as a .xlsx file) by code 2_DemandCurves_simulated_trades.R File Name: WECC_MERIT_5min_v3b_mask.nc; Description: Gridded file that identified which land grid cells are in the WBM model domain, used for processing in code Figure4.py File Name: Table_1.csv; Description: All data in Table 1, reproducible from WBM output files using code table_1.R

Economics↗

A comparative study on deep learning models for condition monitoring of advanced reactor piping systems

Advanced nuclear reactors offer innovative applications due to their portability, reliability, resiliency, and high capacity factors. To operate them on a wider scale, reducing maintenance life-cycle costs while ensuring their integrity is essential. Autonomous operations in advanced nuclear reactors using augmented Digital Twin (DT) technology can serve as a cost-effective solution by increasing awareness about the system’s health. A key component of nuclear DT frameworks is the condition monitoring of safety systems, such as piping-equipment systems, which involves acquiring and monitoring the plant’s sensor data. Here, this research proposes a condition monitoring methodology utilizing deep learning algorithms, such as multilayer perceptions (MLP) and convolutional neural networks (CNNs), to detect degradation and its severity in nuclear piping-equipment systems. Sensor signals are processed to obtain the power spectral density and the Short-Time Fourier transform, and feature extraction methodologies are proposed to develop degradation-sensitive data repositories. The performance of MLP, one-dimensional (1D) CNN, and 2D CNN within the proposed condition monitoring framework is compared using a finite element model of a 3D piping system subjected to seismic loads as the application case study. Various approaches, such as dropout, k-Fold validation, regularization, and early stopping of training the network, are investigated to avoid overfitting the models to the input sensor data. The predictive capability and computational capacity of the deep learning algorithms are also compared to detect degradation in the Z-pipe system of the Experimental Breeder Reactor II (EBRII). The Z-pipe system is subjected to harmonic excitations that represent normal operating loads, such as pump-induced vibrations. The findings of the study indicate that the proposed artificial intelligence (AI)-driven condition monitoring framework demonstrates superior prediction accuracies with a 2D CNN, whereas the MLP exhibits higher computational efficiency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Hourly Electricity Demand Profiles for Each County in the Contiguous United States

This dataset provides estimated hourly electricity demand for each county in the contiguous United States from 2016-2023. The demand profiles represent the sum of two components: (1) Weighted averages of reported hourly demand profiles for North American Electric Reliability Corporation balancing authority (BA) regions and subregions, scaled to match annual estimates of county-level retail sales and direct use of electricity and weighted by the estimated percentage of county load served by each BA region or subregion. (2) Weighted averages of modeled hourly, county- and sector-level distributed photovoltaic (DPV) capacity factor profiles, scaled to match annual estimates of on-site consumption of DPV-generated electricity for each county and weighted by the percentage of consumption attributable to each sector Annual county-level retail sales are estimated by aggregating utility-reported sales to the state level and allocating the results to counties according to each county's share of state population. Annual county-level direct use is calculated by aggregating power plant-reported direct use values. Annual county-level on-site consumption of DPV-generated electricity is estimated by aggregating utility-reported net metering data to determine the amount of DPV-generated electricity sold back to the grid for each state, subtracting those values from modeled state-level DPV generation estimates, and allocating the results to counties according to each county's share of statewide modeled DPV generation. The open-source Python code used to develop this dataset is available at "Historical Load Data Repository" link below.

14 SOLAR ENERGY↗

DAISY Benchmark Performance Data

This repository contains the underlying data from benchmark experiments for Drifting Acoustic Instrumentation SYstems (DAISYs) in waves and currents described in "Performance of a Drifting Acoustic Instrumentation SYstem (DAISY) for Characterizing Radiated Noise from Marine Energy Converters" (https://link.springer.com/article/10.1007/s40722-024-00358-6). DAISYs consist of a surface expression connected to a hydrophone recording package by a tether. Both elements are instrumented to provide metadata (e.g., position, orientation, and depth). Information about how to build DAISYs is available at https://www.pmec.us/research-projects/daisy. The repository's primary content is three compressed archives (.zip format), each containing multiple MATLAB binary data files (.mat format). A table relating individual data files to figures in the paper, as well as the structure of each file, is included in the repository as a Word document (Data Description MHK-DR.docx). Most of the files contain time series information for a single DAISY deployment (file naming convention: [site]_DAISY_[Drift #].mat) consisting of processed hydrophone data and associated metadata. For a limited number of DAISY deployments, the hydrophone package was replaced with an acoustic Doppler velocimeter (file naming convention: [site]_DAISY_[Drift #]_ADV.mat). Data were collected over several years at three locations: (1) Sequim Bay at Pacific Northwest National Laboratory's Marine & Coastal Research Laboratory (MCRL) in Sequim, WA, the energetic tidal channel in Admiralty Inlet, WA (Admiralty Inlet), and the U.S. Navy's Wave Energy Test Site (WETS) in Kaneohe, HI. Brief descriptions of data files at each location follow. - MCRL - (1) Drift #4 and #16 contrast the performance of a DAISY and a reference hydrophone (icListen HF Reson), respectively, in the quiescent interior of Sequim Bay (September 2020). (2) Drift #152 and #153 are velocity measurements for a drifting acoustic Doppler velocimeter in in the tidally-energetic entrance channel inside a flow shield and exposed to the flow, respectively (January 2018). (3) Two non-standard files are also included: DAISY_data.mat corresponds to a subset of a DAISY drift over an Adaptable Monitoring Package (AMP) and AMP_data.mat corresponds to approximately co-temporal data for a stationary hydrophone on the AMP (February 2019). - Admiralty Inlet - (1) Drift #1-12 correspond to tests with flow shielded DAISYs, unshielded DAISYs, a reference hydrophone, and drifting acoustic Doppler velocimeter with 5, 10, and 15 m tether lengths between surface expression and hydrophone recording package (July 2022). (2) Drift #13-20 correspond to tests of flow shielded DAISYs with three different tether materials (rubber cord, nylon line, and faired nylon line) in lengths of 5, 10, and 15 m (July 2022). - WETS - (1) Drift #30-32 correspond to tests with a heave plate incorporated into the tether (standard configuration for wave sites), rubber cord only, and rubber cord, but with a flow shielded hydrophone (November 2022). (2) Drift #49-58 and Drift #65-68 correspond to measurements around mooring infrastructure at the 60 m berth where time-delay-of-arrival localization was demonstrated for different DAISY arrangements and hydrophone depths (November 2022).

16 TIDAL AND WAVE POWER↗

PSU-BSEC Doppler Lidar Processed Scans

This data repository includes the processed vertically staring measurements (stare files) and the angled scans (profile files) necessary to calculate horizontal winds retrieved by the Pennsylvania State University (PSU) Doppler Lidar as a part of the Baltimore Social-Environmental Collaborative Urban Integrated Field Laboratory (BSEC UIFL). The PSU Doppler Lidar measures aerosol backscatter intensity (m-1 sr-1), signal-to-noise ratio, and radial velocity (i.e., vertical velocity in the case of stare files, units: m/s) at approximately 1 Hz temporal resolution and 30 m spatial resolution. Within the "stare_data" subdirectory there exist example figures of all retrieved data. For further information, please email Nicholas Prince, nec5299@psu.edu.

Air Quality↗