Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “document repository”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Data and scripts associated with the manuscript "Organic Molecules are Deterministically Assembled in River Sediments"

This data package is associated with the publication "Organic Molecules are Deterministically Assembled in River Sediments" submitted to Scientific Reports (Stegen et al., 2024). The study applies community ecology methods to dissolved organic matter (DOM) chemistry from variably inundated riverbed sediments to uncover principles governing DOM composition at a reach-scale. This data package documents the workflow used to process and generate the main findings in the manuscript. The R scripts reference the raw, unprocessed Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data from another data package, available on ESS-DIVE at https://data.ess-dive.lbl.gov/view/doi:10.15485/1834208. The scripts then process the raw FTICR-MS data and generate the findings and figures presented in the associated manuscript. In brief, this study demonstrates that DOM assemblages in variably inundated sediments are primarily governed by deterministic variable selection, including sediment moisture effecting the degree of deterministic assembly. See the manuscript for more details pertaining to interpretation and implications of the findings. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/ECA_2020_Sed.This data package is comprised of 6 scripts and 7 folders. The file-level metadata file (file ending in "flmd.csv") lists all files contained in this data package and descriptions for each. The data dictionary (file ending in "dd.csv) describes all tabular data columns and their respective definitions and units. The FTICR_Processing_Scripts produce the outputs found in the "Processed_Data" folder. The remaining scripts (located in the parent directory) produce the outputs found in the following four folders: (1) "MCD_Dendrograms", "MCD_Randomizations", "MCD_bNTI_Outcomes", and "OM_Null_Modeling". The fifth script additionally takes the three comma-separated values (CSV) files found in the parent directory as input ("VGC_texture.csv", "merged_weights.csv", and "ECA2_FTICR_BetaDisp.csv"). The outputs of each of the five scripts serve as the input to the following script, with the final outputs stored in the folder "OM_Null_Modeling".

54 ENVIRONMENTAL SCIENCES↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Mapping and Synthesis of International Biomass Supply Assessments

This report, Mapping and Synthesis of International Biomass Supply Assessments (or Global Biomass Resource Assessment) is the first step in a long-term process to assemble data from around the globe into a virtual repository that can be updated and provide user-friendly access to the data. The Clean Energy Ministerial (CEM) Biofuture Platform Initiative recommended that research be completed to “address the need for internationally accepted benchmarks quantifying sustainable biomass feedstock supplies.” To act upon the CEM Biofuture recommendation, in 2024, the U.S. Department of Energy (DOE) commissioned Oak Ridge National Laboratory (ORNL) to prepare this report as the primary deliverable for a one-year assignment to assemble data into a citable form that could help resolve the persistent question presented related to bioenergy policy, “Is there enough sustainable biomass?” In response to that query, this report includes information received by August 2024 from national CEM representatives, collaborators, and public sources on current and future sustainable biomass supplies in 62 nations,and subsequently documents (a) the approach used by ORNL to analyze and categorize the information received in a manner that enables aggregation and comparability; and (b) recommendations for next steps and guidelines to help others update and harmonize future assessments of global sustainable biomass supplies.

09 BIOMASS FUELS↗

Applying 3D Geologic Modeling Workflows to the Argillite Reference Case (Rev. 1)

The objective of this short report is to document the application of our 3D geologic modeling workflow to an argillite (shale) host rock. Over the past four years, our team at Los Alamos National Laboratory has developed a geologic modeling workflow that can be applied to generic alluvial basins such as those found in the western United States. In “frontier” or “exploratory” basins where data are sparse, the first steps are to collect, evaluate and integrate available subsurface data into conceptual geologic models. Those models form the basis for constructing the geologic framework model, a 3D geocellular model ideally constrained by seismic and borehole data. To date we have constructed our models using “synthetic” well data derived from conceptual models, without the prospect of validating our workflow using “real” subsurface data. We were tasked to investigate whether our workflow designed for alluvial basin sediments could be applied to other potential repository host rocks. This task also provided the opportunity to work with high-quality subsurface data collected specifically for siting and evaluating a nuclear waste repository. Nagra, the Swiss governmental agency responsible for the disposal of the nation’s radioactive waste, generously provided us with data from two deep boreholes drilled through their argillaceous target formation. The aim of our proof-of-concept demonstration is to evaluate whether geostatistical methods offer a viable approach to property modeling in argillaceous rocks. Nagra provided us with the well data on the condition that we maintain confidentiality with all transferred information and results. Fortunately, Nagra posts numerous technical reports on its public website that describe the subsurface geology in great detail. All of the information and illustrations in this report related to the Swiss repository enterprise are taken from the Nagra public website.

58 GEOSCIENCES↗

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram↗

ESS-DIVE Reporting Format for Amplicon Abundance Table

While standardized sequencing data is available in public repositories and efforts such as MIxS for common sample collection and processing metadata are well established, the lack of common bioinformatic processing metadata has hindered the ability to do large-scale metaanalyses and the potential for data re-use by non-experts such as ecosystem, watershed, or earth system modelers. To address this need for Department of Energy researchers, we have developed an amplicon reporting format which captures both sample preparation and bioinformatic processing metadata and stores processed amplicon data as a paired abundance table and sequencing file to maximize the potential for re-use of these data. To aid in the adoption of accessible and reproducible analysis workflows, this reporting format was developed in concert with amplicon functionality within the Department of Energy’s Systems Biology Knowledgebase (KBase) to ensure common data and metadata requirements and facilitate seamless transfer between these platforms.This dataset contains support documentation for the amplicon reporting format (README.md and instructions.md), templates for both bioinformatic and sequencing metadata (amplicon_bioinformatic_metadata_template_2021_10_03.csv and amplicon_sequencing_metadata_template_2021_10_03.csv), a crosswalk indicating how this reporting format relates to the current MIxS format (ESSDIVE-MIxS_crosswalk.csv), a list of available instrument terms (amplicon_seq_instrument_terms_2021_10_03.csv), a map between QIIME2 parameter settings and metadata fields (amplicon_qiime2_plugin_metadata_map.csv), a data dictionary (amplicon_CSV_dd.csv), and file-level metadata (amplicon_FLMD.csv).

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with: “Burn severity and vegetation type control phosphorus concentration, molecular composition, and mobilization”

This data package is associated with the publication “Burn severity and vegetation type control phosphorus concentration, molecular composition, and mobilization” published in European Geophysical Union - Biogeosciences (Barnes et al. 2025). This study investigates how phosphorus (P) biogeochemistry is altered by burn severity in contrasting types of vegetation chars. This data package documents the workflow used to process and generate the main figures and statistics in the manuscript. The R scripts reference minimally processed P nuclear magnetic resonance (P-NMR) and X-ray absorption near edge structure (P-XANES) data, as well as fully processed data including total elemental composition of the solid chars, total elemental composition of the char leachates (particulate and aqueous phases), and leachate aqueous phase molybdate reactive P concentration. These source data and associated metadata can be found on ESS-DIVE at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1894135 (Grieger et al. 2022; v3). Files and scripts included in this data package finish the processing workflow for P-NMR and P-XANES data. These data can be used to gain a better understanding of bulk chemical changes in chars and their leachates, as well as detailed molecular changes to P. This data package is associated with the GitHub repository found at https://github.com/river-corridors-sfa/rcsfa-RC3-BSLE_P. This data package is comprised of a “data” folder and a series of data processing and analysis scripts. Details on how to recreate the workflow can be found in the Critical Details section of the readme and the “workflow_readme.md” file. The file-level metadata file (file ending in “flmd.csv”) lists all files contained in this data package and descriptions for each. The data dictionary (file ending in “dd.csv”) describes all tabular data columns and their respective definitions and units.

54 ENVIRONMENTAL SCIENCES↗

Fuel/Basket Degradation Modeling Summary

This report documents the activities in a preliminary phase of development for three models: 1) waste package breach model, 2) fuel/basket degradation model, and 3) dual-purpose canister (DPC) crush model. The waste package breach model describes the coupling of mass flow, heat transport, and canister shell deformation in response to a heat-generating (criticality) event. The fuel/basket degradation model describes potential weakening and disaggregation of the DPC structure from corrosion, possibly accelerated by seismic ground motion. Progressively degraded three-dimensional (3D) configurations of the fuel, basket, and shell are generated for future analysis of reactivity (with as-loaded DPC fuel characteristics). Another important application for the fuel/basket degradation model is validation of the two stylized degradation cases currently being used by other investigators for analysis of the as-loaded DPC inventory under disposal conditions. The DPC crush model investigates stability of a typical DPC after breach of the disposal overpack allows fluids from the repository near-field environment to penetrate and externally pressurize the canister shell. Preliminary results show that large deformation of the DPC could occur for external pressure on the order of 10 to 15 MPa, or the shell could be stable (not collapse) with pressure of 20 MPa or greater if the basket plates are fully welded at the connections.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Precursor Analysis Report: Doppelpaymer Ransomware Attack on Petroleos Mexicanos (PEMEX) 2019

The DoppelPaymer Ransomware Attack on Petroleos Mexicanos (PEMEX) 2019 Precursor Analysis Report leverages publicly available information about the PEMEX cyber attack and catalogs anomalous observables for each technique employed in the attack. This analysis is based upon the methodology of the Cybersecurity for the Operational Technology Environment (CyOTE) program. The 2019 DoppelPaymer ransomware attack on PEMEX, Mexico’s nationalized petroleum corporation, highlights a unique threat that ransomware and cybercriminal extortion poses to Operational Technology (OT) environments in critical infrastructure. The incident began with an employee downloading commodity malware that allowed adversaries to gain initial access to PEMEX’s enterprise environment. After conducting privilege escalation, tool ingress, and data exfiltration, the adversaries deployed DoppelPaymer ransomware throughout the PEMEX enterprise environment, resulting in the company having to take dozens of systems offline for at least several days. Although PEMEX stated that their operations were not affected, the data exfiltrated from PEMEX was made available for download on DoppelPaymer’s leak site, as well as on other illicit criminal forums. This stolen data included not only company information, but also sensitive OT-specific configuration data. This incident showcases how cybercriminal exfiltration and posting of sensitive OT architecture documentation can pose security concerns for the targeted organization for years due to the long lifespan of OT assets and architectures. Researchers and analysts identified 18 unique techniques utilized during the attack with a total of 190 observables using MITRE ATT&CK® for Industrial Control Systems. The CyOTE program assesses observables accompanying techniques used prior to the triggering event to identify opportunities to detect malicious activity. If observables accompanying the attack techniques are perceived and investigated prior to the triggering event, earlier comprehension of malicious activity can take place. Fifteen of the identified techniques used during the DoppelPaymer ransomware attack were precursors to the triggering event. Analysis identified 163 observables associated with these precursor techniques, 34 of which were assessed to have an increased likelihood of being perceived in the 60 days preceding the triggering event. The response and comprehension time could have been reduced if the observables had been identified earlier. The information gathered in this report contributes to a library of observables tied to a repository of artifacts, data sources, and technique detection references for practitioners and developers to support the comprehension of indicators of attack. Asset owners and operators can use these products if they experience similar observables or to prepare for comparable scenarios.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Workshop to Plan R&D Support of Fuel/Basket Modification for Direct Disposal of Future DPCs

By 2030 about half of all spent nuclear fuel (SNF) arising from the current fleet of commercial power plants will be in dual-purpose canisters (DPCs), which are designed for storage and transportation but not for disposal. As an alternative to complete repackaging of the fuel for disposal, considerable cost savings and lower worker dose could be realized by directly disposing of this SNF in DPCs. The principal technical consideration is criticality control in a geologic repository, because the DPCs are large and depend on neutron absorbing basket components for criticality control. Neutron absorbing materials are generally aluminum-based, and under disposal conditions can degrade after a few hundred years contact with ground water. Simple modifications to the SNF assemblies or the DPC baskets could help to achieve direct disposal, and this is one of the approaches being studied to address the possibility of disposal criticality (SNL 2020a). Five fuel/basket modification concepts have been proposed (SNL 2020b) and a virtual workshop was conducted to solicit review and feedback on these concepts. The proposed solutions are: 1) zone loading of DPCs to limit reactivity, 2) replacing absorber plates with advanced neutron absorbing (ANA) material, 3) adding disposal control rods to pressurized water reactor (PWR) assemblies, 4) rechanneling boiling water reactor (BWR) assemblies with ANA material, and 5) basket insert plates (chevron inserts) made from ANA material. The presentations from the workshop are provided in this report, and the workshop discussions are summarized. This information includes prioritization of the proposed fuel/basket modification solutions, and prioritization of the associated model development, validation testing, and quality assurance activities. Information documented in this report will help to steer research and development efforts at Sandia National Laboratories, Oak Ridge National Laboratory, and Idaho National Laboratory that support the U.S. Department of Energy, Office of Nuclear Energy, Spent Fuel and Waste Science and Technology program

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Next Generation System Analysis Model Recently Added Features and Future Plans - Abstract

The Nuclear Waste Policy Act of 1982, as amended (NWPA 1982), established the federal government’s responsibility to accept spent nuclear fuel (SNF) and high-level radioactive waste (HLW) from waste owners and generators for ultimate disposition. SNF generated by the current fleet of commercial nuclear reactors is being stored at the reactor sites in spent fuel pools (SFPs) and in dry independent spent fuel storage installations (ISFSIs). The US Department of Energy Office of Nuclear Energy (DOE-NE) is developing an Integrated Waste Management Program (IWMP) comprising a suite of options and supporting analyses to enable future informed choices. The IWMP is applying integrated waste management system architecture analysis, system engineering, and decision analysis principles to inform potential future decisions regarding potential nuclear waste management system architectures. Architecture analyses of the IWM system are being conducted to support the future deployment of a comprehensive system for managing nuclear waste that considers all major aspects of the back end of the nuclear fuel cycle (i.e., transportation, storage, and disposal). The Next Generation System Analysis Model (NGSAM) is an agent-based simulation software tool designed for the express purpose of modeling the IWM system. NGSAM imports data from the Oak Ridge National Laboratory (ORNL) Unified Database (e.g., historic assembly information, thermal profiles for assembly heat, at-reactor dry storage loadings) to ensure that the simulation initializes with a realistic representation of the state of commercial SNF in the United States. Recent major enhancements that have been implemented into NGSAM since NGSAM was last presented at the WM2019 conference include: • Tracking of railroad escort and buffer car acquisition. • Addition of heavy haul and barge routes for some sites, as well as support for user-defined inter-modal routes. • Updates to the logic that checks the thermal maps prior to package transport. • Addition of an allocation method that predicts when reactor sites will pack assemblies from their pools for dry storage and allocates packages to those reactor sites in the preceding periods, favoring direct transport packages and reducing the number of packages that reactor sites pack for dry storage at their ISFSIs. • Addition of reactor site family operational limits, which are used to limit the number of loads from the pool and from dry storage at a given reactor site per year. • Support has been added for multiple canister loading maps and packages having multiple compatible transportation overpacks. • Updates in the handling of non-commercial fuel, including a new database containing data to support the updates. • Support for repackaging at reactor sites. • Implementing additional output reports or modifying existing reports. • User edits can now be created and edited via the NGSAM website. • Ability to load packages for dry storage at ISF pools. • Same-type package blending at DOE sites. • Support for multi-mode transloading at reactor sites. These new features have improved NGSAM capabilities and/or improve the user experience with the model and will be discussed in more detail. The initial NGSAM requirements for advanced reactor fuels, reprocessing, treatment, and conditioning are preliminary and are described at a high level in this paper: analysts will provide more specific requirements to the NGSAM team in the future. Additionally, there are many data needs associated with modeling advanced reactors in NGSAM, but many of the data or plans are still in progress and/or yet to be fully defined. However, this document describes an initial exploration of the data relevant to this program. Advanced reactor data will likely require revision as concepts evolve and new considerations are made. This is a technical paper that does not take into account contractual limitations or obligations under the Standard Contract for Disposal of Spent Nuclear Fuel and/or High-Level Radioactive Waste (Standard Contract) (10 CFR Part 961). For example, under the provisions of the Standard Contract, spent nuclear fuel in multi-assembly canisters is not an acceptable waste form, absent a mutually agreed to contract amendment. To the extent discussions or recommendations in this paper conflict with the provisions of the Standard Contract, the Standard Contract governs the obligations of the parties, and this paper in no manner supersedes, overrides, or amends the Standard Contract. This paper reflects technical work which could support future decision making by DOE. No inferences should be drawn from this paper regarding future actions by DOE, which are limited both by the terms of the Standard Contract and Congressional appropriations for the Department to fulfill its obligations under the Nuclear Waste Policy Act including licensing and construction of a spent nuclear fuel repository.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Sensitivity Analysis of Drivers Water Shortage in the Los Angeles Region During Drought

The code and detailed step-by-step instructions for generating the model output data, processing results, and analysis and plotting are provided at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. The PyArtes model is a python adaptation of the Artes model. PyArtes uses many of the same input data and optimization model architecture as Artes. Documentation for the PyArtes model is provided in the Supplement to the paper. The primary data product are simulated monthly water shortages for indoor and outdoor demand under a large ensemble of drought scenarios (>13,000). The droughts are hypothetical and are not based on historical time series data of supply sources - though historical data did help inform ranges explored for supply parameters. Demands are informed by recent 2017-2021 water supply data. Demands used for the model can be accessed at https://github.com/IMMM-SFA/Ferencz_et_al_2026_ER_Water. Simulations resolve demand for over 90 water providers in the study region. The results report 36 months of water shortage data for each indoor and outdoor demand node. The study also developed a multilayer perceptron (MLP) neural network trained on a subset of the simulated shortage ensemble to emulate worst annual water shortage for a given set of parameter multipliers -- provided the parameter values fall within the ranges sampled in the ensemble. Emulated water shortages for synthetic ensembles are in the MLP-generated shortages folder. The MLP model was used to generate larger ensembles to support Sobol analysis that would have been extremely computationally expensive to simulate. Datasets provided in this repository*: Simulated shortages. These results are used for the analysis for Figures 5, 8, and 9 in the paper, and also to train the MLP emulator. .zip file containing outputs for the 13,312 scenario ensemble. Separate .csv files for indoor and outdoor shortage for each scenario. Rows = demand ids (~100), Columns = months (36) Units = acre-feet/month of shortage (shortage = monthly demand - supply). 1 acft = 1233.48 m^3 .csv files of aggregated shortages derived from the 13,312 ensemble Rows = scenarios (13,312), Columns = demand ids (~100) Units = acre-feet/year (either worst annual shortage or total shortage over the 3-year drought) .csv file of the parameter multipliers scenarios for the ensemble .csv file of the parameter ranges and baseline values the multipliers were applied to MLP-generated shortages. These results are used for Figures 4, 6, and 7 in the paper. mwd higher folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results Emulated shortages. Rows = scenarios, columns = demand ids, units acft Sobol results. Rows = demand ids, columns Sobol (S1, ST, or 95% confidence interval) value for each parameter mwd lower folder: scenario ensembles, emulated worst year total shortages (acft), and Sobol results same organization as mwd higher MLP performance: performance metrics (R^2, RMSE, BIAS, MAPE) for the testing subset (20% or 2,662 scenarios) and simulated vs emulated worst year shortage (acre-feet/year) for every demand node, MWD wholesale regions, and the entire study region (LAC). Supporting data for figures. Figure plotting scripts in the associated GitHub repo. These files support analysis and visualization. Geospatial Data used for plotting simulated water shortages and Sobol results. Dictionary of full names for demand nodes in the model and estimates of water supply by source type informed by Artes input files and California Urban Water Management Planning data: https://water.ca.gov/Programs/Water-Use-And-Efficiency/Urban-Water-Use-Efficiency/Urban-Water-Management-Plans *Readme files provided for each folder.

drought↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

Integrated GW Farm ABM

This Data Repository includes data used for the integrated groundwater- farm ABM model, raw model output from scenario ensemble, and processed outputs that isolate the groundwater storage depletion outcomes for the 35,000 farm cells. Model Inputs: Farm ABM Inputs: This folder contains the input data used by the integrated groundwater - farm ABM modelling script (Python file) used for the high performance computing (HPC) experiments. The sub-folder "data inputs" contains all of the farm attribute data, while the three files in the folder have the hydrogeological data lookup table (NLDAS Cost Curve Attributes.csv), a lookup table (Theis well function table.csv) for the groundwater cost curve function, and the farm indexes and corresponding NLDAS ids for all of the cells run in this experiment (nldas farms subset final.csv). NLDAS Cost curve hydrogeological data: Hydrogeological data aggregated to 1/8 degree resolution and aligned with the NLDAS grid. Parameters include: water depth below ground surface [meters], subsurface porosity [unitless], aquifer depth from ground surface to aquifer bottom [meters], annual average recharge (USGS: mm, Doll: meters), and three different hydraulic conductivity (K) values (meters/day). The three K values represent the mean value from Gleeson et al. (2018), one standard deviation above the mean from Gleeson et al. (2018), and the de Graaf et al. 2020 modifications to certain lithologies. Additional information about these datasets and their processing are documented in the supplement to Yoon et al. 2025 (in review). Output: Raw outputs: This folder contains a .zip file that has model outputs for the entire scenario ensemble. There is one csv for each farm id, using the format "farm farmid cases.csv". The relationship between the farm id and NLDAS id is defined by the "nldas farms subset final.csv" located in the Farm ABM Inputs folder. Each csv has 625 rows, corresponding to 625 combinations of different scenario parameter values. Each row (scenario) represents the outcome of a 100 year simulation. Columns define scenario settings and summary statistics for each scenario. The first four columns define the scenario settings: "hydro ratio," "econ ratio," "K scenario," and "gamma scenario." The hydro and econ ratios are values passed to the modeling script that influence multipliers for other model parameters, as documented in the supplement to Yoon et al. 2025 (in review). The gamma multiplier is a coefficient multiplier applied to the baseline gamma values (values below 1 represent lower unobserved costs compared to baseline, values above 1 represent higher costs). The K scenario names represent K values of: "low": 0.5 m/d, "int 1": 2.5 m/d, "int 2": 10 m/d, "high": 50 m/d, and "gleeson": mean Gleeson K value. "Perc vol depleted" is the fraction of groundwater depleted at the end of the 100 simulation. Processed Output: Derived depletion outcomes from raw outputs: All of the individual csv files from the Raw outputs were aggregated into a single file that has the scenario settings and fraction depletion "Perc vol depleted" for every farm cell, for every scenario. The other two files define relationships between the farm id, NLDAS id, and local and major aquifer units, used for aquifer-level depletion analysis.

Agent based modeling↗

SEGUID v2: Extending SEGUID checksums for circular, linear, single- and double-stranded biological sequences

Background Synthetic biology involves combining different DNA fragments, each containing functional biological parts, to address specific problems. Fundamental gene-function research often requires cloning and propagating DNA fragments, such as those from the iGEM Parts Registry or Addgene, typically distributed as circular plasmids. Addgene’s repository alone offers around 150,000 plasmids. To ensure data integrity, cryptographic checksums can be calculated for the sequences. Each sequence has a unique checksum, making checksums useful for validation and quick lookups of associated annotations. For example, the SEGUID checksum uniquely identifies protein sequences with a 27-character string. Objectives The original SEGUID, while effective for protein sequences and single-stranded DNA (ssDNA), is not suitable for circular DNA since there is no natural starting position nor for double-stranded DNA (dsDNA) since two separate sequences are present. Challenges include how to uniquely represent linear dsDNA, circular ssDNA, and circular dsDNA. To meet these needs, we propose SEGUID v2, which extends the original SEGUID to handle additional types of sequences. Conclusions SEGUID v2 produces orientation and rotation invariant checksums for single-stranded, double-stranded, possibly staggered, linear, and circular DNA and RNA sequences. Customizable alphabets allow for other types of sequences. In contrast to the original SEGUID, which uses Base64, SEGUID v2 uses Base64url to encode the SHA-1 hash. This ensures SEGUID v2 checksums can be used as-is in filenames, regardless of platform, and in URLs, with minimal friction. Availability SEGUID v2 is readily available for major programming languages, distributed under the MIT license. JavaScript package seguid is available on npm, Python package seguid on PyPi, R package seguid on CRAN, and a Tcl script on GitHub. These tools, along with documentation, examples, and an online SEGUID Calculator , can be found at https://www.seguid.org .

Pereira, Humberto↗

Consolidated Hydropower Data Repository: Value and Opportunities

Hydropower is one of several types of generating assets that provides energy, capacity, and services to electric power systems. It does so under rubrics and objectives—market driven and regulated, internal and external to asset and fleet owners—that address reliability, cost, price, and, increasingly, flexibility of output. The aggregation of data from multiple hydropower units can provide insights into asset operations and maintenance practices and needs and assist in meeting hydropower objectives. This paper examines the concept and potential benefits of aggregating hydropower asset data—primarily supervisory control and data acquisition (SCADA) information—with examples of insights developed from data aggregated by the Hydropower Research Institute (HRI). Data aggregation as discussed herein, and as implemented by the HRI, extends beyond multiple units in a powerhouse and beyond multiple hydropower facilities in an electric utility fleet or river system. Examples of research and analytics from such aggregated datasets range from unit load dependency analyses to modeling sensor measurements to detect and diagnose anomalies in assets. These examples showed the benefits of utilizing the entire dataset for insights into how the sensor layout of a single unit or set of units compares to the hydropower industry overall. Such insights include whether additional sensors are needed to complete analyses or to make decisions. In addition, utilizing multiple sensors of the same kind within a unit can provide an indication of possible current or upcoming problems with equipment. Although other analyses are possible, their use requires the development of complex models and, potentially, access to types of data that are currently not included with the example dataset used in this study. However, the examples studied herein confirmed the value of data aggregation in the fleet and unit contexts, and the value extends beyond multiple units in a powerhouse and beyond multiple hydropower facilities in an electric utility fleet or river system. The assessments also provided insights into potential extensions to the data aggregation concept that could further add to their value to the hydropower community, and these are included in this document as a set of recommendations.

13 HYDRO ENERGY↗

IDAES-PSE 2.1.0 Release

The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost, most environmentally sustainable solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.1.0 Release Highlights New IDAES Examples Repository Starting with this release, the IDAES examples are developed in the new IDAES/examples repository. Along with many content and usability improvements, the most significant changes are: To install the examples, after installing IDAES, run pip install idaes-examples The idaes get-examples command, previously used for this, has been removed The HTML version is now available at https://idaes-examples.readthedocs.io The previous URL, https://idaes.github.io/examples-pse, will not be updated and may be removed at some point in the future For more details, refer to the resources available at IDAES/examples. Removal of Non-Functional Apps A review of the code in the idaes/apps and idaes/models_extra folders was undertaken, and a number of tools were identified as being outdated or non-functional and no longer supported by their development teams. Due to this, the following tools have been removed: idaes/apps/alamopy_depr (note that the new ALAMOpy interface remains avaialble in idaes/core/surrogates) idaes/apps/helmet idaes/apps/ripe idaes/apps/roundingRegression idaes/models_extra/carbon_capture Pyomo 6.6 This version of IDAES is the first requiring Pyomo 6.6. This version of Pyomo contains multiple internal improvements and refactorings. While for the majority of cases this should have positive or no impact on solvability of IDAES models, we are aware of a small number of models that have been affected as a result of these changes. For more information, refer to the Pyomo 6.6.1 release notes. Other highlights Model Initialization A prototype API for a new approach to initializing IDAES models is now available which makes available some new techniques for initializing models. This is documented in the Initializing Models Reference Guide Modular Properties Framework Support for some transport properties Helmholtz Equation of State properties Better error checking for case where unit models are set to include phase equilibrium but the property package is set to support only a single phase Multi-Stream Contactor model: a new base model for systems involving contacting of two or more streams with mass transfer. This model is intended to be used as the foundation for models such as membrane separators, solvent extraction and other similar processes. This is documented in the Multi-Stream Contactor Reference Guide idaes/models_extra/power_generation report() methods for unit models using Helmholtz equation of state General Code Maintenance Streamlining of dependencies and creation of new optional dependency groupings to support non-core tools General linting of codebase to ensure compliance with most pylint checks Spell checking of all code and doc strings Removal of backward compatibility code for Python 2

IDAES↗

Supply Chain Risk Management: Data Structuring

Supply chain risk management (SCRM) is an area of research that addresses both logistics concepts to maximize efficiency, reliability, and revenue as well as risk features, such as potential weak points, break points, and vulnerabilities within the supply chain. SCRM is used to find risks introduced at each node in a supply chain and how these risks can impact a company’s products, individuals, customers, and reputation. SCRM is a relatively new field, so standardized processes including data structuring are not fully documented. This paper explains the importance of a standard data structuring methodology and how it can enhance current SCRM efforts. Data ingest, structuring, and analysis are predominantly managed by humans. Automating some of the less complex steps can positively impact SCRM by allowing human analysts to focus on more strategic analyses. Types of data to be collected and structured are collected via publicly available information related to hardware, software, and corporate entities. After the data has been collected, the information is formatted in a specific manner, conforming to a schema, to allow for more effective and efficient ingest for further analysis. This paper outlines data structures used by Pacific Northwest National Laboratory for SCRM research and analysis purposes. These structures have been used for hundreds of analyses and have been successful in developing a common baseline. Data structuring is one of the first steps in data standardization, which will further mature and enhance the SCRM research area.

supply chain risk management, data structuring, re↗