Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Modelling Framework For Fire And Smoke Detection In Imagery

This framework was developed to ingest images/video images to train and test artificial neural network structures for image-based detection of fire and smoke. The code includes data preprocessing, model development, and testing. The framework is designed to work with RGB video imagery.

Griffel, LloydM.↗

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

HydraGNN_GFM_FineTuning4Materials v1.0

This repository enables fine-tuning of the HydraGNN Predictive GFM 2026 — an open-source ensemble of pre-trained graph foundation models for atomistic materials modeling, developed at Oak Ridge National Laboratory. The GFM 2026 is freely available and downloadable via Globus from the OLCF Data Constellation (DOI: 10.13139/OLCF/2562660). Starting from these pre-trained weights, this repository provides a complete transfer learning pipeline for adapting the GFM ensemble to domain-specific molecular and materials property prediction tasks. It includes: 1) Utilities for ensemble fine-tuning with task-specific output heads 2) Example pipelines for eight widely-used materials and molecular datasets 3) Tools for model adaptation and head configuration 4) Data preprocessing utilities for each supported dataset 5) Benchmarking and evaluation scripts

Ungerboeck, Linda↗

REDI – Readiness Engine for Data Integration

The Readiness Engine for Data Integration (REDI) is an open-source framework for automating, standardizing, and assessing the process of preparing scientific data for AI training. REDI implements a five-stage pipeline (ingest, preprocess, transform, structure, output) with per-stage provenance instrumentation via Flowcept, domain-aware transformation logic (PII anonymization, regridding, graph encoding, and more), and built-in readiness assessment and validation modes. REDI has been evaluated across climate, proteomics, materials science, and nuclear fusion datasets, demonstrating near-ideal parallel scaling to 100 nodes on OLCF's Frontier system. REDI is deployable as an agent-callable skill in coding environments such as Claude Code and OpenAI Codex, and is complemented by SetGo for FAIR compliance and catalog publication.

Brewer, Wesley [Oak Ridge National Laboratory (ORN↗

Asi Nuclear Energy Sensors Data Portal Chatbot And Data Structuring Tool

The Idaho National Laboratory (INL) is advancing the development of an AI-powered chatbot and data structuring tool specifically designed to accelerate data mining processes for sensor-related information and seamlessly integrate the results into the ASI Sensors Data Portal (https://nes.energy.gov/). By doing so, the software aims to enhance the accessibility, usability, and organization of sensor data for nuclear energy applications. The software initial phase focuses on retrieving comprehensive datasets, prioritizing the past five years of publicly available information from the Office of Scientific and Technical Information (OSTI). These datasets will be meticulously processed to ensure compatibility, employing cleaning and preprocessing steps to eliminate irrelevant, incomplete, or corrupted information, thus establishing a robust foundation for subsequent AI use. The data will serve as the backbone for training an AI model and chatbot, which will act as an interactive tool enabling users to ask complex, context-specific questions and receive accurate, validated answers derived from constrained literature. In parallel, the project incorporates a data structuring process supported by AI to organize sensor information from multiple sources into a standardized format. This structured data will include detailed sensor specifications, such as measurement range, applications, accuracy, and operating conditions, generated and documented with AI. These specifications will be systematically integrated into the sensor portal. To maintain the highest levels of accuracy and relevance, all AI-generated outputs will be reviewed and validated by subject matter experts (SMEs), with additional fields or parameters added as needed. Future stages of the project aim to expand the dataset beyond OSTI to include other sources and potentially incorporate unclassified controlled information (UCI) with restricted access protocols to address security and confidentiality requirements.

Mapes, NormanJ. [Idaho National Laboratory (INL), ↗

Let’s Unleash the Network Judgment: A Self-Supervised Approach for Cloud Image Analysis

Accurate cloud type identification and coverage analysis are crucial in understanding the Earth’s radiative budget. Traditional computer vision methods rely on low-level visual features of clouds for estimating cloud coverage or sky conditions. Several handcrafted approaches have been proposed; however, scope for improvement still exists. Newer deep neural networks (DNNs) have demonstrated superior performance for cloud segmentation and categorization. These methods, however, need expert engineering intervention in the preprocessing steps—in the traditional methods—or human assistance in assigning cloud or clear sky labels to a pixel for training DNNs. Such human mediation imposes considerable time and labor costs. We present the application of a new self-supervised learning approach to autonomously extract relevant features from sky images captured by ground-based cameras, for the classification and segmentation of clouds. We evaluate a joint embedding architecture that uses self-knowledge distillation plus regularization. We use two datasets to demonstrate the network’s ability to classify and segment sky images—one with ~85,000 images collected from our ground-based camera and another with 400 labeled images from the WSISEG database. We find that this approach can discriminate full-sky images based on cloud coverage, diurnal variation, and cloud base height. Additionally, it semantically segments the cloud areas without labels. The approach shows competitive performance in all tested tasks, suggesting a new alternative for cloud characterization.

54 ENVIRONMENTAL SCIENCES↗

Classification of Cloud Particle Imagery from Aircraft Platforms Using Convolutional Neural Networks

Abstract A vast amount of ice crystal imagery exists from a variety of field campaign initiatives that can be utilized for cloud microphysical research. Here, nine convolutional neural networks are used to classify particles into nine regimes on over 10 million images from the Cloud Particle Imager probe, including liquid and frozen states and particles with evidence of riming. A transfer learning approach proves that the Visual Geometry Group (VGG-16) network best classifies imagery with respect to multiple performance metrics. Classification accuracies on a validation dataset reach 97% and surpass traditional automated classification. Furthermore, after initial model training and preprocessing, 10 000 images can be classified in approximately 35 s using 20 central processing unit cores and two graphics processing units, which reaches real-time classification capabilities. Statistical analysis of the classified images indicates that a large portion (57%) of the dataset is unusable, meaning the images are too blurry or represent indistinguishable small fragments. In addition, 19% of the dataset is classified as liquid drops. After removal of fragments, blurry images, and cloud drops, 38% of the remaining ice particles are largely intersecting the image border (≥10% cutoff) and therefore are considered unusable because of the inability to properly classify and dimensionalize. After this filtering, an unprecedented database of 1 560 364 images across all campaigns is available for parameter extraction and bulk statistics on specific particle types in a wide variety of storm systems, which can act to improve the current state of microphysical parameterizations.

54 ENVIRONMENTAL SCIENCES↗

Monitoring Noble Gases (Xe and Kr) and Aerosols (Cs and Rb) in a Molten Salt Reactor Surrogate Off-Gas Stream Using Laser-Induced Breakdown Spectroscopy (LIBS)

In this study with surrogate materials we show that laser-induced breakdown spectroscopy (LIBS) is a robust tool with promising capability toward monitoring gaseous (Xe and Kr) and aerosol (Cs and Rb) species in an off-gas stream from a molten salt reactor (MSR). MSRs will continually evolve fission products into the cover gas flowing across the reactor headspace. The cover gas entrains Xe and Kr gases, along with aerosol particles, before passing into an off-gas treatment system. Univariate models of Xe and Kr peaks showed a strong correlation to concentration indicated by their coefficients of determination of 0.983 and 0.997, respectively. Multivariate models were built for all four analytes using partial least squares regression coupled with preprocessing steps including normalization, trimming, and/or genetic algorithm derived filters. The models were evaluated by predicting the concentrations of the analytes in four validation samples, in which all calibration models were successfully validated at a confidence interval of 99.9%. Finally, pressure controllers were used to regulate the mass flow rate of Kr flowing into the measurement cell in sinusoidal and stepwise waveforms to test the real-time monitoring capabilities of the regression models. Both univariate and partial least squares Kr models were able to successfully quantify the gas concentration in the real-time evaluation. The root mean squared error of prediction (RMSEP) values for these real-time tests were calculated to be 0.051, 0.060, and 0.121 mol% demonstrating the measurement systems’ capability to perform online monitoring with acceptable accuracy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

The Artificial Intelligence Ontology: LLM-Assisted Construction of AI Concept Hierarchies

The Artificial Intelligence Ontology (AIO) is a systematization of artificial intelligence (AI) concepts, methodologies, and their interrelations. Developed via manual curation, with the additional assistance of large language models (LLMs), AIO aims to address the rapidly evolving landscape of AI by providing a comprehensive framework that encompasses both technical and ethical aspects of AI technologies. The primary audience for AIO includes AI researchers, developers, and educators seeking standardized terminology and concepts within the AI domain. We use the term “branches” for classes, and their subclasses, in our ontology that are subclasses of owl:Thing. AIO contains eight branches: Bias, Layer, Machine Learning Task, Mathematical Function, Model, Network, Preprocessing, and Training Strategy, each designed to support the modular composition of AI methods and facilitate a deeper understanding of deep learning architectures and ethical considerations in AI. AIO uses the Ontology Development Kit (ODK) for its creation and maintenance, with its content being more easily updated through AI-driven curation support. This approach not only ensures the ontology's relevance amidst the fast-paced advancements in AI but also significantly enhances its utility for researchers, developers, and educators by simplifying the integration of new AI concepts and methodologies. The ontology's utility is demonstrated through the annotation of AI methods data in a catalog of AI research publications and the integration into the BioPortal ontology resource, highlighting its potential for cross-disciplinary research. The AIO ontology is open source and is available on GitHub ( https://w3id.org/aio/ ) and BioPortal ( https://bioportal.bioontology.org/ontologies/AIO ).

Joachimiak, Marcin P. [Biosystems Data Science Dep↗

Transplatformer: translating toxicogenomic profiles between generations of platforms

Background Transcriptomic profiling technologies have advanced the analysis of biological and toxicological responses. However, substantial differences in probe design, dynamic range, gene coverage, and preprocessing pipelines across platforms introduce artifacts that limit cross-study integration and hinder the reuse of historical datasets. We aim to develop computational methods for accurate cross-platform translation to maximize the value of legacy resources. Results We present TransPlatformer a deep learning framework for translating gene expression profiles across heterogeneous toxicogenomics platforms. TransPlatformer employs a novel attention-based architecture to map high-dimensional fold-change vectors from legacy microarray technologies to current platforms. Models are trained and evaluated using DrugMatrix, spanning three technological generations. We investigate mixed-tissue, single-tissue, and cross-tissue training paradigms and benchmark performance against multilayer perceptron and matrix-completion baselines. In mixed-tissue training, TransPlatformer achieves a greater than 50% reduction in mean absolute error (0.043 vs. 0.09) and nearly doubles Pearson correlation ( ≈ 0.71 vs. 0.37) relative to baseline methods. Importantly, TransPlatformer preserves rare but biologically meaningful over- and under-expressed signals, with mean absolute error below 0.22. Single-tissue models yield further improvements for well-represented organs, such as a 10% reduction in liver mean absolute error, while underscoring the need for data augmentation strategies in low-sample tissues.ra Conclusions TransPlatformer provides an effective and scalable computational solution for cross-platform transcriptomic translation. By enabling biologically faithful harmonization of gene expression data, the proposed approach facilitates the reuse of legacy toxicogenomics datasets, enhances downstream biomarker discovery, and supports more reproducible predictive modeling in toxicology.

59 BASIC BIOLOGICAL SCIENCES↗

Improving multiwell petrophysical interpretation from well logs via machine learning and statistical models

Well-log interpretation estimates in situ rock properties along well trajectory, such as porosity, water saturation, and permeability, to support reserve-volume estimation, production forecasts, and decision making in reservoir development. However, due to measurement errors, variability of well logs caused by multiple measurement vendors, different borehole tools, and nonuniform drilling/borehole conditions, estimations of rock properties with original well logs without proper preprocessing may not be accurate, especially in the context of multiwell estimation. Well-log normalization techniques such as two-point scaling and mean-variance normalization are commonly used to improve the robustness of multiwell rock-property estimation. However, these techniques do not consider the correlation between well logs and require subjective knowledge for their effective implementation. To reduce uncertainties and processing time associated with multiwell rock-property estimation from well logs, we develop discriminative adversarial (DA) and linear constraint models for well-log normalization and rock-property estimation. The DA neural network model developed for well-log normalization and interpretation can perform linear and nonlinear well-log normalization while considering the joint distribution of each well log and rock properties. However, the linear constraint model uses an ensemble of predictions from linear models to constrain well-log normalization and rock-property estimation. We also develop a divergence-based type well identification method to select type (training) wells for a test well based on the statistical similarity of associated well-log distributions instead of the interwell distance. We apply the DA model to perform well-log normalization and prediction of permeability for the Seminole San Andres Unit carbonate reservoir. Compared with the permeability predicted with the classical machine learning model without well-log normalization and models with two-point scaling normalization, the DA model yields the most accurate permeability prediction by decreasing the mean-squared error of permeability prediction by 20%–50%.

Geochemistry & Geophysics↗

Chapter 4: "Waste"-to-Energy for Decarbonization - Transforming Nut Shells Into Carbon-Negative Electricity

This chapter presents a study demonstrating waste pistachio nut shells as a renewable feedstock for climate-friendly electricity generation via industrial gasification technology. The study includes biomass feedstock characterization (i.e., pistachio waste critical material attributes), process variability (i.e., bulk material handling), and overall operational reliability and conversion performance through extended testing. Additionally, techno-economic analysis (TEA) and life cycle assessment (LCA) were performed to assess the economic feasibility and environmental impact of the technology to transform agricultural waste to biopower. For processing pistachio waste material, among critical material attributes, fines content in the biomass (<1/4") had the largest potential to reduce the operating time of the gasifiers due to plugging. Pelletizing fines and co-feeding them with the mixed pistachio waste increased the average feed density, feed rate, and biochar production. Compared to pine wood chips, mixed pistachio waste yielded higher biochar quantity but slightly reduced quality. In general, a systematic Quality by Design methodology is the preferred approach for designing preprocessing and material conveyance systems, where a downstream technology (end user) for the produced intermediate is specified at the outset. TEA results show that the biochar production rate and selling price had an overwhelming impact on the modeled Minimum Electricity Selling Price (MESP), which ranged from 35.5 to 39.9 cents/kWh for the cases studied (16 h/day operational basis). Moreover, LCA results show that the valorization of pistachio shells for biopower generation is a "carbon negative" process that can help decarbonize the U.S. electricity grid. The specific carbon intensity was -0.29 to -0.71 kg CO2e/kWh, compared to 0.45 kg CO2e/kWh for the average U.S. electricity mix. Biochar production from pistachio waste as a potential means for carbon sequestration was a significant driver for the LCA. The highly stable biochar permanently sequesters a considerable fraction of biochar carbon in the ground, more than enough to offset the life cycle emissions, and can be a complementary climate change mitigation strategy.

bio-char↗

Capturing Travel Mode Adoption in Designing On-Demand Multimodal Transit Systems

This paper studies how to integrate rider mode preferences into the design of on-demand multimodal transit systems (ODMTSs). It is motivated by a common worry in transit agencies that an ODMTS may be poorly designed if the latent demand, that is, new riders adopting the system, is not captured. This paper proposes a bilevel optimization model to address this challenge, in which the leader problem determines the ODMTS design, and the follower problems identify the most cost efficient and convenient route for riders under the chosen design. The leader model contains a choice model for every potential rider that determines whether the rider adopts the ODMTS given her proposed route. To solve the bilevel optimization model, the paper proposes an exact decomposition method that includes Benders optimal cuts and no-good cuts to ensure the consistency of the rider choices in the leader and follower problems. Moreover, to improve computational efficiency, the paper proposes upper and lower bounds on trip durations for the follower problems, valid inequalities that strengthen the no-good cuts, and approaches to reduce the problem size with problem-specific preprocessing techniques. The proposed method is validated using an extensive computational study on a real data set from the Ann Arbor Area Transportation Authority, the transit agency for the broader Ann Arbor and Ypsilanti region in Michigan. The study considers the impact of a number of factors, including the price of on-demand shuttles, the number of hubs, and access to transit systems criteria. The designed ODMTSs feature high adoption rates and significantly shorter trip durations compared with the existing transit system and highlight the benefits of ensuring access for low-income riders. Finally, the computational study demonstrates the efficiency of the decomposition method for the case study and the benefits of computational enhancements that improve the baseline method by several orders of magnitude. Funding: This research was partly supported by National Science Foundation [Leap HI Proposal NSF-1854684] and the Department of Energy [Research Award 7F-30154].

Operations Research & Management Science↗

Data Projection of the High Temperature Electrolysis System in the Dynamic Energy Transport and Integration Laboratory using Dynamic System Scaling

For nuclear power to be flexible in a functioning Integrated Energy System (IES), excess produced heat must be stored or utilized during times of low power demand to ensure a load factor of 1 while load balancing. The Dynamic Energy Transport and Integration Laboratory (DETAIL) is one facility that is under development to emulate IES conditions on the engineering-scale, planned to conduct virtual real time operations with industry-scale facilities, and is currently testing thermal storage and high temperature electrolysis. As part of the study to develop a method to preprocess input signals or postprocess output signals between systems of different scales via Dynamical System Scaling (DSS), the current research is one of the continued efforts branching from the data projection activity conducted for the Thermal Energy Distribution System and currently engages the High Temperature Electrolysis (HTE) System in DETAIL. The HTE SOEC electrical, fluid, and thermal dynamics Figure of Merits (FOM) were identified, governing equations and closure relations were successfully scaled, and relations between FOM scaling ratios were determined. Setting the scaling objectives to reform existing data to project a data set that doubly accelerated the electrolysis process while preserving the produced amount of hydrogen was generated for the full transient. The calculated boundary conditions were inlet temperature, stack current, and inlet steam mass flow rate at 1470 K, 121.1 A, and 1.886 g/s, respectively. The research outcomes demonstrated an output signal postprocessing case accelerating the hydrogen production without changing geometry, number of cells, and partial pressures.

08 HYDROGEN↗

Jefferson Laboratory C100 Superconducting Radio-Frequency Cavity Fault Data, 2020

The dataset was created to train machine learning models for the task of identifying the (1) cavity and (2) fault type from C100-type cryomodules at the Thomas Jefferson National Accelerator Facility (Jefferson Lab), thereby replacing the time-consuming efforts of a subject matter expert. Superconducting radio-frequency (SRF) cavity trips represent a significant source of accelerator downtime. Real-time – rather than post-mortem – identification of the offending cavity and classification of the fault type would give control room operators valuable feedback for corrective action planning. The anticipated benefit is increased beam-on-target time for users and provides performance metrics that can be used to improve future cavity designs. A series of 17 RF signals are recorded for each of the 8 cavities in a C100 cryomodule every time a cavity trips. These time-series signals are written to file using a specially designed data acquisition system. The dataset represents fault events recorded during Continuous Electron Beam Accelerator Facility (CEBAF) beam operations between January 18, 2019 and March 9, 2020. The following filtering steps were applied to collected data; (1) only 4 of the 17 signals per cavity are retained (GMES, GASK, CRFP, DETA2) (2) only events with data from each of the eight cavities in the cryomodule are kept, (3) only events that were sampled at 5 kHz were kept, (4) events from cryomodule 0L04 were neglected, (5) events occurring between February 4, 3PM and February 5, 12PM were neglected. As a result of preprocessing, the dataset is comprised of 2,375 unique events. The full dataset is comprised of three files: features.csv, cavity_labels.csv, fault_labels.csv. Each instance in faults.csv includes a timestamp (“date_time”), a label for the cryomodule which experienced the trip (“zone_label”), and 192 features (“feature_1”, “feature_2”... “feature_192”). The features correspond to 6 autoregressive features for each of 4 signals per cavity for each of the 8 cavities (6 × 4 signals/cavity × 8 cavities/cryomodule = 192). To deal with the large variation of signal amplitudes, time-series standardization via the z-score (standard score) function was applied prior to computing the features. For each instance, there is an associated label for the (1) cavity which faulted first (cavity_labels.csv) and (2) the type of fault that caused the trip (fault_labels.csv). The cavity identification can take values of [0, 1, 2, 3, 4, 5, 6, 7, 8] and the fault type can take values of [‘Microphonics’, ‘Quench_100ms’, ‘Controls_Fault’, ‘E_Quench’, ‘Quench_3ms’, ‘Single_Cav_Turn_Off’ , ‘Heat_Riser_Choke’, ‘Multi_Cav_Turn_Off’].

43 PARTICLE ACCELERATORS↗

Dataset for 'Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning', Water 2022

This data package presents forcing data, model code, and model output for classical machine learning models that predict monthly stream water temperature as presented in the manuscript ‘Stream Temperature Predictions for River Basin Management in the Pacific Northwest and Mid-Atlantic Regions Using Machine Learning’, Water (Weierbach et al., 2022). Specifically, for input forcing datasets we include two files each generated using the BASIN-3D data integration tool (Varadharajan et al., 2022) for stations in the Pacific Northwest and Mid Atlantic Hydrologic regions. Model code (written in python with the use of jupyter notebooks) includes codes for data preprocessing, training Multiple Linear Regression, Support Vector Regression, and Extreme Gradient Boosted Tree models, and additional notebooks for analysis of model output. We include specific model output files which represent modeling configurations presented in the manuscript also presented in an hdf5 format. Together, these data make up the workflow for predictions across three scenarios (single station, regional, and predictions in unmonitored basins) presented in the manuscript and allow for reproducibility of modeling procedures.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with the manuscript "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning"

This package contains the data and scripts used in "Encoding Diel Hysteresis and the Birch Effect in Dryland Soil Respiration Models through Knowledge-Guided Deep Learning" (Jiang et al., 2022). The data.zip file contains the flux tower and automated chamber observations used for developing the deep learning model for modeling soil respiration. The scripts.zip file contains the Jupyter notebooks and python scripts for preprocessing the data, training the deep learning models, and postprocessing the results. The src.zip contains the source code for training the deep learning model, performing mutual information analysis, and plotting functions. The trained_models.zip contains multiple folders used for hosting the trained deep-learning models and the associated soil respiration predictions. The whole process is performed using python. We include the REAMD.md to document the python package requirements.Soil respiration in dryland ecosystems is challenging to model due to its complex interactions with environmental drivers. Knowledge-guided deep learning provides a much more effective means of accurately representing these complex interactions than traditional Q10-based models. Mutual information analysis revealed that future soil temperature shares more information with soil respiration than past soil temperature, consistent with their clockwise diel hysteresis. We explicitly encoded diel hysteresis, soil drying, and soil rewetting effects on soil respiration dynamics in a newly designed Long Short Term Memory (LSTM) model. The model takes both past and future environmental drivers as inputs to predict soil respiration. The new LSTM model substantially outperformed three Q10-based models and the Community Land Model when reproducing the observed soil respiration dynamics in a semi-arid ecosystem. The new LSTM model clearly demonstrated its superiority for temporally extrapolating soil respiration dynamics, such that the resulting correlation with observational data is up to 0.7 while the correlations of both Q10-based models and the Community Land Model (CLM) are less than 0.4. Our results underscore the high potential for knowledge-guided deep learning to replace Q10-based soil respiration modules in Earth system models.

54 ENVIRONMENTAL SCIENCES↗

Classification of River Catchments in the Contiguous United States: Code, Dataset, Similarity Patterns, and Resulting Classes

This dataset serves as supplementary information for the paper by Ciulla F. and Varadharajan C. A Network Approach for Multiscale Catchment Classification using Traits (see reference 1). It contains environmental and physical catchment traits, such as temperatures, precipitation, land use and human interference, from 9067 sites across the contiguous United States (CONUS). The purpose of this dataset is to provide information for a better trait-based categorization of river catchments in the CONUS using networks as an analytical tool. The traits variables match the ones present in the GAGES-II dataset and the preprocessing steps are described in the Methods section (processed_dataset.csv). Additionally we include the topologies (nodes, edges and clusters, also referred as classes) of the catchment network and traits network generated by said dataset (csv and json files). A series of tables support the information carried by the network providing more detailed descriptions of cluster components (SI1.pdf). A summary of all the plots of clusters of catchments with at least 50 nodes is provided (SI2.pdf). The characteristic traits for each cluster of catchments is presented as z-score (traits_categories_zscores_per_catchment_class.csv). The link to the hydrological behavior of clusters of catchments is displayed by boxplots, each describing a particular river discharge index (SI3.pdf). Both csv and json files can be read by common text editors but the data contained into them can be better handled using programming languages like python and database oriented libraries like pandas. Pdf files can be read by any pdf reader software.[02-23-2024] Update: The code and datasets necessary to reproduce the results of the study are available as a zipped repository (code_datasets_catchments_similarity.zip).

54 ENVIRONMENTAL SCIENCES↗