Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Advanced Terrestrial Simulator (ATS) evaluation dataset at 7 catchments across the continental United States

This dataset comprises of the input files and other files required for Advanced Terrestrial Simulator (ATS) simulations at 7 catchments across the continental United States. ATS is an integrated surface-subsurface hydrology model. We include Jupyter notebooks (within scripts folder) for individual catchments showing information (including data sources, river network, soil, geology, landuse types etc.) on preparing the machine readable input files. ATS observation output files are provided in the output folder. Figures and analyses (.xlsx sheets) are also provided. The catchments include, Taylor River Upstream (Colorado); (b) Cossatot River (Arkansas); (c) Panther Creek (Alabama); (d) Little Tennessee River (North Carolina and Georgia); (e) Mayo River (Virginia); (f) Flat Brook (New Jersey); (g) Neversink River headwaters (New York). Readme files are provided inside the directories providing more details. Files types include: .xml, .h5, .xlsx, .png, .ipynb, .py, .nc, .txt. All of the files types can be accessed by open source software, details on software requirements are following: .xml (any text editors including notepad and textedit), .h5 (in python using hdf libraries), .xlsx (WPS Office Spreadsheets, OpenOffice Calc, LibreOffice Calc, Microsoft Office etc.), .png (any image viewer), .ipynb (Jupyter notebook), .py (any text editors including notepad and textedit), .nc (using python or other open source software).

54 ENVIRONMENTAL SCIENCES↗

Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)

This data release provides all data and code used in the paper " "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)" to model stream temperature, evaluate, and assess results. The associated manuscript explores current open questions in prediction in ungauged and unmonitored basins concerning top-down versus bottom-up approaches, tradeoffs between data available and input requirements, and the appropriate representation of catchment attributes as inputs to deep learning models. Modeling was done primarily with long short-term memory (LSTM) models, and stream site coverage spans 1362 locations across the conterminous United States. The data is organized into these items items:Code repository and data for the paper " "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models" Willard et al. (2024)".Code: stream_temp_ml_regionalization.zip contains the code repositoryData to run the code: - data_dir.zip -- contains all files that should be moved to the "DATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- metadata_dir.zip -- contains all files that should be moved to the "METADATA_DIR" variable defined in the "set_env_vars.sh" script in the code repository- error_analysis_attribute_and_groundwater_dir.zip - workflows for the extended error analysis by stream attribute and groundwater influenceData produced by the code and used in the paper:- outputs_dir.zip - contains model output and results (outputs_dir/results), model weights (outputs_dir/models), and all other outputs used for the paper including feature importances.To cite this code, please use the following BibTeX or MLA entries:bibtex:@misc{willard2024streamdata, author = {Jared Willard and Fabio Ciulla and Helen Weierbach and Vipin Kumar and Charuleka Varadharajan}, title = {Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models"}, year = {2024}, doi = {10.15485/2448016}, publisher = {ESS-DIVE Repository}, url = {https://doi.org/10.15485/2448016}}MLA: Willard, Jared, et al. Dataset for "Evaluating Deep Learning Approaches for Predictions in Unmonitored Basins with Continental-scale Stream Temperature Models". 2024. ESS-DIVE Repository, doi:10.15485/2448016.

54 ENVIRONMENTAL SCIENCES↗

Dataset for Evaluation of Extreme Weather Impacts on Utility-Scale Photovoltaic Plant Performance in the United States

This dataset is a fusion of three data types (operations and maintenance tickets, weather data, and production data) that was used to support machine learning analysis and evaluation of drivers for low performance at photovoltaic (PV) sites during compound, extreme weather events. After being processed with machine learning, the data was used in the "Evaluation of Extreme Weather Impacts on Utility-scale Photovoltaic Plant Performance in the United States" manuscript. Additional details are captured in the associated manuscript.

AI↗

Development of a Benchmark Eddy Flux Evapotranspiration Dataset for Evaluation of Satellite-Driven Evapotranspiration Models Over the CONUS

A large sample of ground-based evapotranspiration (ET) measurements made in the United States, primarily from eddy covariance systems, were post-processed to produce a benchmark ET dataset. The dataset was produced primarily to support the intercomparison and evaluation of the OpenET satellite-based remote sensing ET (RSET) models and could also be used to evaluate ET data from other models and approaches. OpenET is a web-based service that makes field-delineated and pixel-level ET estimates from well-established RSET models readily available to water managers, agricultural producers, and the public. The benchmark dataset is composed of flux and meteorological data from a variety of providers covering native vegetation and agricultural settings. Flux footprint predictions were developed for each station and included static flux footprints developed based on average wind direction and speed, as well as dynamic hourly footprints that were generated with a physically based model of upwind source area. The two footprint prediction methods were rigorously compared to evaluate their relative spatial coverage. Data from all sources were post-processed in a consistent and reproducible manner including data handling, gap-filling, temporal aggregation, and energy balance closure correction. The resulting dataset included 243,048 daily and 5,284 monthly ET values from 194 stations, with all data falling between 1995 and 2021. We assessed average daily energy imbalance using 172 flux sites with a total of 193,021 days of data, finding that overall turbulent fluxes were understated by about 12% on average relative to available energy. Multiple linear regression analyses indicated that daily average latent energy flux may be typically understated slightly more than sensible heat flux. This dataset was developed to provide a consistent reference to support evaluation of RSET data being developed for a wide range of applications related to water accounting and water resources management at field to watershed scales.

54 ENVIRONMENTAL SCIENCES↗

Data assimilation and model evaluation experiment datasets

The Institute for Naval Oceanography, in cooperation with Naval Research Laboratories and universities, executed the Data Assimilation and Model Evaluation Experiment (DAMEE) for the Gulf Stream region during fiscal years 1991-1993. Enormous effort has gone into the preparation of several high-quality and consistent datasets for model initialization and verification. This paper describes the preparation process, the temporal and spatial scopes, the contents, the structure, etc., of these datasets. The goal of DAMEE and the need of data for the four phases of experiment are briefly stated. The preparation of DAMEE datasets consisted of a series of processes: (1) collection of observational data; (2) analysis and interpretation; (3) interpolation using the Optimum Thermal Interpolation System package; (4) quality control and re-analysis; and (5) data archiving and software documentation. The data products from these processes included a time series of 3D fields of temperature and salinity, 2D fields of surface dynamic height and mixed-layer depth, analysis of the Gulf Stream and rings system, and bathythermograph profiles. To date, these are the most detailed and high-quality data for mesoscale ocean modeling, data assimilation, and forecasting research. Feedback from ocean modeling groups who tested this data was incorporated into its refinement. Suggestions for DAMEE data usages include (1) ocean modeling and data assimilation studies, (2) diagnosis and theoretical studies, and (3) comparisons with locally detailed observations.

Lai, Chung-Cheng A.↗

A globally sampled high-resolution hand-labeled validation dataset for evaluating surface water extent maps

Effective monitoring of global water resources is increasingly critical due to climate change and population growth. Advancements in remote sensing technology, specifically in spatial, spectral, and temporal resolutions, are revolutionizing water resource monitoring, leading to more frequent and high-quality surface water extent maps using various techniques such as traditional image processing and machine learning algorithms. However, satellite imagery datasets contain trade-offs that result in inconsistencies in performance, such as disparities in measurement principles between optical (e.g., Sentinel-2) and radar (e.g., Sentinel-1) sensors and differences in spatial and spectral resolutions among optical sensors. Therefore, developing accurate and robust surface water mapping solutions requires independent validations from multiple datasets to identify potential biases within the imagery and algorithms. However, high-quality validation datasets are expensive to build, and few contain information on water resources. For this purpose, we introduce a globally sampled, high-spatial-resolution dataset labeled using 3 m PlanetScope imagery. Our surface water extent dataset comprises 100 images, each with a size of 1024×1024 pixels, which were sampled using a stratified random sampling strategy covering all 14 biomes. We highlighted urban and rural regions, lakes, and rivers, including braided rivers and coastal regions. We evaluated two surface water extent mapping methods using our dataset – Dynamic World, based on Sentinel-2, and the NASA IMPACT model, based on Sentinel-1. Dynamic World achieved a mean intersection over union (IoU) of 72.16 % and F1 score of 79.70 %, while the NASA IMPACT model had a mean IoU of 57.61 % and F1 score of 65.79 %. Performance varied substantially across biomes, highlighting the importance of evaluating models on diverse landscapes to assess their generalizability and robustness. Our dataset can be used to analyze satellite products and methods, providing insights into their advantages and drawbacks. Our dataset offers a unique tool for analyzing satellite products, aiding the development of more accurate and robust surface water monitoring solutions. The dataset can be accessed via https://doi.org/10.25739/03nt-4f29.

54 ENVIRONMENTAL SCIENCES↗

VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images

Images are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of 12 state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of 469K question8 answer pairs involving 30K images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images

Maruf, M [Virginia Tech, Blacksburg]↗

BCARS Simulated Phantom Dataset for Evaluation of Processing Pipelines

Broadband coherent anti-Stokes Raman scattering (BCARS) microscopy is a powerful label-free biological imaging technique, but the raw signal requires careful processing. The vibrationally resonant (Raman) fingerprint signal is usually small compared with instrumental noise sources and the nonresonant background (NRB) inherent in the BCARS signal. Fortunately, the NRB exhibits a systematic phase relationship with the coherent Raman response, acting as a heterodyne amplifier for the weak fingerprint signal. Due to this heterodyne effect, the Raman response can be recovered quantitatively and invariantly across different instruments, provided the NRB shape is known. Even with heterodyne amplification, the amplitudes of fingerprint signal components are often comparable to system noise. Singular value decomposition (SVD), which utilizes spatial information, is often employed for additional noise filtering. Consequently, finding optimal processing parameters to properly distinguish the NRB and Raman responses and suppress noise in the complex BCARS signal requires a reference system that realistically represents the spectral and spatial properties of BCARS signals obtained from biological samples. We present a digital tissue phantom that meets these criteria as a tool for testing candidate signal processing pipelines. The digital phantom is generated with simulated hyperspectral Raman images having system-specific noise and background characteristics. Here, we analyze phantom datasets with differing background and signal-to-noise conditions to evaluate their impact on the performance of multiple signal processing pipelines. Specifically, we investigate the application of a Butterworth filter-based routine to directly estimate the NRB from the BCARS signal. Additionally, we evaluate a Lorentzian wavelet transform as an alternative to the Hilbert transform for extracting the Raman spectrum from the BCARS signal. While we demonstrate this phantom for BCARS, it can be used for any spectroscopic Raman imaging approach.

Dixon, Jessica Z. [Georgia Institute of Technology↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

Diagnosis of convective organization and cold pools using ARM datasets and evaluation of a unified convection parameterization (UNICON)

Tropical thunderstorms often cluster together. Studies have suggested that the degree to which the tropical thunderstorms are clustered impacts Earth's energy balance and water cycle, as well as extreme precipitation events. However, the processes controlling the spatial distribution of the thunderstorms are poorly understood and are not properly represented in most computer models for weather and climate prediction. Under the goals of better understanding how convection organizes at the mesoscale and advancing the representation of mesoscale convective organization in global climate models, we i) objectively quantified the degrees of convective organization and diagnosed cold pool processes using ARM field campaign observations, ii) examined the organization processes in storm-resolving model simulations, and iii) evaluated the impacts of mesoscale convective organization in global model simulations. The project yielded a firm reference against which the global model representation of mesoscale convective organization and cold pools can be evaluated against and shed new light into the role of parameterized convective organization in the global model simulation of the basic state and variability. Our results revealed two distinct phases of convective clustering during the two-day rain episodes (Cheng et al. 2018) and a new mechanism through which vertical wind shear in the low-troposphere can aid convective organization over tropical oceans (Cheng et al. 2020). It was demonstrated that the interactive representation of cold pools and mesoscale convective organization is key for global models to successfully simulate both the mean state and intraseasonal variability in the tropics (Ahn et al. 2019; 2020).

54 ENVIRONMENTAL SCIENCES↗

Leveraging NREL's ResStock & ComStock Dataset to Evaluate Building Stock Electrification: Preprint

Residential and commercial buildings accounted for 40% of U.S. energy consumption in 2022 and represent a significant opportunity for decarbonization through energy efficiency and electrification, and for grid planning. Building stock energy modeling is a powerful tool that can evaluate what-if scenarios as utilities, municipalities, policymakers, building owners and others work towards equitable building decarbonization and climate goals. This presentation will highlight several high-impact use cases of the National Renewable Energy Laboratory (NREL)'s highly granular, bottom-up building stock energy modeling tools, ResStock and ComStock. These use cases cover a wide range of project scale, from neighborhood electrification analysis and municipality long-term energy planning, to state energy code development and national policy evaluation. This presentation will showcase specific real-world applications for which ResStock and ComStock have been utilized across the country, including California codes and standards cost-effectiveness analysis, New York City affordable housing electrification cost gap analysis, and California targeted electrification and gas decommissioning analysis. For each use case, this presentation will illustrate how ResStock and ComStock played a crucial role in accurately characterizing regional building stocks, providing discrete and aggregated end-use load shapes, and calculating lifecycle consumption, emissions, and costs for a variety of building electrification strategies and scenarios. Finally, this presentation will demonstrate how the data provided by ResStock and ComStock can help unlock significant outcomes for these use cases, including but not limited to, customer bill impact, incentive and program design, and energy equity analyses.

building stock modeling↗

Evaluation of remote sensing-based evapotranspiration products at low-latitude eddy covariance sites

Remote sensing-based evapotranspiration (ET) products have been evaluated primarily using data from northern middle latitudes; therefore, little is known about their performance at low latitudes. To address this bias, an evaluation dataset was compiled using eddy covariance data from 40 sites between latitudes 30° S and 30° N. The flux data were obtained from the emerging network in Mexico (MexFlux) and from openly available databases of FLUXNET, AsiaFlux, and OzFlux. This unique reference dataset was then used to evaluate remote sensing-based ET products in environments that have been underrepresented in earlier studies. The evaluated products were: MODIS ET (MOD16, both the discontinued collection 5 (C5) and the latest collection (C6)), Global Land Evaporation Amsterdam Model (GLEAM) ET, and Atmosphere-Land Exchange Inverse (ALEXI) ET. Products were compared with unadjusted fluxes (ETorig) and with fluxes corrected for the lack of energy balance closure (ETebc). Three common statistical metrics were used: coefficient of determination (R2), root mean square error (RMSE), and percent bias (PBIAS). The effect of a vegetation mismatch between pixel and site on product evaluation results was investigated by examining the relationship between the statistical metrics and product-specific vegetation match indexes. Evaluation results of this study and those published in the literature were used to examine the performance of the products across latitudes. Differences between the MOD16 collection 5 and 6 datasets were generally smaller than differences with the other products. Performance and ranking of the evaluated products depended on whether ETorig or ETebc was used. When using ETorig, GLEAM generally had the highest R2, smallest PBIAS, and best RMSE values across the studied land cover types and climate zones. Neither MOD16 nor ALEXI performed consistently better than the other. When using ETebc, none of the products stood out in terms of both low bias and strong correlations. The use of ETebc instead of ETorig affected the biases more than the correlations. The product evaluation results showed no significant relationship with the degree of match between the vegetation at the pixel and site scale. The latitudinal comparison showed tendencies of lower R2 (all products) but better PBIAS and normalized RMSE values (MOD16 and GLEAM) for forests at low latitudes than for forests at northern middle latitudes. For non-forest vegetation, the products showed no clear latitudinal differences in performance.

Diego Salazar-Martínez↗

HydraGNN_Predictive_GFM_2026 - Ensemble of predictive graph foundation models for atomistic materials modeling

This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0

36 MATERIALS SCIENCE↗

Comparison of Dust Optical Depth From Multi-Sensor Products and MONARCH (Multiscale Online Non-hydrostatic AtmospheRe CHemistry) Dust Reanalysis Over North Africa, the Middle East, and Europe

Aerosol reanalysis datasets are model-based, observationally constrained, continuous 3D aerosol fields with a relatively high temporal frequency that can be used to assess aerosol variations and trends, climate effects, and impacts on socioeconomic sectors, such as health. Here we compare and assess the recently published MONARCH (Multiscale Online Non-hydrostatic AtmospheRe CHemistry) high-resolution regional desert dust reanalysis over northern Africa, the Middle East, and Europe (NAMEE) with a combination of ground-based observations and space-based dust retrievals and products. In particular, we compare the total and coarse dust optical depth (DOD) from the new reanalysis with DOD products derived from MODIS (MODerate resolution Imaging Spectroradiometer), MISR (Multi-angle Imaging SpectroRadiometer), and IASI (Infrared Atmospheric Sounding Interferometer) spaceborne instruments. Despite the larger uncertainties, satellite-based datasets provide a better geographical coverage than ground-based observations, and the use of different retrievals and products allows at least partially overcoming some single-product weaknesses in the comparison. Nevertheless, limitations and uncertainties due to the type of sensor, its operating principle, its sensitivity, its temporal and spatial resolution, and the methodology for retrieving or further deriving dust products are factors that bias the reanalysis assessment. We, therefore, also use ground-based DOD observations provided by 238 stations of the AERONET (AErosol RObotic NETwork) located within the NAMEE region as a reference evaluation dataset. In particular, prior to the reanalysis assessment, the satellite datasets were evaluated against AERONET, showing moderate underestimations in the vicinities of dust sources and downwind regions, whereas small or significant overestimations, depending on the dataset, can be found in the remote regions. Taking these results into consideration, the MONARCH reanalysis assessment shows that total and coarse-DOD simulations are consistent with satellite- and ground-based data, qualitatively capturing the major dust sources in the area in addition to the dust transport patterns. Moreover, the MONARCH reanalysis reproduces the seasonal dust cycle, identifying the increased dust activity that occurred in the NAMEE region during spring and summer. The quantitative comparison between the MONARCH reanalysis DOD and satellite multi-sensor products shows that the reanalysis tends to slightly overestimate the desert dust that is emitted from the source regions and underestimate the transported dust over the outflow regions, implying that the model's removal of dust particles from the atmosphere, through deposition processes, is too effective. More specifically, small positive biases are found over the Sahara desert (0.04) and negative biases over the Atlantic Ocean and the Arabian Sea (−0.04), which constitute the main pathways of the long-range dust transport. Considering the DOD values recorded on average there, such discrepancies can be considered low, as the low relative bias in the Sahara desert (< 50 %) and over the adjacent maritime regions (< 100 %) certifies. Similarly, over areas with intense dust activity, the linear correlation coefficient between the MONARCH reanalysis simulations and the ensemble of the satellite products is significantly high for both total and coarse DOD, reaching 0.8 over the Middle East, the Atlantic Ocean, and the Arabian Sea and exceeding it over the African continent. Moreover, the low relative biases and high correlations are associated with regions for which large numbers of observations are available, thus allowing for robust reanalysis assessment.

Michail Mytilinaios↗

MOFLUX Intensified Soil Moisture Extremes Decrease Soil Organic Carbon Decomposition: Modeling Archive

This Modeling Archive is in support the publication “Intensified Soil Moisture Extremes Decrease Soil Organic Carbon Decomposition: A Mechanistic Modeling Analysis” (Liang et al., 2021). Here we provide model code, inputs, outputs and evaluation datasets for the Microbial ENzyme Decomposition (MEND) model for the Missouri Ozarks AmeriFlux eddy covariance measurement site (MOFLUX) near Ashland, Missouri USA. The MEND model was developed with explicit representation of microbial and enzyme pools to mechanistically simulate the role of microbial organisms and extracellular enzymes in soil organic carbon (SOC) decomposition. Long-term SOC dynamics under intensified moisture extremes are studied using the MEND model that is parameterized with 11 years of measurements from the MOFLUX forest. The model explicitly represents microbial dormancy and resuscitation, different types of SOC-degrading enzymes, and how they vary with changes in soil moisture (Wang et al. 2015, 2019). A combination of two levels of frequency and severity of soil moisture, as well as a control with normal interannual variability, are used to simulate a range of moisture scenarios over 100 years. The code of Microbial-ENzyme Decomposition (MEND) as well as the input and output data are included in the archive. A user’s manual (MEND_Readme.pdf) is included with instructions for compiling and running the model to simulate soil organic carbon decomposition under various moisture scenarios. This dataset contains the modelling archive contained within a compressed (*.zip) file, a file-level metadata file in comma separate (*.csv) format, and two instructional files in PDF (*.pdf) format.

54 ENVIRONMENTAL SCIENCES↗