Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Citation network datasets for benchmarking spiking graph neural networks on experimental neuromorphic hardware

Spiking neural networks (SNNs) running on neuromorphic computers offer an energy-efficient alternative for AI tasks. Recently, spiking graph neural networks (S-GNNs) have been shown to produce encouraging results on benchmark citation network datasets such as Cora, CiteSeer, and PubMed for node classification tasks. These S-GNNs were run on SNN simulators only because they contain up to tens of thousands of neurons and up to millions of synapses, translating poorly to neuromorphic hardware. Therefore, in this paper, we create a suite of benchmark datasets from the CiteSeer dataset that can be accommodated on current neuromorphic hardware platforms. Our contribution consists of a collection of three datasets. First, we have an induced subgraph of CiteSeer, which we call MiniSeer, containing 2110 papers, 3604 binary features, and 6 topics. Second, MicroSeer is a very small dataset consisting of 84 papers, 1227 features, and 6 topics. Lastly, BiteSeer is a collection of 15 binary classification datasets. We present creation of these datasets along with accuracies, running times, and spike counts when simulated. We believe that our results in this paper will be used by the neuromorphic community to benchmark, test, and develop neuromorphic hardware and simulators.

Zhu, Kevin [George Mason University, Virginia]↗

Open Power System Datasets and Open Simulation Engines: A Survey Toward Machine Learning Applications

A major factor behind the success of machine learning (ML) models in multiple domains is the availability and accessibility of large, labeled, and well-organized datasets for training and benchmarking. In comparison, power grid datasets face three major challenges: (i) real-world data is often restricted by regulatory constraints, privacy reasons, or security concerns, making it difficult to obtain and work with; (ii) synthetic datasets, which are created to address these limitations, often have incomplete information and are released using specialized tools, making them inaccessible to the broader community; and, (iii) input-output datasets are difficult to generate through simulation for non-experts because open-source simulators are not known outside the power system community. This survey addresses these challenges by serving as an entry point to publicly available datasets and simulators for researchers venturing in this area. We review the current landscape of open-source power network data, machine models, consumer demand profiles, renewable generation data, and inverter models. We also examine open-source power system simulators, which are crucial for generating high-quality, high-fidelity power grid datasets. We aim to provide a foundation for overcoming data scarcity and advance towards a structured web of datasets and simulators to support the development of ML for power systems.

42 ENGINEERING↗

Layer-wise Imaging Dataset from Powder Bed Additive Manufacturing Processes for Machine Learning Applications (Peregrine v2022-10.1)

This release consists of six datasets which together include multi-modal layer-wise powder bed images from two different powder bed printing technologies. These datasets are designed primarily to facilitate the development and testing of new computer vision and machine learning based anomaly and defect detection algorithms. The authors provide both training data with corresponding ground truth pixel masks and evaluation data with corresponding baseline prediction pixel masks made by a trained neural network. The laser powder bed fusion (L-PBF) datasets are sourced from EOS M290 and AddUp FormUp 350 printers and the binder jet (BJ) dataset is sourced from an ExOne M-Flex printer. The materials represented in these datasets include 17-4 PH Stainless Steel, GammaPrint-700, Inconel 718, Maraging Steel, and H13 Steel. The sensor imaging modalities represented include visible-light (VL), temporally-integrated (i.e., long duration exposure) near-infrared (TI-NIR), and wide-band infrared (IR). To download the dataset: (1) Create a Globus account. (2) Create a Globus Endpoint on your computer. (3) Transfer the dataset from the OLCF DOI-DOWNLOADS Collection to your Collection. Common troubleshooting steps: (a) Confirm that the transfer is going from OLCF DOI-DOWNLOADS to your Collection. (b) Create an exception for Globus in your antivirus software so that it can create an Endpoint. (c) Manually create a Globus access directory (where the data will be downloaded) by going to the Preferences > Access tab.

36 MATERIALS SCIENCE↗

A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process

Recent progress in sensing techniques and data analytics tools have significantly accelerated the development of Wire Arc Additive Manufacturing (WAAM) systems. This data centric approach emphasizes leveraging available data throughout the production process to optimize performance. Integration of extensive data analysis provides the opportunity to improve precision, reduce waste, and enhance the quality of produced parts. This method relies on AI/ML models and optimization techniques, which are developed using the data collected from various sources, including in-situ sensors, ex-situ imaging, and manufacturing process parameters. The quality and diversity of this data, along with the alignment between different data streams (achieved through spatiotemporal registration) are critical for the successful development of AI/ML and optimization models. In this work, we present a spatiotemporally registered dataset generated during the WAAM process of deposition of a rectangular block. The dataset includes the comprehensive description of deposition process, process parameters, in-situ collected welding characteristics, acoustic data, and X-Ray Computed Tomography analysis data for the build. Dataset A Co-Registered In-Situ and Ex-Situ Dataset of Electrical, Acoustic, and CT Characteristics from Wire Arc Additive Manufacturing Process has arisen under UT-Battelle, LLC’s Prime Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy (DOE) to manage and operate the Oak Ridge National Laboratory. UT-Battelle, LLC will not assert any rights under United States law or under the Prime Contract it has in the dataset against any user of the dataset, including any copyrights or patent rights. UT-Battelle, LLC requests that attribution to the dataset is provided as academically appropriate.

42 ENGINEERING↗

FY2021 Derivative Datasets LDRD Report

In this pilot project, the feasibility around the creation, management, and deployment of derivative datasets as a potential new class of intellectual property was explored. The objective of these derivative datasets was to address environmental, industrial, and commercial challenges in a novel way that could provide additional impact based on the research performed at PNNL. As a result of efforts in FY2021, two new derivative datasets were created using newly created code that would enable the facile creation of additional derivative datasets of varying complexity and integration from pre-existing and published originating datasets. In addition, the process workflows were documented, detailing how the creation, management, and deployment of the derivative datasets could integrate with existing processes within PNNL. Finally, exploratory interviews with industry allowed for the scoping of future collaboration and commercialization opportunities around derivative datasets.

97 MATHEMATICS AND COMPUTING↗

WTK-LED: The WIND Toolkit Long-Term Ensemble Dataset

To satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have, we aimed at developing an updated version of the meteorological WIND Toolkit, named WIND Toolkit Long-term Ensemble Dataset (WTK-LED), which is a meteorological dataset providing time series every 5 min and 2 km, including model uncertainty of wind speed at every modeling grid point so that users are provided with a range of possible wind speeds every 2 km. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: 1) Numerical simulations covering the continental United States, Alaska, and Hawaii, with high-resolution data being available for 3 years (2018-2020). 2) Climate simulations from Argonne National Laboratories covering the North American continent, including Alaska, Canada, and most of Mexico and the Caribbean Islands. These simulations complement the new WTK-LED to offer a 4-km dataset covering 20 years, from 2001-2020. 3) Specific long-term,high-resolution offshore simulations have been conducted separately for the US coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind data set. This report focuses on a description of the land-based WTK-LED for CONUS, Hawaii, and Alaska, for the 3-year 2-km/5-min dataset and the 20-year 4-km/hourly dataset, as well as the uncertainty quantification method. We also provide limited validation results. Based on our results to date, we suggest use cases and applications for each dataset of the WTK-LED.

17 WIND ENERGY↗

WIND Toolkit Long-Term Ensemble Dataset

WIND Toolkit Long-term Ensemble Dataset (WTK-LED), an updated version of the meteorological WIND Toolkit, is a meteorological dataset providing high-resolution time series, including interannual variability and model uncertainty of wind speed at every modeling grid point to indicate ranges of possible wind speeds. The data were produced using the Weather Research and Forecasting Model (WRF). The vertical grid used in WTK-LED includes many vertical layers in the atmospheric boundary layer to provide information of atmospheric quantities across the rotor layer of utility scale and distributed wind turbines. The WTK-LED includes: (1) Numerical simulations of wind speed and other meteorological variables covering the contiguous United States (CONUS) and Alaska, with high-resolution (5-minute [min], 2-kilometer [km]) data for 3 years (2018-2020): WTK-LED CONUS, WTK-LED Alaska. (2) Climate simulations from Argonne National Laboratory covering North America, including Alaska, Canada, and most of Mexico and the Caribbean islands. These simulations complement the new WTK-LED to offer a 4-km, hourly dataset covering 20 years (2001-2020): WTK-LED Climate. (3) Specific long-term, high-resolution offshore simulations have been conducted separately for the U.S. coasts, Hawaii, and the Great Lakes, leading to the 2023 National Offshore Wind dataset: NOW-23. The data for Hawaii include land-based data and are part of WTK-LED Hawaii. Because the accuracy of simulations from a mesoscale model, such as WRF, varies depending on the location and weather situation, and can reach up to several m/s for wind speed, we provide simulated wind speed uncertainty estimates to the community to be used in conjunction with the deterministic model simulations. This dataset was developed to satisfy a wide group of stakeholders across various wind energy disciplines, including but not limited to stakeholders in the distributed and utility scale wind industry, the new emerging airborne wind energy field, grid integration, power systems modeling, environmental modeling, and researchers in academia, and to close some of the gaps that current public datasets have. Based on our validation results to date, we suggest use cases and applications for each dataset of the WTK-LED as shown in "WTK-LED Use Cases" resource below.

Array↗

Physics-Based Method for Generating Fully Synthetic IV Curve Training Datasets for Machine Learning Classification of PV Failures

Classification machine learning models require high-quality labeled datasets for training. Among the most useful datasets for photovoltaic array fault detection and diagnosis are module or string current-voltage (IV) curves. Unfortunately, such datasets are rarely collected due to the cost of high fidelity monitoring, and the data that is available is generally not ideal, often consisting of unbalanced classes, noisy data due to environmental conditions, and few samples. In this paper, we propose an alternate approach that utilizes physics-based simulations of string-level IV curves as a fully synthetic training corpus that is independent of the test dataset. In our example, the training corpus consists of baseline (no fault), partial soiling, and cell crack system modes. The training corpus is used to train a 1D convolutional neural network (CNN) for failure classification. The approach is validated by comparing the model’s ability to classify failures detected on a real, measured IV curve testing corpus obtained from laboratory and field experiments. Results obtained using a fully synthetic training dataset achieve identical accuracy to those obtained with use of a measured training dataset. When evaluating the measured data’s test split, a 100% accuracy was found both when using simulations or measured data as the training corpus. When evaluating all of the measured data, a 96% accuracy was found when using a fully synthetic training dataset. The use of physics-based modeling results as a training corpus for failure detection and classification has many advantages for implementation as each PV system is configured differently, and it would be nearly impossible to train using labeled measured data.

Hopwood, Michael W. (ORCID:0000000161901767)↗

Pavement condition and climatic data in southeast Texas: A dataset for evaluating flood impacts on pavement performance

Effective pavement maintenance is essential for economic stability, optimal network performance, and roadway safety. Achieving this requires thorough evaluation of pavement conditions, including structural integrity, surface roughness, and distress characteristics. Pavement performance indicators play a critical role in influencing vehicle safety and ride quality. Recent advances have emphasized the use of data-driven modeling to anticipate pavement behavior, with the goal of optimizing resource allocation and refining Maintenance and Rehabilitation (M&R) strategies through accurate condition assessment. A foundational requirement for these modeling efforts is the availability of standardized, high-quality datasets that can support robust and reproducible infrastructure analysis. This data article presents a comprehensive dataset assembled to facilitate pavement performance prediction, with a geographic focus on Southeast Texas, particularly the flood-vulnerable area of Beaumont. The dataset encompasses pavement and traffic attributes, meteorological records, flood simulation outputs, ground deformation measurements, and topographic indices, enabling detailed examination of both load-associated and non-load-associated degradation mechanisms. Data preprocessing was performed using ArcGIS Pro, Microsoft Excel, and Python to ensure consistency and usability in data-driven modeling applications, including machine learning workflows. Key contributions of this dataset include its utility in analyzing the climatic and environmental factors affecting pavement conditions, identifying critical predictive features, and enabling in-depth correlation analysis across diverse variables. By filling existing gaps in input variable selection resources, this dataset supports the development of predictive tools for estimating future maintenance demand and enhancing the resilience of pavement networks in flood-impacted areas. The resource highlights the importance of standardized datasets for advancing pavement management practices and provides a robust foundation for ongoing infrastructure performance modeling.

42 ENGINEERING↗

ResOpsUS, a dataset of historical reservoir operations in the contiguous United States

Abstract There are over 52,000 dams in the contiguous US ranging from 0.5 to 243 meters high that collectively hold 600,000 million cubic meters of water. These structures have dramatically affected the river dynamics of every major watershed in the country. While there are national datasets that document dam attributes, there is no national dataset of reservoir operations. Here we present a dataset of historical reservoir inflows, outflows and changes in storage for 679 major reservoirs across the US, called ResOpsUS. All of the data are provided at a daily temporal resolution. Temporal coverage varies by reservoir depending on construction date and digital data availability. Overall, the data spans from 1930 to 2020, although the best coverage is for the most recent years, particularly 1980 to 2020. The reservoirs included in our dataset cover more than half of the total storage of large reservoirs in the US (defined as reservoirs with storage greater 0.1 km 3 ). We document the assembly process of this dataset as well as its contents. Historical operations are also compared to static reservoir attribute datasets for validation.

54 ENVIRONMENTAL SCIENCES↗

Two excited-state datasets for quantum chemical UV-vis spectra of organic molecules

Abstract We present two open-source datasets that provide time-dependent density-functional tight-binding (TD-DFTB) electronic excitation spectra of organic molecules. These datasets represent predictions of UV-vis absorption spectra performed on optimized geometries of the molecules in their electronic ground state. The GDB-9-Ex dataset contains a subset of 96,766 organic molecules from the original open-source GDB-9 dataset. The ORNL_AISD-Ex dataset consists of 10,502,904 organic molecules that contain between 5 and 71 non-hydrogen atoms. The data reveals the close correlation between the magnitude of the gaps between the highest occupied molecular orbital (HOMO) and the lowest unoccupied molecular orbital (LUMO), and the excitation energy of the lowest singlet excited state energies quantitatively. The chemical variability of the large number of molecules was examined with a topological fingerprint estimation based on extended-connectivity fingerprints (ECFPs) followed by uniform manifold approximation and projection (UMAP) for dimension reduction. Both datasets were generated using the DFTB+ software on the “Andes” cluster of the Oak Ridge Leadership Computing Facility (OLCF).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Typical and extreme weather datasets for studying the resilience of buildings to climate change and heatwaves

We present unprecedented datasets of current and future projected weather files for building simulations in 15 major cities distributed across 10 climate zones worldwide. The datasets include ambient air temperature, relative humidity, atmospheric pressure, direct and diffuse solar irradiance, and wind speed at hourly resolution, which are essential climate elements needed to undertake building simulations. The datasets contain typical and extreme weather years in the EnergyPlus weather file (EPW) format and multiyear projections in comma-separated value (CSV) format for three periods: historical (2001–2020), future mid-term (2041–2060), and future long-term (2081–2100). The datasets were generated from projections of one regional climate model, which were bias-corrected using multiyear observational data for each city. The methodology used makes the datasets among the first to incorporate complex changes in the future climate for the frequency, duration, and magnitude of extreme temperatures. These datasets, created within the IEA EBC Annex 80 “Resilient Cooling for Buildings”, are ready to be used for different types of building adaptation and resilience studies to climate change and heatwaves.

54 ENVIRONMENTAL SCIENCES↗

Search for ultralight dark matter in the SuperMAG high-fidelity dataset

Ultralight dark matter, such as kinetically mixed dark-photon dark matter (DPDM) or axion-like-particle dark matter (axion DM), can source an oscillating magnetic-field signal at Earth’s surface. Previous work searched for this signal in a publicly available dataset of global magnetometer measurements maintained by the SuperMAG collaboration. This “low-fidelity” dataset reported measurements with a 1-min time resolution, allowing the search to set leading direct constraints on DPDM and axion DM with Compton frequencies f DM ≤ 1 / ( 1 min ) (corresponding to masses m DM ≤ 7 × 10 − 17 eV ). More recently, a dedicated experiment undertaken by the SNIPE Hunt collaboration has also searched for this same signal at higher frequencies f DM ≥ 0.5 Hz (or m DM ≥ 2 × 10 − 15 eV ). In this work, we search for this signal of ultralight DM in the SuperMAG “high-fidelity” dataset, which features a 1-sec time resolution, allowing us to probe the gap in parameter space between the low-fidelity dataset and the SNIPE Hunt experiment. The high-fidelity dataset exhibits lower geomagnetic noise than the low-fidelity dataset and features more data than the SNIPE Hunt experiment, making it a powerful probe of ultralight DM. Our search finds no robust DPDM or axion DM candidates. We set constraints on DPDM and axion DM parameter space for 10 − 3 Hz ≤ f DM ≤ 0.98 Hz (or 4 × 10 − 18 eV ≤ m DM ≤ 4 × 10 − 15 eV ). Our results are the leading direct constraints on both DPDM and axion DM in this mass range, and our DPDM constraint surpasses the leading astrophysical constraint in a narrow range around m A ′ ≈ 2 × 10 − 15 eV . Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data Fusion for the Development of a Multimodal Freight Transload Facilities Dataset in the U.S.

To withstand the growing demand of commodity volume and its strain on the transportation infrastructure, it is necessary to identify the flow of commodities by route and mode. However, a national multimodal freight routing model does not exist for the U.S. The development of such model requires multiple building blocks, such as virtual representations of roadway, railway, and waterway networks, transload facilities (TFs), and access/egress links. Most of these blocks have a robust database in the U.S., except for the TFs. Here, this paper presents the fusion of dispersed and heterogeneous representations of multimodal TFs into a single, comprehensive, geospatial freight TF dataset. The TF dataset is derived from several sources, including the U.S. Army Corps of Engineers Master Docks Plus, the National Transportation Atlas Database, the Intermodal Association of North America, industry publications, and other public information. First, individual datasets were queried and reconciled. A geocoding/reverse geocoding process was applied to get the best street address and latitude/longitude location for each terminal. Then, duplicate terminals were identified by a fuzzy match algorithm based on terminal name and location, and removed. Validation was performed by visual inspection of random facilities. The main contributions of this work are: a publicly available version of the TF dataset, including facility location and multimodal transfer capability of 9,003 facilities, and an enterprise-version with the same facilities but including commodity handling capabilities. The main purpose of developing the TF dataset is to inform multimodal routing algorithms. The proposed TF dataset allows for credibly modeling the multimodal transfer of commodities within shipment routes.

Commodity Routing↗

2021 Smoky Mountains Conference Data Challenge Synthetic-to-Real Domain Adaptation for Autonomous Driving Dataset

The dataset is comprised of both real and synthetic images from a vehicle's forward-facing camera. Each camera image is accompanied by a corresponding pixel-level semantic segmentation image (all files are .png files). In total, the dataset contains 5600 images in the training/validation set and 1400 images in the testing set. The training dataset contains mostly synthetic RGB images collected with a wide range of weather and lighting conditions using the CARLA simulator [1]. In addition, the training data also includes a small pre-selected subset of data from the Cityscapes training dataset – which is comprised of RGB-segmentation image pairs from driving scenarios in various European cities [2]. The testing data is split into three sets. The first set contains synthetic CARLA images with weather/lighting conditions that were not present in the training set. The second set is a subset of the Cityscapes testing dataset. Finally, the third set is an unknown testing set which will not be revealed to the participants until after the submission deadline. [1] Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017, October). CARLA: An open urban driving simulator. In Conference on robot learning (pp. 1-16). PMLR. [2] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... and Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213-3223).

99 GENERAL AND MISCELLANEOUS↗

DEEPEN 3D PFA Weights for Exploration Datasets in Magmatic Environments

DEEPEN stands for DE-risking Exploration of geothermal Plays in magmatic ENvironments. As part of the development of the DEEPEN 3D play fairway analysis (PFA) methodology for magmatic plays (conventional hydrothermal, superhot EGS, and supercritical), weights needed to be developed for use in the weighted sum of the different favorability index models produced from geoscientific exploration datasets. This GDR submission includes those weights. The weighting was done using two different approaches: one based on expert opinions, and one based on statistical learning. The weights are intended to describe how useful a particular exploration method is for imaging each component of each play type. They may be adjusted based on the characteristics of the resource under investigation, knowledge of the quality of the dataset, or simply to reduce the impact a single dataset has on the resulting outputs. Within the DEEPEN PFA, separate sets of weights are produced for each component of each play type, since exploration methods hold different levels of importance for detecting each play component, within each play type. The weights for conventional hydrothermal systems were based on the average of the normalized weights used in the DOE-funded PFA projects that were focused on magmatic plays. This decision was made because conventional hydrothermal plays are already well-studied and understood, and therefore it is logical to use existing weights where possible. In contrast, a true PFA has never been applied to superhot EGS or supercritical plays, meaning that exploration methods have never been weighted in terms of their utility in imaging the components of these plays. To produce weights for superhot EGS and supercritical plays, two different approaches were used: one based on expert opinion and the analytical hierarchy process (AHP), and another using a statistical approach based on principal component analysis (PCA). The weights are intended to provide standardized sets of weights for each play type in all magmatic geothermal systems. Two different approaches were used to investigate whether a more data-centric approach might allow new insights into the datasets, and also to analyze how different weighting approaches impact the outcomes. The expert/AHP approach involved using an online tool (https://bpmsg.com/ahp/) with built-in forms to make pairwise comparisons which are used to rank exploration methods against one-another. The inputs are then combined in a quantitative way, ultimately producing a set of consensus-based weights. To minimize the burden on each individual participant, the forms were completed in group discussions. While the group setting means that there is potential for some opinions to outweigh others, it also provides a venue for conversation to take place, in theory leading the group to a more robust consensus then what can be achieved on an individual basis. This exercise was done with two separate groups: one consisting of U.S.-based experts, and one consisting of Iceland-based experts in magmatic geothermal systems. The two sets of weights were then averaged to produce what we will from here on refer to as the "expert opinion-based weights," or "expert weights" for short. While expert opinions allow us to include more nuanced information in the weights, expert opinions are subject to human bias. Data-centric or statistical approaches help to overcome these potential human biases by focusing on and drawing conclusions from the data alone. More information on this approach along with the dataset used to produce the statistical weights may be found in the linked dataset below.

15 GEOTHERMAL ENERGY↗

QA/QC of the East River, Colorado, discharge and geochemical time series datasets (Almont, BCC, and Pump House) to be used for modeling of hydrogeochemical balance

The following datasets were QA/QC-ed (Quality Assurance/Quality Control): 1. Brush Creek Confluence (BCC) discharge data (from Helen Malenda, USGS, Colorado School of Mines), which were calculated using the pressure transducer data and rating curves. The original 15 min time series data were presented as mean daily discharge. 2. Almont discharge data from United States Geological Survey (USGS). The original data were in 15 min time intervals, and were averaged to mean daily discharge time series. 3. Pump House discharge data as mean daily discharge (downloaded from the SFA portal). 4. BCC and Pump House chemistry data from SFA data portal and/or original spreadsheets provided by Roelof Versteeg. The following challenging QA/QC problems of the datasets were resolved: Missing data with the duration of gaps up to >1 month; Duplicated dates; Anomalies and outliers of discharge and concentrations; Time stamps of measurements of the discharge and concentrations are not aligned (hydrogeochemical balance calculations require the timestamps to be aligned). All QA/QC-ed datasets are given as csv files. The csv files were prepared using the xts files with multiple worksheets, which are also included in the data packages. Figures of the QA/QC-ed datasets are given in the jpeg and pdf formats. The QA/QC-ed datasets have been used to quantify discharge and chemical concentrations in river water in order to understand riverine exports of water and dissolved constituents in the East River watershed. These datasets served as a basis in the presentation given by P. Fox et al. at the 2021 Goldschmidt Conference.

54 ENVIRONMENTAL SCIENCES↗

U.S. Freight Transload Facilities Dataset

The U.S. Freight Transload Facilities Dataset provides location information (latitude, longitude, zip, city, county, state)for more than 9,000 facilities across 50 U.S. States where freight may be transferred between waterways, railways, and roadways. The dataset lists the known modes and available direction(s) for freight transfers at each facility as of 2024. The U.S. Freight Transload Facilities dataset was built by mining and fusing several public sources, such as the USACE Master Docks Plus, the USDOT National Transportation Atlas Database (NTAD), files from the Intermodal Association of North America (IANA), and the industry publication Bulk Transloader. The dataset constitutes a key piece of a multimodal freight transportation network and routing algorithm developed by USACE-ERDC. The dataset is shared as a .csv file. The dataset is published for research purposes and should not be considered exhaustive or authoritative.

Peterson, Steven [ORNL] (ORCID:0000000287672998)↗