Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Development of a Benchmark Eddy Flux Evapotranspiration Dataset for Evaluation of Satellite-Driven Evapotranspiration Models Over the CONUS

A large sample of ground-based evapotranspiration (ET) measurements made in the United States, primarily from eddy covariance systems, were post-processed to produce a benchmark ET dataset. The dataset was produced primarily to support the intercomparison and evaluation of the OpenET satellite-based remote sensing ET (RSET) models and could also be used to evaluate ET data from other models and approaches. OpenET is a web-based service that makes field-delineated and pixel-level ET estimates from well-established RSET models readily available to water managers, agricultural producers, and the public. The benchmark dataset is composed of flux and meteorological data from a variety of providers covering native vegetation and agricultural settings. Flux footprint predictions were developed for each station and included static flux footprints developed based on average wind direction and speed, as well as dynamic hourly footprints that were generated with a physically based model of upwind source area. The two footprint prediction methods were rigorously compared to evaluate their relative spatial coverage. Data from all sources were post-processed in a consistent and reproducible manner including data handling, gap-filling, temporal aggregation, and energy balance closure correction. The resulting dataset included 243,048 daily and 5,284 monthly ET values from 194 stations, with all data falling between 1995 and 2021. We assessed average daily energy imbalance using 172 flux sites with a total of 193,021 days of data, finding that overall turbulent fluxes were understated by about 12% on average relative to available energy. Multiple linear regression analyses indicated that daily average latent energy flux may be typically understated slightly more than sensible heat flux. This dataset was developed to provide a consistent reference to support evaluation of RSET data being developed for a wide range of applications related to water accounting and water resources management at field to watershed scales.

54 ENVIRONMENTAL SCIENCES↗

A dataset of eco-evidence tools to inform early-stage environmental impact assessments of hydropower development

The datasets described herein provide the foundation for a decision support prototype (DSP) toolkit aimed at assisting stakeholders in determining evidence of which aspects of river ecosystems have been impacted by hydropower. The DSP toolkit and its application are presented and described in the article “Evidence-based indicator approach to guide preliminary environmental impact assessments of hydropower development” [1]. Development of the DSP and the output for decision support centralize around 42 river function indicators describing the dimensionality of river ecosystems through six main categories: biota and biodiversity, water quality, hydrology, geomorphology, land cover, and river connectivity. Three main tools are represented in the DSP: A science-based questionnaire (SBQ), an environmental envelope model (EEM), and a river function linkage assessment tool (RFLAT). The SBQ is a structured survey-style questionnaire whose objective is to provide evidence of which indicators have been impacted by hydropower. Based on a global literature review, 140 questions were developed from general hypotheses regarding the impacts of dams on rivers. The EEM is a model to predict the likelihood of hydropower impacting indicators based on a several variables. The intended use of the EEM is for situations of new hydropower development where results of the SBQ are incomplete or highly uncertain. The EEM was developed through the compilation of a dataset containing attributes of dams, reservoirs, and geospatial information on environmental concerns, which was combined with data on ecological indicators documented at those sites through literature review. The model operates through 247 “envelopes” and weighting factors, representing the individual effect of each variable on each indicator, all available through spreadsheets. Finally, the RFLAT is a tool to examine causal relationships amongst indicators. Inter-indicator relationships were hypothesized based on literature review and summarized into node and edge datasets to represent the structure of a graphical network. Bayes theorem was used estimate conditional probabilities of inter-indicator relationships based on the output of the SBQ. Nodes and edges were imported into R programming environment to visualize ecological indicator networks. The datasets can be expanded upon and enriched with more detailed questions for the SBQ, building upon the EEM with to develop more sophisticated models, and identifying new relationships for the RFALT. Additionally, once the tools are applied to numerous hydropower developments, the output of the tools (e.g. evidence of impacted indicators) becomes a very useful dataset for meta-analyses of hydropower impacts.

13 HYDRO ENERGY↗

Detect and correct bias in multi-site neuroimaging datasets

The desire to train complex machine learning algorithms and to increase the statistical power in association studies drives neuroimaging research to use ever-larger datasets. The most obvious way to increase sample size is by pooling scans from independent studies. However, simple pooling is often ill-advised as selection, measurement, and confounding biases may creep in and yield spurious correlations. In this work, we combine 35,320 magnetic resonance images of the brain from 17 studies to examine bias in neuroimaging. In the first experiment, Name That Dataset, we provide empirical evidence for the presence of bias by showing that scans can be correctly assigned to their respective dataset with 71.5% accuracy. Given such evidence, we take a closer look at confounding bias, which is often viewed as the main shortcoming in observational studies. In practice, we neither know all potential confounders nor do we have data on them. Hence, we model confounders as unknown, latent variables. Kolmogorov complexity is then used to decide whether the confounded or the causal model provides the simplest factorization of the graphical model. Finally, we present methods for dataset harmonization and study their ability to remove bias in imaging features. In particular, we propose an extension of the recently introduced ComBat algorithm to control for global variation across image features, inspired by adjusting for unknown population stratification in genetics. Overall, our results demonstrate that harmonization can reduce dataset-specific information in image features. Further, confounding bias can be reduced and even turned into a causal relationship. However, harmonization also requires caution as it can easily remove relevant subject-specific information. Code is available at https://github.com/ai-med/Dataset-Bias.

42 ENGINEERING↗

A data-driven global soil heterotrophic respiration dataset and the drivers of its inter-annual variability

Soil heterotrophic respiration (SHR), one of the primary carbon fluxes from terrestrial ecosystems to the atmosphere, is important for carbon-climate feedbacks because of its sensitivity to available litter and soil carbon, climatic conditions, and nutrient availability. However, until recently limited SHR data were available, and most published global SHR estimates have either a short time span, coarse spatial resolution, or reply on overly-simple model formulations. To better understand and quantify the global distribution of SHR and its sensitivity to climate variability, we produced a new global SHR dataset using Random Forest algorithms, up-scaling 455 point data from the Global Soil Respiration Database (SRDB 4.0) with gridded fields of climatic, edaphic and productivity as explanatory variables. We estimated a global total SHR of 46.8 Pg C yr-1 over 1985-2013 (95% confidence interval: 38.6-56.3 Pg C yr-1), with a significant increasing trend of 0.03 Pg C yr-2 during this period. We found that the choice of soil moisture datasets contributes more to the difference among these data-driven SHR members rather than that of productivity, temperature and precipitation data sources. We also analyzed the influence of climatic variables on the inter-annual variability (IAV) of our SHR product. Water availability was the dominant driver of IAV at global scales, although the inferred sensitivity depends on the choice of the soil moisture gridded dataset. At the ecosystem scale, temperature strongly controls the IAV of SHR in tropical forests, while water availability dominates in extra-tropical forest and semi-arid regions. Our machine-learning gridded SHR dataset and outputs from process-based land surface models (TRENDYv6) show agreement for a strong association between water variability and SHR IAV at the global scale, but the two approaches lead to different temporal trend globally and different controlling variables for IAV at the ecosystem scale. Our study provides evidence for the pervasive and important role of water availability in driving SHR, indicating both a direct effect limiting decomposition rates and an indirect effect through the amount of fresh organic matter made available to SHR from productivity. In consideration of potential limitations and uncertainties remaining in our data-driven SHR datasets, we call for a more scientifically designed observation network for SHR, more observation data compilation, and increased use of deep learning methods making maximum use of observation data in hand. This will benefit process-based models, and improve our understanding of SHR response to future anomalous environmental conditions.

Yao, Yitong↗

Small-wedge synchrotron and serial XFEL datasets for Cysteinyl leukotriene GPCRs

Structural studies of challenging targets such as G protein-coupled receptors (GPCRs) have accelerated during the last several years due to the development of new approaches, including small-wedge and serial crystallography. Here, we describe the deposition of seven datasets consisting of X-ray diffraction images acquired from lipidic cubic phase (LCP) grown microcrystals of two human GPCRs, Cysteinyl leukotriene receptors 1 and 2 (CysLT 1 R and CysLT 2 R), in complex with various antagonists. Five datasets were collected using small-wedge synchrotron crystallography (SWSX) at the European Synchrotron Radiation Facility with multiple crystals under cryo-conditions. Two datasets were collected using X-ray free electron laser (XFEL) serial femtosecond crystallography (SFX) at the Linac Coherent Light Source, with microcrystals delivered at room temperature into the beam within LCP matrix by a viscous media microextrusion injector. All seven datasets have been deposited in the open-access databases Zenodo and CXIDB. Here, we describe sample preparation and annotate crystallization conditions for each partial and full datasets. We also document full processing pipelines and provide wrapper scripts for SWSX and SFX data processing. A Correction to this paper has been published: https://doi.org/10.1038/s41597-020-00759-w

97 MATHEMATICS AND COMPUTING↗

The FLUXNET2015 dataset and the ONEFlux processing pipeline for eddy covariance data

The FLUXNET2015 dataset provides ecosystem-scale data on CO2, water, and energy exchange between the biosphere and the atmosphere, and other meteorological and biological measurements, from 212 sites around the globe (over 1500 site-years, up to and including year 2014). These sites, independently managed and operated, voluntarily contributed their data to create global datasets. Data were quality controlled and processed using uniform methods, to improve consistency and intercomparability across sites. The dataset is already being used in a number of applications, including ecophysiology studies, remote sensing studies, and development of ecosystem and Earth system models. FLUXNET2015 includes derived-data products, such as gap-filled time series, ecosystem respiration and photosynthetic uptake estimates, estimation of uncertainties, and metadata about the measurements, presented for the first time in this paper. In addition, 206 of these sites are for the first time distributed under a Creative Commons (CC-BY 4.0) license. This paper details this enhanced dataset and the processing methods, now made available as open-source codes, making the dataset more accessible, transparent, and reproducible.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Large Ensemble Global Dataset for Climate Impact Assessments

We present a self-consistent, large ensemble, high-resolution global dataset of long-term future climate, which accounts for the uncertainty in climate system response to anthropogenic emissions of greenhouse gases and in geographical patterns of climate change. The dataset is developed by applying an integrated spatial disaggregation (SD) - bias-correction (BC) method to climate projections from the MIT Integrated Global System Model (IGSM). Four emission scenarios are considered that represent energy and environmental policies and commitments of potential future pathways, namely, Reference, Paris Forever, Paris 2 °C and Paris 1.5 °C. The dataset contains nine key meteorological variables on a monthly scale from 2021 to 2100 at a spatial resolution of 0.5°x 0.5°, including precipitation, air temperature (mean, minimum and maximum), near-surface wind speed, shortwave and longwave radiation, specific humidity, and relative humidity. We demonstrate the dataset’s ability to represent climate-change responses across various regions of the globe. This dataset can be used to support regional-scale climate-related impact assessments of risk across different applications that include hydropower, water resources, ecosystem, agriculture, and sustainable development.

54 ENVIRONMENTAL SCIENCES↗

Dataset about Warming Effects on Carbon Cycling and Greenhouse Gas Fluxes in Permafrost Ecosystems

Field observations provide direct evidence of how does carbon cycling in permafrost ecosystems respond to climate change. This study provides a comprehensive dataset on the impact of warming on carbon cycling and greenhouse gas (GHG) fluxes in permafrost ecosystems. The dataset is extracted and integrated from 132 peer-reviewed studies with 1430 paired observations across eight major permafrost ecosystems, including Arctic and subarctic tundra and wetland, and alpine meadow, steppe, tundra and wetland. This dataset includes 17 variables from experiments conducted during the growing season, covering the plant and soil carbon pools, soil nitrogen pool, and GHG (i.e., CO 2 , CH 4 , and N 2 O) fluxes, among others. Background information on site climate conditions, vegetation and soil characteristics, and details of the warming experiments, including timing, methods, and warming magnitude, are also contained in the dataset. This dataset facilitates a comprehensive understanding of the impact of warming on carbon cycling and GHG fluxes in permafrost ecosystems, and provides supports for meta-analyses and literature reviews, remote sensing data validation, and land model development and parameterization.

Bao, Tao [Chinese Academy of Sciences (CAS), Beiji↗

Using multiple high-resolution datasets to benchmark the energy exascale earth system model (E3SM) for renewable resource assessment

The United States is accelerating its shift toward a renewable energy system. However, renewable resources, which harness energy from the Earth system, are susceptible to both present-day climate variability and future climate change. For example, variations in regional climate can alter renewable energy production patterns and site viability. The use of high-resolution climate model projections can therefore facilitate and may be critical to long-term planning of renewable energy investments. However, climate models must first be validated for renewable resource assessment. This research employs multiple high-spatiotemporal-resolution datasets to assess the capability of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 North American Regionally Refined Model (E3SMv2-NARRM) for predicting multi-year climatological values of solar and wind energy capacity factors in the continental U.S., with a focus on regional and seasonal variability. Present-day E3SMv2-NARRM simulations are compared with reported utility-scale production data obtained from the Energy Information Administration (EIA). In addition, E3SMv2-NARRM data are evaluated against non-climate benchmark models from the National Renewable Energy Laboratory, including the Wind Integration National Dataset Toolkit and the National Solar Radiation Database (NSRDB), as well as three wind energy datasets from PLUSWIND. Our analysis indicates that solar capacity factors from E3SM closely match those from the NSRDB dataset. However, both datasets tend to overestimate values by 10% in comparison to EIA data. Furthermore, biases in wind capacity factors within E3SM are notably pronounced in the West Coast regions, where the seasonal cycle diverges from EIA data.

Energy forecasting, Capacity factor, Renewable ene↗

PERSIANN Dynamic Infrared–Rain Rate (PDIR-Now): A Near-Real-Time, Quasi-Global Satellite Precipitation Dataset

This study presents the Precipitation Estimation from Remotely Sensed Information Using Artificial Neural Networks–Dynamic Infrared Rain Rate (PDIR-Now) near-real-time precipitation dataset. This dataset provides hourly, quasi-global, infrared-based precipitation estimates at 0.04° × 0.04° spatial resolution with a short latency (15–60 min). It is intended to supersede the PERSIANN–Cloud Classification System (PERSIANN-CCS) dataset previously produced as the near-real-time product of the PERSIANN family. We first provide a brief description of the algorithm’s fundamentals and the input data used for deriving precipitation estimates. Second, we provide an extensive evaluation of the PDIR-Now dataset over annual, monthly, daily, and subdaily scales. Last, the article presents information on the dissemination of the dataset through the Center for Hydrometeorology and Remote Sensing (CHRS) web-based interfaces. The evaluation, conducted over the period 2017–18, demonstrates the utility of PDIR-Now and its improvement over PERSIANN-CCS at all temporal scales. Specifically, PDIR-Now improves the estimation of rain/no-rain days as demonstrated by a critical success index (CSI) of 0.53 compared to 0.47 of PERSIANN-CCS. In addition, PDIR-Now improves the estimation of seasonal and diurnal cycles of precipitation as well as regional precipitation patterns erroneously estimated by PERSIANN-CCS. Finally, an evaluation is carried out to examine the performance of PDIR-Now in capturing two extreme events, Hurricane Harvey and a cluster of summer thunderstorms that occurred over the Netherlands, where it is shown that PDIR-Now adequately represents spatial precipitation patterns as well as subdaily precipitation rates with a correlation coefficient (CORR) of 0.64 for Hurricane Harvey and 0.76 for the Netherlands thunderstorms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Layer-wise Imaging Dataset from Powder Bed Additive Manufacturing Processes for Machine Learning Applications (Peregrine v2022-10)

This release consists of six datasets which together include multi-modal layer-wise powder bed images from two different powder bed printing technologies. These datasets are designed primarily to facilitate the development and testing of new computer vision and machine learning based anomaly and defect detection algorithms. The authors provide both training data with corresponding ground truth pixel masks and evaluation data with corresponding baseline prediction pixel masks made by a trained neural network. The laser powder bed fusion (L-PBF) datasets are sourced from EOS M290 and AddUp FormUp 350 printers and the binder jet (BJ) dataset is sourced from an ExOne M-Flex printer. The materials represented in these datasets include 17-4 PH Stainless Steel, DMREF, Inconel 718, Maraging Steel, and H13 Steel. The sensor imaging modalities represented include visible-light (VL), temporally-integrated (i.e., long duration exposure) near-infrared (TI-NIR), and wide-band infrared (IR).

36 MATERIALS SCIENCE↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

INGENIOUS - Great Basin Regional Dataset Compilation

This is the regional dataset compilation for the INnovative Geothermal Exploration through Novel Investigations Of Undiscovered Systems (INGENIOUS) project. The primary goal of this project is to accelerate discoveries of new, commercially viable hidden geothermal systems while reducing the exploration and development risks for all geothermal resources. These datasets will be used in INGENIOUS as input features for predicting geothermal favorability throughout the Great Basin study area. Datasets consist of shapefiles, geotiffs, tabular spreadsheets, and metadata that describe: 2-meter temperature probe surveys, quaternary faults and volcanic features, geodetic shear and dilation models, heat flow, magnetotellurics (conductance), magnetics, gravity, paleogeothermal features (such as sinter and tufa deposits), seismicity, spring and well temperatures, spring and well aqueous geochemistry analyses, thermal conductivity, and fault slip and dilation tendency. For additional project information, see the INGENIOUS project site linked in the submission. Terms of use: These datasets are provided "as is", and the contributors assume no responsibility for any errors or omissions. The user assumes the entire risk associated with their use of these data and bears all responsibility in determining whether these data are fit for their intended use. These datasets may be redistributed with attribution (see citation information below). Please refer to the license information on this page for full licensing terms and conditions.

15 GEOTHERMAL ENERGY↗

Mauka Energy FEVER Tool Dataset

Mauka Energy’s dataset, developed under the Forestry Electric Vehicle Energy Routing (FEVER) project and funded by the U.S. Department of Energy’s Small Business Innovation Research program, is a high-resolution geospatial resource designed to support energy modeling for electric log trucks in complex forestry environments. The dataset integrates detailed spatial and road network data to enable accurate simulation of vehicle performance across varied terrain. At its core, the dataset incorporates lidar-derived elevation models, road alignments, and surface classifications from Oregon State University’s McDonald-Dunn Research Forest. These data capture fine-scale variations in slope, curvature, and surface conditions across forest road systems, allowing for vehicle-level analysis of energy consumption and recovery. The dataset also includes data collected on the surrounding public and private road networks in Benton County, Oregon, used in real-world haul routes. These connecting segments provide critical context for modeling transitions between forest operations and regional transportation infrastructure, incorporating attributes such as grade profiles, elevation change, and speed constraints. This combined dataset underpins the development of Mauka Energy’s rolldown tool, which quantifies energy use and regenerative braking potential on downhill and variable-grade segments. By leveraging high-resolution terrain and road data, the FEVER project enables more accurate assessment of electric vehicle feasibility and performance in forestry applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Downscaled precipitation and mean air temperature datasets; East-Taylor subbasin; 2008-2019; daily temporal resolution; 400 m spatial resolution

This dataset provides gridded meteorological forcing data (specifically, daily precipitation and daily mean air temperature). The dataset has been generated by downscaling the Parameter-elevation Regressions on Independent Slopes Model (PRISM) dataset from a spatial resolution of 800 m to 400 m. The time period is 2008-2019, and the mapped area is East Taylor subbasin in Upper Colorado. The data are in the form of NetCDF files, arranged by year. Compared to the PRISM dataset, the dates of the downscaled dataset are shifted backwards by one (e.g., downscaled data for May 25 corresponds to PRISM data for May 26). This temporal shifting makes the meteorological forcing correspond more closely to the prescribed date. The NetCDF format is a standard raster format that can be read using any Geographic Information System (GIS) software; plenty of modules exist in popular scripting languages (such as Python, R and Matlab) that can also be used to read NetCDF files.East_Taylor_Tavg_PRISM400 corresponds to mean air temperature.East_Taylor_Precip_PRISM400.zip corresponds to precipitation.The proprietary PRISM data (800 m resolution) were purchased with funding from the Watershed Function Scientific Focus Area supported by U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research under award no. DE-AC02-05CH11231.Dataset update on March 9, 2022:Uploaded new data files that are identical to the earlier files but are now in the NetCDF format, arranged by year. Earlier, the files were in the GeoTiff format, arranged by date.

54 ENVIRONMENTAL SCIENCES↗

Classification of River Catchments in the Contiguous United States: Code, Dataset, Similarity Patterns, and Resulting Classes

This dataset serves as supplementary information for the paper by Ciulla F. and Varadharajan C. A Network Approach for Multiscale Catchment Classification using Traits (see reference 1). It contains environmental and physical catchment traits, such as temperatures, precipitation, land use and human interference, from 9067 sites across the contiguous United States (CONUS). The purpose of this dataset is to provide information for a better trait-based categorization of river catchments in the CONUS using networks as an analytical tool. The traits variables match the ones present in the GAGES-II dataset and the preprocessing steps are described in the Methods section (processed_dataset.csv). Additionally we include the topologies (nodes, edges and clusters, also referred as classes) of the catchment network and traits network generated by said dataset (csv and json files). A series of tables support the information carried by the network providing more detailed descriptions of cluster components (SI1.pdf). A summary of all the plots of clusters of catchments with at least 50 nodes is provided (SI2.pdf). The characteristic traits for each cluster of catchments is presented as z-score (traits_categories_zscores_per_catchment_class.csv). The link to the hydrological behavior of clusters of catchments is displayed by boxplots, each describing a particular river discharge index (SI3.pdf). Both csv and json files can be read by common text editors but the data contained into them can be better handled using programming languages like python and database oriented libraries like pandas. Pdf files can be read by any pdf reader software.[02-23-2024] Update: The code and datasets necessary to reproduce the results of the study are available as a zipped repository (code_datasets_catchments_similarity.zip).

54 ENVIRONMENTAL SCIENCES↗

Machine Learning meets Algebraic Combinatorics: A Suite of Benchmark Datasets to Accelerate AI for Mathematics Research

The use of benchmark datasets has become an important engine of progress in machine learning (ML) over the past 15 years. Recently there has been growing interest in utilizing machine learning to drive advances in research-level mathematics. However, off-the-shelf solutions often fail to deliver the types of insights required by mathematicians. This suggests the need for new ML methods specifically designed with mathematics in mind. The question then is: what benchmarks should the community use to evaluate these? On the one hand, toy problems such as learning the multiplicative structure of small finite groups have become popular in the mechanistic interpretability community whose perspective on explainability aligns well with the needs of mathematicians. While toy datasets are a useful benchmark for initial work, they lack the scale, complexity, and sophistication of many of the principal objects of study in modern mathematics. To address this, we introduce a new collection of benchmark datasets, Algebraic Combinatorics Benchmarks (ACBench), representing either classic or open problems in algebraic combinatorics, a subfield of mathematics that studies discrete structures arising from abstract algebra. After describing the datasets, we discuss the challenges involved in constructing “good” mathematics benchmarks, describe baseline model performance, and discuss some of the insights these datasets can provide that may be of interest even to those who are not interested in mathematics research itself.

97 MATHEMATICS AND COMPUTING↗

Comparison of Radiosonde Datasets: SondeHub and Integrated Global Radiosonde Archive

SondeHub aggregates radiosonde telemetry data uploaded from community-run radiosonde receiver stations. This radiosonde telemetry dataset is open-source, available to anyone through Amazon S3. There are also other public radiosonde datasets such as National Centers for Environmental Information (NCEI)’s Integrated Global Radiosonde Archive (IGRA). While there are many similarities between the two datasets, there are many differences as well due to the nature of the two datasets: one is community-run, while the other is managed by a government agency. This report presents the result of analyzing and comparing the two datasets.

54 ENVIRONMENTAL SCIENCES↗