Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Dataset about Warming Effects on Carbon Cycling and Greenhouse Gas Fluxes in Permafrost Ecosystems

Field observations provide direct evidence of how does carbon cycling in permafrost ecosystems respond to climate change. This study provides a comprehensive dataset on the impact of warming on carbon cycling and greenhouse gas (GHG) fluxes in permafrost ecosystems. The dataset is extracted and integrated from 132 peer-reviewed studies with 1430 paired observations across eight major permafrost ecosystems, including Arctic and subarctic tundra and wetland, and alpine meadow, steppe, tundra and wetland. This dataset includes 17 variables from experiments conducted during the growing season, covering the plant and soil carbon pools, soil nitrogen pool, and GHG (i.e., CO 2 , CH 4 , and N 2 O) fluxes, among others. Background information on site climate conditions, vegetation and soil characteristics, and details of the warming experiments, including timing, methods, and warming magnitude, are also contained in the dataset. This dataset facilitates a comprehensive understanding of the impact of warming on carbon cycling and GHG fluxes in permafrost ecosystems, and provides supports for meta-analyses and literature reviews, remote sensing data validation, and land model development and parameterization.

Bao, Tao [Chinese Academy of Sciences (CAS), Beiji

Using multiple high-resolution datasets to benchmark the energy exascale earth system model (E3SM) for renewable resource assessment

The United States is accelerating its shift toward a renewable energy system. However, renewable resources, which harness energy from the Earth system, are susceptible to both present-day climate variability and future climate change. For example, variations in regional climate can alter renewable energy production patterns and site viability. The use of high-resolution climate model projections can therefore facilitate and may be critical to long-term planning of renewable energy investments. However, climate models must first be validated for renewable resource assessment. This research employs multiple high-spatiotemporal-resolution datasets to assess the capability of the Department of Energy’s (DOE) Energy Exascale Earth System Model version 2 North American Regionally Refined Model (E3SMv2-NARRM) for predicting multi-year climatological values of solar and wind energy capacity factors in the continental U.S., with a focus on regional and seasonal variability. Present-day E3SMv2-NARRM simulations are compared with reported utility-scale production data obtained from the Energy Information Administration (EIA). In addition, E3SMv2-NARRM data are evaluated against non-climate benchmark models from the National Renewable Energy Laboratory, including the Wind Integration National Dataset Toolkit and the National Solar Radiation Database (NSRDB), as well as three wind energy datasets from PLUSWIND. Our analysis indicates that solar capacity factors from E3SM closely match those from the NSRDB dataset. However, both datasets tend to overestimate values by 10% in comparison to EIA data. Furthermore, biases in wind capacity factors within E3SM are notably pronounced in the West Coast regions, where the seasonal cycle diverges from EIA data.

Energy forecasting, Capacity factor, Renewable ene

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES

Mauka Energy FEVER Tool Dataset

Mauka Energy’s dataset, developed under the Forestry Electric Vehicle Energy Routing (FEVER) project and funded by the U.S. Department of Energy’s Small Business Innovation Research program, is a high-resolution geospatial resource designed to support energy modeling for electric log trucks in complex forestry environments. The dataset integrates detailed spatial and road network data to enable accurate simulation of vehicle performance across varied terrain. At its core, the dataset incorporates lidar-derived elevation models, road alignments, and surface classifications from Oregon State University’s McDonald-Dunn Research Forest. These data capture fine-scale variations in slope, curvature, and surface conditions across forest road systems, allowing for vehicle-level analysis of energy consumption and recovery. The dataset also includes data collected on the surrounding public and private road networks in Benton County, Oregon, used in real-world haul routes. These connecting segments provide critical context for modeling transitions between forest operations and regional transportation infrastructure, incorporating attributes such as grade profiles, elevation change, and speed constraints. This combined dataset underpins the development of Mauka Energy’s rolldown tool, which quantifies energy use and regenerative braking potential on downhill and variable-grade segments. By leveraging high-resolution terrain and road data, the FEVER project enables more accurate assessment of electric vehicle feasibility and performance in forestry applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Machine Learning meets Algebraic Combinatorics: A Suite of Benchmark Datasets to Accelerate AI for Mathematics Research

The use of benchmark datasets has become an important engine of progress in machine learning (ML) over the past 15 years. Recently there has been growing interest in utilizing machine learning to drive advances in research-level mathematics. However, off-the-shelf solutions often fail to deliver the types of insights required by mathematicians. This suggests the need for new ML methods specifically designed with mathematics in mind. The question then is: what benchmarks should the community use to evaluate these? On the one hand, toy problems such as learning the multiplicative structure of small finite groups have become popular in the mechanistic interpretability community whose perspective on explainability aligns well with the needs of mathematicians. While toy datasets are a useful benchmark for initial work, they lack the scale, complexity, and sophistication of many of the principal objects of study in modern mathematics. To address this, we introduce a new collection of benchmark datasets, Algebraic Combinatorics Benchmarks (ACBench), representing either classic or open problems in algebraic combinatorics, a subfield of mathematics that studies discrete structures arising from abstract algebra. After describing the datasets, we discuss the challenges involved in constructing “good” mathematics benchmarks, describe baseline model performance, and discuss some of the insights these datasets can provide that may be of interest even to those who are not interested in mathematics research itself.

97 MATHEMATICS AND COMPUTING

Comparison of Radiosonde Datasets: SondeHub and Integrated Global Radiosonde Archive

SondeHub aggregates radiosonde telemetry data uploaded from community-run radiosonde receiver stations. This radiosonde telemetry dataset is open-source, available to anyone through Amazon S3. There are also other public radiosonde datasets such as National Centers for Environmental Information (NCEI)’s Integrated Global Radiosonde Archive (IGRA). While there are many similarities between the two datasets, there are many differences as well due to the nature of the two datasets: one is community-run, while the other is managed by a government agency. This report presents the result of analyzing and comparing the two datasets.

54 ENVIRONMENTAL SCIENCES

PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE

This dataset is an open-source repository of spectral data measured at Pacific Northwest National Laboratory (PNNL). This database provides quantitative values for the complex index of refraction for seven polycyclic aromatic hydrocarbon (PAH) solids. A list of the chemicals is available in the readme file. These spectra consist of the optical constants, i.e., the real, n(ν), and imaginary, k(ν), refractive indices, over the spectral range from 7,800 to 400 cm-1 (1.28 – 25 μm). The conditions under which the individual data were acquired are described in the associated metadata files, and the user is strongly encouraged to read and understand this information to ensure the data are used appropriately for your application. Recommended Citation for Dataset Jessica M Salcido, Jeremy D. Erickson, Ashley M. Bradley, Russell G. Tonkyn, Timothy J. Johnson and Tanya L. Myers. 2026. PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE. [Data Set] PNNL DataHub. INSERT DOI License Information This work is marked with CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/. The authors do request that you appropriately cite the dataset when referencing or using the dataset.

Salcido, Jessica Marie Ortola

Discrete global grid system-based flow routing datasets in the Amazon and Yukon basins

Abstract. Discrete global grid systems (DGGS) are emerging spatial data structures widely used to organize geospatial datasets across scales. While DGGS have found applications in various scientific disciplines, including atmospheric science and ecology, their integration into physically based hydrological models and Earth system models (ESMs) has been hindered by the lack of flow routing datasets based on DGGS. In response to this gap, this study pioneers the development of new flow routing datasets using icosahedral Snyder equal-area (ISEA) DGGS and a novel mesh-independent flow direction model. We present flow routing datasets for two large basins, the tropical Amazon River basin and the Arctic Yukon River basin. These datasets (1) facilitate the adoption of DGGS for hydrological models and (2) provide flow routing inputs for evaluation of DGGS-based flow routing in the Amazon and Yukon river basins. The data are available at https://doi.org/10.5281/zenodo.8377765 (Liao, 2023).

54 ENVIRONMENTAL SCIENCES

U-Surf: a global 1 km spatially continuous urban surface property dataset for kilometer-scale urban-resolving Earth system modeling

High-resolution urban climate modeling has faced substantial challenges due to the absence of a globally consistent, spatially continuous, and accurate dataset to represent the spatial heterogeneity of urban surfaces and their biophysical properties. This deficiency has long obstructed the development of urban-resolving Earth system models (ESMs) and ultra-high-resolution urban climate modeling, over large domains. Here, we present U-Surf, a first-of-its-kind 1 km resolution present-day (circa 2020) global continuous urban surface parameter dataset. Using the urban canopy model (UCM) in the Community Earth System Model as a base model for satisfying dataset requirements, U-Surf leverages the latest advances in remote sensing, machine learning, and cloud computing to provide the most relevant urban surface biophysical parameters, including radiative, morphological, and thermal properties, for UCMs at the facet and canopy level. Generated using a systematically unified workflow, U-Surf ensures internal consistency among key parameters, making it the first globally coherent urban canopy surface dataset. U-Surf significantly improves the representation of the urban land heterogeneity both within and across cities globally; provides essential, high-fidelity surface biophysical constraints to urban-resolving ESMs; enables detailed city-to-city comparisons across the globe; and supports next-generation kilometer-resolution Earth system modeling across scales. U-Surf parameters can be easily converted or adapted to various types of UCMs, such as those embedded in weather and regional climate models, as well as air quality models. The fundamental urban surface constraints provided by U-Surf can also be used as features for machine learning models and can have other broad-scale applications for socioeconomic, public health, and urban planning contexts. We expect U-Surf to advance the research frontier of urban system science, climate-sensitive urban design, and coupled human–Earth systems in the future. The dataset is publicly available at https://doi.org/10.5281/zenodo.11247598 (Cheng et al., 2024).

Cheng, Yifan [Univ. of Illinois at Urbana-Champaig

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat

Legacy Survey of Space and Time Data Preview 1: calibrations dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the calibrations dataset type. These are a collection of calibration datasets such as biases, darks, and flats used to construct the data release. This release contains 496 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS

Reference Site Condition Datasets for Floating Wind Arrays in the United States

Floating offshore wind farm design is highly site-specific, requiring detailed information about the specific conditions of a project area for realistic design studies. Unfortunately, publicly available site condition data for potential floating offshore wind project sites in the United States is scarce. To support U.S. offshore wind research, we developed reference site condition datasets, including metocean and seabed information, for four potential floating wind project areas in the U.S.: Humboldt Bay, Morro Bay, the Gulf of Maine, and the Gulf of Mexico. These datasets were compiled using publicly available data. Our metocean analysis, covering wind, waves, and surface currents, utilized measurement data from 2000 to 2020. Sources included the National Renewable Energy Laboratory’s National Offshore Wind Dataset for wind data, National Data Buoy Center buoys for wave data, and the High Frequency Radar Network for surface currents. These data were integrated into hourly time series used to compute extreme return periods up to 500 years, monthly statistics, and joint probability clusters for fatigue analysis. Soil conditions were evaluated using the usSEABED database and bathymetry grids were interpolated from the NCEI Digital Elevation Model Global Mosaic. Further information on the datasets and how they were created can be found in: Biglu, M., M. Hall, E. Lozon, S. Housner. 2024. Reference Site Conditions for Floating Wind Arrays in the United States. Golden, CO: National Renewable Energy Laboratory (NREL). NREL/TP-5000-89897. The data are also available at: https://github.com/FloatingArrayDesign/SiteConditions The content of each dataset is as follows: _NOW23_wind.txt: Hourly NOW-23 wind data up to a height of 400 meter. _metocean_1hr.txt: Hourly time series including wind, wave, surface current and temperature data. _Summary.xlsx: Metocean data, including extreme values, joint probability distributions and monthly statistics. _usSEABED_soil.csv: Extract of the usSEABED database for this specific site. _bathymetry_200m.txt (and 500m, 1000m): Gridded seabed depth data.

16 TIDAL AND WAVE POWER

Predicting cutoff L-shells of solar protons using the GPPSn particle dataset

Solar energetic protons (SEPs) arriving at the Earth trigger severe radiation storms in the near-Earth space, directly impacting space missions operating at various altitudes. Therefore, monitoring SEP events and predicting the penetration depths of solar protons are critical for aerospace sectors. Building on previous efforts, here we demonstrate the feasibility of using proton measurements from the Global Prompt Proton Sensor network (GPPSn), enabled by Los Alamos National Laboratory developed combined X-ray dosimeters aboard GPS satellites, to characterize and predict the penetration of solar protons into the geomagnetic field. The inclined medium-Earth-orbits (MEOs) of the global GPS constellation offer a unique advantage of allowing simultaneous measurements of penetrating solar protons inside both open- and closed-field line regions. Therefore, the L-profiles of ∼10s–100 MeV solar protons and their associated cutoff L-shells can be determined from the GPPSn dataset, using predefined threshold proton flux values rather than traditional flux ratios. After examining a list of SEP event intervals across solar cycles 23, 24 and 25—including the 2024 Mother’s Day superstorm, we showcase how the latest GPPSn proton dataset (release v1.10), reprocessed and calibrated, can not only be used to monitor solar proton distributions inside the dynamic geomagnetic field for individual events, but also to derive a new empirical model linking cutoff L-shells with several key space weather parameters. This newly developed SEPCL-MEO model demonstrates high predictive performance; for example, predictions for > 30 MeV solar protons yield a correlation coefficient of 0.85 and performance efficiency of 0.67 when validated against GPPSn observations. Results from this pilot study underscores the scientific and operational value of the GPPSn dataset, and this dataset—when paired with machine-learning techniques—can play a critical role in observing and predicting the effects of future incoming SEP events, including extreme ones.

58 GEOSCIENCES

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David

The high explosives & affected targets (HEAT) dataset

Artificial Intelligence (AI) surrogate models offer a computationally efficient alternative to full-physics simulations, yet no existing datasets are publicly available for training, testing, and validation of machine learning models of the dynamics of high-explosive driven shocks through multiple materials. Shock propagation through materials is a computationally challenging problem because simulations must include material-specific equations of state (EOS) along with descriptions of other physical processes such as plastic deformation, phase change, damage processes, fluid instabilities, and multi-material interactions. Shocks are typically initiated by high-velocity impacts or explosive loading. The latter case necessitates the addition of models of reactive materials to represent high-explosive (HE) detonation. Here, to address the lack of an expansive dataset for multi-material shock propagation in the AI/ML community, we present the High-Explosives and Affected Targets (HEAT) Dataset. HEAT is a physics-rich collection of two-dimensional, cylindrically symmetric, simulations generated using an Eulerian, multi-material, shock-propagation code developed at Los Alamos National Laboratory. The dataset includes two partitions: (1) the expanding shock-cylinder (CYL) simulations, Figs. 1, and (2) the Perturbed Layered Interface (PLI) simulations, Fig. 2. Entries in both partitions consist of time series of arrays of thermodynamic fields (pressure, density, and temperature), kinematic fields (position and velocity), and additional fields that depend on thermodynamic and/or kinematic fields (e.g., material stress). Materials in the CYL partition include solids (aluminium, copper, depleted uranium, stainless steel, tantalum, and a generic polymer), a liquid (water), gases (air, nitrogen), and a generic detonating material (high explosive, HE). The PLI partition spans a highly varying geometry but consists of fixed materials across entries: Copper, aluminium, stainless steel, generic polymer, and generic HE. HEAT captures critical phenomena such as momentum transfer, shock propagation, plastic deformation, and thermal effects, making HEAT a valuable benchmark for development of AI/ML emulation of multi-material shock propagation.

36 MATERIALS SCIENCE

Systematic Evaluation of Atmospheric Forcing, Surface Datasets, and Mesh Effects on Kilometer-Scale Land Surface and River Modeling

Earth system models are advancing toward kilometer-scale resolution to capture local climate impacts and extremes. High-resolution land and river modeling depends on multiple factors, including mesh, surface datasets, and atmospheric forcing, but their relative effects at kilometer scales remain unquantified. We evaluated five Energy Exascale Earth System Model land and river configurations over the Mid-Atlantic region using two mesh (1/8° structured versus variable-resolution unstructured mesh), two surface datasets (default versus newly developed), and three atmospheric forcings (NLDAS2, MSWX, GSWP). Evaluation against satellite, reanalysis, and in situ benchmarks across water, energy, and carbon cycles quantifies how these factors affect model performance. Forcing selection produces the largest bias reductions (12-99% across variables), followed by surface datasets (7-75%) and mesh (up to 21%). Forcing effects vary by variable, with MSWX reducing biases for snow water equivalent, evapotranspiration, albedo, temperature, and gross primary productivity, GSWP for snow cover and runoff, and NLDAS for soil moisture and streamflow. The use of newly developed surface datasets improves gross primary productivity (58% bias reduction) and evapotranspiration but increase soil moisture and albedo biases due to current modeling limitations. Variable-resolution unstructured mesh improves the simulation of small-basin streamflow through better capturing drainage networks, though mesh minimally affects other land variables. These findings provide important guidance for high-resolution modeling development and actionable science.

Land and River modeling

A global urban heat island intensity dataset: Generation, comparison, and analysis

The urban heat island (UHI) effect, a phenomenon of local warming over urban areas, is the most well-known impact of urbanization on climate. Globally consistent estimates of the UHI intensity (UHII) are crucial for examining this phenomenon across time and space. However, publicly available UHII datasets are limited and have several constraints: (1) they are for clear-sky surface UHII, not all-sky surface UHII and canopy (air temperature) UHII; (2) the estimation methods often neglect anthropogenic disturbance, introducing uncertainties in the estimated UHII. To address these issues, this study proposes a new dynamic equal-area (DEA) method that can minimize the influence of various confounding factors on UHII estimates through a dynamic cyclic process. Utilizing the DEA method and leveraging various gridded temperature data, we develop a global-scale (>10,000 cities), long-term (over 20 years by month), and multi-faceted (clear-sky surface, all-sky surface, and canopy) UHII dataset. Further, based on these estimates, we provide a comprehensive analysis of the UHII and its trends in global cities. The UHII is found to be greater than zero in >80% of cities, with global annual average magnitudes around 1.0 °C (day) and 0.8 °C (night) for surface UHII, and close to 0.5 °C for canopy UHII. Furthermore, an interannual upward trend in UHII is observed in >60% of cities, with global annual average trends exceeding 0.1 °C/decade (day) and over 0.06 °C/decade (night) for surface UHII, and slightly surpassing 0.03 °C/decade for canopy UHII. Notably, there exists a positive correlation between the magnitude and trend of UHII, suggesting that cities with stronger UHII tend to experience faster growth in UHII. Additionally, discrepancies in UHII are found between different temperature data, stemming not only from distinctions in data types (surface or air temperature) but also from differences in data acquisition times (Terra or Aqua), weather conditions (clear-sky or all-sky), and processing methodologies (with or without gap filling). Overall, our proposed method, dataset, and analysis results have the potential to provide valuable insights for future urban climate studies. The UHII dataset is publicly available at https://doi.org/10.6084/m9.figshare.24821538.

54 ENVIRONMENTAL SCIENCES

RadioGalaxyNET: Dataset and novel computer vision algorithms for the detection of extended radio galaxies and infrared hosts

Abstract Creating radio galaxy catalogues from next-generation deep surveys requires automated identification of associated components of extended sources and their corresponding infrared hosts. In this paper, we introduce RadioGalaxyNET, a multimodal dataset, and a suite of novel computer vision algorithms designed to automate the detection and localization of multi-component extended radio galaxies and their corresponding infrared hosts. The dataset comprises 4 155 instances of galaxies in 2 800 images with both radio and infrared channels. Each instance provides information about the extended radio galaxy class, its corresponding bounding box encompassing all components, the pixel-level segmentation mask, and the keypoint position of its corresponding infrared host galaxy. RadioGalaxyNET is the first dataset to include images from the highly sensitive Australian Square Kilometre Array Pathfinder (ASKAP) radio telescope, corresponding infrared images, and instance-level annotations for galaxy detection. We benchmark several object detection algorithms on the dataset and propose a novel multimodal approach to simultaneously detect radio galaxies and the positions of infrared hosts.

Astronomy & Astrophysics