Engineering PapersSearch

SEARCH · Engineering Papers

Results for “dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Comparison of Radiative Energy Flows in Observational Datasets and Climate Modeling

This study examines radiative flux distributions and local spread of values from three major observational datasets (CERES, ISCCP, and SRB) and compares them with results from climate modeling (CMIP3). Examinations of the spread and differences also differentiate among contributions from cloudy and clear-sky conditions. The spread among observational datasets is in large part caused by noncloud ancillary data. Average differences of at least 10Wm(exp -2) each for clear-sky downward solar, upward solar, and upward infrared fluxes at the surface demonstrate via spatial difference patterns major differences in assumptions for atmospheric aerosol, solar surface albedo and surface temperature, and/or emittance in observational datasets. At the top of the atmosphere (TOA), observational datasets are less influenced by the ancillary data errors than at the surface. Comparisons of spatial radiative flux distributions at the TOA between observations and climate modeling indicate large deficiencies in the strength and distribution of model-simulated cloud radiative effects. Differences are largest for lower-altitude clouds over low-latitude oceans. Global modeling simulates stronger cloud radiative effects (CRE) by +30Wmexp -2) over trade wind cumulus regions, yet smaller CRE by about -30Wm(exp -2) over (smaller in area) stratocumulus regions. At the surface, climate modeling simulates on average about 15Wm(exp -2) smaller radiative net flux imbalances, as if climate modeling underestimates latent heat release (and precipitation). Relative to observational datasets, simulated surface net fluxes are particularly lower over oceanic trade wind regions (where global modeling tends to overestimate the radiative impact of clouds). Still, with the uncertainty in noncloud ancillary data, observational data do not establish a reliable reference.

Raschke, Ehrhard

The Cumulus and Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD)

Low clouds continue to contribute greatly to the uncertainty in cloud feedback estimates. Depending on whether a region is dominated by cumulus (Cu) or stratocumulus (Sc) clouds, the interannual low-cloud feedback is somewhat different in both spaceborne and large-eddy simulation studies. Therefore, simulating the correct amount and variation of the Cu and Sc cloud distributions could be crucial to predict future cloud feedbacks. Here we document spatial distributions and profiles of Sc and Cu clouds derived from Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) and CloudSat measurements. For this purpose, we create a new dataset called the Cumulus And Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD), which identifies Sc, broken Sc, Cu under Sc, Cu with stratiform outflow and Cu. To separate the Cu from Sc, we design an original method based on the cloud height, horizontal extent, vertical variability and horizontal continuity, which is separately applied to both CALIPSO and combined CloudSat–CALIPSO observations. First, the choice of parameters used in the discrimination algorithm is investigated and validated in selected Cu, Sc and Sc–Cu transition case studies. Then, the global statistics are compared against those from existing passive- and active-sensor satellite observations. Our results indicate that the cloud optical thickness – as used in passive-sensor observations – is not a sufficient parameter to discriminate Cu from Sc clouds, in agreement with previous literature. Using clustering-derived datasets shows better results although one cannot completely separate cloud types with such an approach. On the contrary, classifying Cu and Sc clouds and the transition between them based on their geometrical shape and spatial heterogeneity leads to spatial distributions consistent with prior knowledge of these clouds, from ground-based, ship-based and field campaigns. Furthermore, we show that our method improves existing Sc–Cu classifications by using additional information on cloud height and vertical cloud fraction variation. Finally, the CASCCAD datasets provide a basis to evaluate shallow convection and stratocumulus clouds on a global scale in climate models and potentially improve our understanding of low-level cloud feedbacks. The CASCCAD dataset (Cesana, 2019, https://doi.org/10.5281/zenodo.2667637) is available on the Goddard Institute for Space Studies (GISS) website at https://data.giss.nasa.gov/clouds/casccad/ (last access: 5 November 2019) and on the zenodo website at https://zenodo.org/record/2667637 (last access: 5 November 2019).

Cesana, Gregory V.

Version 2 Ozone Monitoring Instrument SO2 Product (OMSO2 V2): New Anthropogenic SO2 Vertical Column Density Dataset

The Ozone Monitoring Instrument (OMI) has been providing global observations of SO2 pollution since 2004. Here we introduce the new anthropogenic SO2 vertical column density (VCD) dataset in the version 2 OMI SO2 product (OMSO2 V2). As with the previous version (OMSO2 V1.3), the new dataset is generated with an algorithm based on principal component analysis of OMI radiances, but features several updates. The most important among those is the use of expanded lookup tables and model a priori profiles to estimate SO2 Jacobians for individual OMI pixels, in order to better characterize pixel-to-pixel variations in SO2 sensitivity, including over snow and ice. Additionally, new data screening and spectral fitting schemes have been implemented to improve the quality of the spectral fit. As compared with the planetary boundary layer SO2 dataset in OMSO2 V1.3, the new dataset has substantially better data quality, especially over areas that are relatively clean or affected by the south Atlantic anomaly. The updated retrievals over snow/ice yield more realistic seasonal changes in SO2 at high latitudes and offer enhanced sensitivity to sources during wintertime. An error analysis has been conducted to assess uncertainties in SO2 VCDs from both the spectral fit and Jacobian calculations. The uncertainties from spectral fitting are reflected in SO2 slant column densities (SCDs) and largely depend on the signal-to-noise ratio of the measured radiances, as implied by the generally smaller SCD uncertainties over clouds or for smaller solar zenith angles. The SCD uncertainties for individual pixels are estimated to be~0.15-0.3 DU (Dobson Units) between ~40°S and ~40°N and to be~0.2-0.5 DU at higher latitudes. The uncertainties from the Jacobians are approximately ~50-100% over polluted areas, and primarily attributed to errors in SO2 a priori profiles and cloud pressures, as well as the lack of explicit treatment for aerosols. Finally, the daily mean and median SCDs over the presumably SO2-free equatorial East Pacific have increased by only~0.0035 DU and ~0.003 DU respectively over the entire 15-year OMI record; while the standard deviation of SCDs has grown by only~0.02 DU or ~10%. Such remarkable long-term stability makes the new dataset particularly suitable for detecting regional changes in SO2 pollution.

OMI, SO2, Remote Sensing

PERSIANN Dynamic Infrared–Rain Rate (PDIR-Now): A Near-Real-Time, Quasi-Global Satellite Precipitation Dataset

This study presents the Precipitation Estimation from Remotely Sensed Information Using Artificial Neural Networks–Dynamic Infrared Rain Rate (PDIR-Now) near-real-time precipitation dataset. This dataset provides hourly, quasi-global, infrared-based precipitation estimates at 0.04° × 0.04° spatial resolution with a short latency (15–60 min). It is intended to supersede the PERSIANN–Cloud Classification System (PERSIANN-CCS) dataset previously produced as the near-real-time product of the PERSIANN family. We first provide a brief description of the algorithm’s fundamentals and the input data used for deriving precipitation estimates. Second, we provide an extensive evaluation of the PDIR-Now dataset over annual, monthly, daily, and subdaily scales. Last, the article presents information on the dissemination of the dataset through the Center for Hydrometeorology and Remote Sensing (CHRS) web-based interfaces. The evaluation, conducted over the period 2017–18, demonstrates the utility of PDIR-Now and its improvement over PERSIANN-CCS at all temporal scales. Specifically, PDIR-Now improves the estimation of rain/no-rain days as demonstrated by a critical success index (CSI) of 0.53 compared to 0.47 of PERSIANN-CCS. In addition, PDIR-Now improves the estimation of seasonal and diurnal cycles of precipitation as well as regional precipitation patterns erroneously estimated by PERSIANN-CCS. Finally, an evaluation is carried out to examine the performance of PDIR-Now in capturing two extreme events, Hurricane Harvey and a cluster of summer thunderstorms that occurred over the Netherlands, where it is shown that PDIR-Now adequately represents spatial precipitation patterns as well as subdaily precipitation rates with a correlation coefficient (CORR) of 0.64 for Hurricane Harvey and 0.76 for the Netherlands thunderstorms.

Rainfall

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda

AssistTaxi: A Comprehensive Dataset for Taxiway Analysis and Autonomous Operations

The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems. This poster presents AssistTaxi, which is a comprehensive novel dataset which is a collection of images for runway and taxiway analysis. The dataset comprises of more than 300,000 frames of diverse and carefully collected data, gathered from Melbourne (MLB) and Grant-Valkaria (X59) general aviation airports. The importance of AssistTaxi lies in its potential to advance autonomous operations, enabling researchers and developers to train and evaluate algorithms for efficient and safe taxiing. Researchers can utilize AssistTaxi to benchmark their algorithms, assess performance, and explore novel approaches for runway and taxiway analysis. Additionally, the dataset serves as a valuable resource for validating and enhancing existing algorithms as well as facilitating innovation in autonomous operations for aviation. We also propose an initial approach to label the dataset using a contour based detection and line extraction technique.

Data Collection

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source

Predicting cutoff L-shells of solar protons using the GPPSn particle dataset

Solar energetic protons (SEPs) arriving at the Earth trigger severe radiation storms in the near-Earth space, directly impacting space missions operating at various altitudes. Therefore, monitoring SEP events and predicting the penetration depths of solar protons are critical for aerospace sectors. Building on previous efforts, here we demonstrate the feasibility of using proton measurements from the Global Prompt Proton Sensor network (GPPSn), enabled by Los Alamos National Laboratory developed combined X-ray dosimeters aboard GPS satellites, to characterize and predict the penetration of solar protons into the geomagnetic field. The inclined medium-Earth-orbits (MEOs) of the global GPS constellation offer a unique advantage of allowing simultaneous measurements of penetrating solar protons inside both open- and closed-field line regions. Therefore, the L-profiles of ∼10s–100 MeV solar protons and their associated cutoff L-shells can be determined from the GPPSn dataset, using predefined threshold proton flux values rather than traditional flux ratios. After examining a list of SEP event intervals across solar cycles 23, 24 and 25—including the 2024 Mother’s Day superstorm, we showcase how the latest GPPSn proton dataset (release v1.10), reprocessed and calibrated, can not only be used to monitor solar proton distributions inside the dynamic geomagnetic field for individual events, but also to derive a new empirical model linking cutoff L-shells with several key space weather parameters. This newly developed SEPCL-MEO model demonstrates high predictive performance; for example, predictions for > 30 MeV solar protons yield a correlation coefficient of 0.85 and performance efficiency of 0.67 when validated against GPPSn observations. Results from this pilot study underscores the scientific and operational value of the GPPSn dataset, and this dataset—when paired with machine-learning techniques—can play a critical role in observing and predicting the effects of future incoming SEP events, including extreme ones.

58 GEOSCIENCES

A Large Dataset of Fluvial Hydraulic and Geometry Attributes Derived From USGS Field Measurement Records

Accurate representation of river channel geometry is important for hydrologic and hydraulic modeling of fluvial systems. Often, channel geometry is estimated using simple rating curves that can be applied across various spatial scales. However, such methods are limited to power law relations that do not employ many potentially relevant catchment and river attributes. This paper introduce a new dataset, IFMHA (Inventory of Field Measurement of Hydraulic Attributes), to enable research studies on channel geometry and streamflow characteristics. IFMHA is derived from the National Water Information System (NWIS) site inventory for surface water field measurements and stream attributes from the National Hydrography Dataset (NHD). IFMHA includes 2,802,532 records from 10,050 sites (NWIS streamgaging stations). The dataset utility is demonstrated here by presenting a series of conceptual models for estimating channel geometry parameters (i.e., channel mean depth, channel maximum depth, wetted perimeter, and roughness) based on the available field attributes within IFMHA. Such a dataset and attributed channel geometry parameters can enhance the performance of operational flood forecasting frameworks (e.g. National Water Model) by providing more accurate initial conditions used in hydrologic and hydraulic routing models.

Hydrology

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David

The high explosives & affected targets (HEAT) dataset

Artificial Intelligence (AI) surrogate models offer a computationally efficient alternative to full-physics simulations, yet no existing datasets are publicly available for training, testing, and validation of machine learning models of the dynamics of high-explosive driven shocks through multiple materials. Shock propagation through materials is a computationally challenging problem because simulations must include material-specific equations of state (EOS) along with descriptions of other physical processes such as plastic deformation, phase change, damage processes, fluid instabilities, and multi-material interactions. Shocks are typically initiated by high-velocity impacts or explosive loading. The latter case necessitates the addition of models of reactive materials to represent high-explosive (HE) detonation. Here, to address the lack of an expansive dataset for multi-material shock propagation in the AI/ML community, we present the High-Explosives and Affected Targets (HEAT) Dataset. HEAT is a physics-rich collection of two-dimensional, cylindrically symmetric, simulations generated using an Eulerian, multi-material, shock-propagation code developed at Los Alamos National Laboratory. The dataset includes two partitions: (1) the expanding shock-cylinder (CYL) simulations, Figs. 1, and (2) the Perturbed Layered Interface (PLI) simulations, Fig. 2. Entries in both partitions consist of time series of arrays of thermodynamic fields (pressure, density, and temperature), kinematic fields (position and velocity), and additional fields that depend on thermodynamic and/or kinematic fields (e.g., material stress). Materials in the CYL partition include solids (aluminium, copper, depleted uranium, stainless steel, tantalum, and a generic polymer), a liquid (water), gases (air, nitrogen), and a generic detonating material (high explosive, HE). The PLI partition spans a highly varying geometry but consists of fixed materials across entries: Copper, aluminium, stainless steel, generic polymer, and generic HE. HEAT captures critical phenomena such as momentum transfer, shock propagation, plastic deformation, and thermal effects, making HEAT a valuable benchmark for development of AI/ML emulation of multi-material shock propagation.

36 MATERIALS SCIENCE

Systematic Evaluation of Atmospheric Forcing, Surface Datasets, and Mesh Effects on Kilometer-Scale Land Surface and River Modeling

Earth system models are advancing toward kilometer-scale resolution to capture local climate impacts and extremes. High-resolution land and river modeling depends on multiple factors, including mesh, surface datasets, and atmospheric forcing, but their relative effects at kilometer scales remain unquantified. We evaluated five Energy Exascale Earth System Model land and river configurations over the Mid-Atlantic region using two mesh (1/8° structured versus variable-resolution unstructured mesh), two surface datasets (default versus newly developed), and three atmospheric forcings (NLDAS2, MSWX, GSWP). Evaluation against satellite, reanalysis, and in situ benchmarks across water, energy, and carbon cycles quantifies how these factors affect model performance. Forcing selection produces the largest bias reductions (12-99% across variables), followed by surface datasets (7-75%) and mesh (up to 21%). Forcing effects vary by variable, with MSWX reducing biases for snow water equivalent, evapotranspiration, albedo, temperature, and gross primary productivity, GSWP for snow cover and runoff, and NLDAS for soil moisture and streamflow. The use of newly developed surface datasets improves gross primary productivity (58% bias reduction) and evapotranspiration but increase soil moisture and albedo biases due to current modeling limitations. Variable-resolution unstructured mesh improves the simulation of small-basin streamflow through better capturing drainage networks, though mesh minimally affects other land variables. These findings provide important guidance for high-resolution modeling development and actionable science.

Land and River modeling

A global urban heat island intensity dataset: Generation, comparison, and analysis

The urban heat island (UHI) effect, a phenomenon of local warming over urban areas, is the most well-known impact of urbanization on climate. Globally consistent estimates of the UHI intensity (UHII) are crucial for examining this phenomenon across time and space. However, publicly available UHII datasets are limited and have several constraints: (1) they are for clear-sky surface UHII, not all-sky surface UHII and canopy (air temperature) UHII; (2) the estimation methods often neglect anthropogenic disturbance, introducing uncertainties in the estimated UHII. To address these issues, this study proposes a new dynamic equal-area (DEA) method that can minimize the influence of various confounding factors on UHII estimates through a dynamic cyclic process. Utilizing the DEA method and leveraging various gridded temperature data, we develop a global-scale (>10,000 cities), long-term (over 20 years by month), and multi-faceted (clear-sky surface, all-sky surface, and canopy) UHII dataset. Further, based on these estimates, we provide a comprehensive analysis of the UHII and its trends in global cities. The UHII is found to be greater than zero in >80% of cities, with global annual average magnitudes around 1.0 °C (day) and 0.8 °C (night) for surface UHII, and close to 0.5 °C for canopy UHII. Furthermore, an interannual upward trend in UHII is observed in >60% of cities, with global annual average trends exceeding 0.1 °C/decade (day) and over 0.06 °C/decade (night) for surface UHII, and slightly surpassing 0.03 °C/decade for canopy UHII. Notably, there exists a positive correlation between the magnitude and trend of UHII, suggesting that cities with stronger UHII tend to experience faster growth in UHII. Additionally, discrepancies in UHII are found between different temperature data, stemming not only from distinctions in data types (surface or air temperature) but also from differences in data acquisition times (Terra or Aqua), weather conditions (clear-sky or all-sky), and processing methodologies (with or without gap filling). Overall, our proposed method, dataset, and analysis results have the potential to provide valuable insights for future urban climate studies. The UHII dataset is publicly available at https://doi.org/10.6084/m9.figshare.24821538.

54 ENVIRONMENTAL SCIENCES

RadioGalaxyNET: Dataset and novel computer vision algorithms for the detection of extended radio galaxies and infrared hosts

Abstract Creating radio galaxy catalogues from next-generation deep surveys requires automated identification of associated components of extended sources and their corresponding infrared hosts. In this paper, we introduce RadioGalaxyNET, a multimodal dataset, and a suite of novel computer vision algorithms designed to automate the detection and localization of multi-component extended radio galaxies and their corresponding infrared hosts. The dataset comprises 4 155 instances of galaxies in 2 800 images with both radio and infrared channels. Each instance provides information about the extended radio galaxy class, its corresponding bounding box encompassing all components, the pixel-level segmentation mask, and the keypoint position of its corresponding infrared host galaxy. RadioGalaxyNET is the first dataset to include images from the highly sensitive Australian Square Kilometre Array Pathfinder (ASKAP) radio telescope, corresponding infrared images, and instance-level annotations for galaxy detection. We benchmark several object detection algorithms on the dataset and propose a novel multimodal approach to simultaneously detect radio galaxies and the positions of infrared hosts.

Astronomy & Astrophysics

Information-entropy-driven generation of material-agnostic datasets for machine-learning interatomic potentials

In contrast to their empirical counterparts, machine-learning interatomic potentials (MLIAPs) promise to deliver near-quantum accuracy over broad regions of configuration space. However, due to their generic functional forms and extreme flexibility, they can catastrophically fail to capture the properties of novel, out-of-sample configurations, making the quality of the training set a determining factor, especially when investigating materials under extreme conditions. We propose a novel automated dataset generation method based on the maximization of the information entropy of the feature distribution, aiming at an extremely broad coverage of the configuration space in a way that is agnostic to the properties of specific target materials. The ability of the dataset to capture unique material properties is demonstrated on a range of unary materials, including elements with the FCC (Al), BCC (W), HCP (Be, Re and Os), graphite (C), and trigonal (Sb, Te) ground states. MLIAPs trained to this dataset are shown to be accurate over a range of application-relevant metrics, as well as extremely robust over very broad swaths of configurations space, even without dataset fine-tuning or hyper-parameter optimization, making the approach extremely attractive to rapidly and autonomously develop general-purpose MLIAPs suitable for simulations in extreme conditions.

36 MATERIALS SCIENCE

A unified ensemble soil moisture dataset across the continental United States

Abstract A unified ensemble soil moisture (SM) package has been developed over the Continental United States (CONUS). The data package includes 19 products from land surface models, remote sensing, reanalysis, and machine learning models. All datasets are unified to a 0.25-degree and monthly spatiotemporal resolution, providing a comprehensive view of surface SM dynamics. The statistical analysis of the datasets leverages the Koppen-Geiger Climate Classification to explore surface SM’s spatiotemporal variabilities. The extracted SM characteristics highlight distinct patterns, with the western CONUS showing larger coefficient of variation values and the eastern CONUS exhibiting higher SM values. Remote sensing datasets tend to be drier, while reanalysis products present wetter conditions. In-situ SM observations serve as the basis for wavelet power spectrum analyses to explain discrepancies in temporal scales across datasets facilitating daily SM records. This study provides a comprehensive soil moisture data package and an analysis framework that can be used for Earth system model evaluations and uncertainty quantification, quantifying drought impacts and land–atmosphere interactions and making recommendations for drought response planning.

54 ENVIRONMENTAL SCIENCES

LandScan Global 30 Arcsecond Annual Global Gridded Population Datasets from 2000 to 2022

Abstract Oak Ridge National Laboratory (ORNL) annually develops the LandScan Global (LSG) dataset, a 30 arcsecond global gridded population dataset representing global ambient human population distribution. This multivariable dasymetric model disaggregates census counts within administrative boundaries using ancillary data. Each country’s distribution reflects cultural and socioeconomic patterns; manual validations yield a unique global dataset for assessing populations at risk. For over two decades, LSG has been a standard for estimating populations at risk, aiding U.S. federal government, academia and humanitarian organizations. During disasters such as the 2004 Indian Ocean tsunami and the 2010 Haiti earthquake and geopolitical crises such as the Syrian civil war and the 2022 Russian invasion of Ukraine, LSG supported scientific and operational communities in emergency response and recovery. In 2022, LSG datasets from 2000 onward were made publicly available through ORNL’s LandScan Portal. This data descriptor details our methodology and the application of geospatial science and machine learning to geographic and demographic data, highlighting uses in urban resiliency, emergency management, disaster response, and human health and security.

Science & Technology - Other Topics

A global dataset of terrestrial biological nitrogen fixation

Biological nitrogen fixation (BNF) is the main natural source of new nitrogen inputs in terrestrial ecosystems, supporting terrestrial productivity, carbon uptake, and other Earth system processes. We assembled a comprehensive global dataset of field measurements of BNF in all major N-fixing niches across natural terrestrial biomes derived from the analysis of 376 BNF studies. The dataset comprises 32 variables, including site location, biome type, N-fixing niche, sampling year, quantification method, BNF rate (kg N ha −1 y −1 ), the percentage of nitrogen derived from the atmosphere (%N dfa ), N fixer or N-fixing substrate abundance, BNF rate per unit of N fixer abundance, and species identity. Overall, the dataset combines 1,207 BNF rates for trees, shrubs, herbs, soil, leaf litter, woody litter, dead wood, mosses, lichens, and biocrusts, 152 herb %N dfa values, 1,005 measurements of N fixer or N-fixing substrate abundance, and 762 BNF rates per unit of N fixer abundance for a total of 424 species across 66 countries. This dataset facilitates synthesis, meta-analysis, upscaling, and model benchmarking of BNF fluxes at multiple spatial scales.

Reis Ely, Carla R. [Oregon State Univ., Corvallis,