Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data products”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Comparisons of the v11.1 Orbiting Carbon Observatory‐2 (OCO‐2) X CO2 Measurements With GGG2020 TCCON

The Orbiting Carbon Observatory 2 (OCO-2) is NASA's first Earth observation satellite mission dedicated to studying the sources and sinks of carbon dioxide (CO 2 ) on a global scale. The observations of reflected sunlight are inverted in a retrieval algorithm to produce estimates of the dry air mole-fractions of CO 2 (X CO2 ). The OCO-2 Level 2 data release, version 11.1 (v11.1) retrievals from the Atmospheric Carbon Observations from Space (ACOS) algorithm, includes significant improvements in the X CO2 data product compared to older OCO-2 data versions. This work compares the v11.1 X CO2 from OCO-2 against X CO2 estimates collected from a global ground-based network known as the Total Carbon Column Observing Network (TCCON), OCO-2's primary validation source. The OCO-2 project provides a version of the Level 2 data product, called “lite” files that include calibrated and bias-corrected XCO2 values, accessible together with all OCO-2 data products through the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). This work shows that OCO-2 X CO2 observations made between September 2014 and December 2023, after quality filtering and the application of an averaging kernel correction, agree well with coincident TCCON data for all OCO-2 observational modes of land (nadir, glint, target) and ocean (glint). The aggregated, bias-corrected, and quality-filtered absolute average bias values are less than or equal to 0.20 parts per million (ppm) globally for all OCO-2 observation modes, where the biases do not indicate a statistically significant time dependence. The land nadir/glint mode has the lowest bias value of −0.03 ± 0.85 ppm.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Impacts of benchmarking choices on inferred model skill of the Arctic–Boreal terrestrial carbon cycle

Abstract Land surface models require continuous validation against observations to improve and reduce simulation uncertainty. However, inferred model performance can be heavily influenced by subjective choices made in the selection and application of observational data products. A key area often misrepresented by models is the Arctic–Boreal region, which is a potential tipping point region in Earth’s climate system due to large permafrost carbon stocks that are vulnerable to release with climate warming. We use the International Land Model Benchmarking (ILAMB) framework to evaluate how the model skill of TRENDY-v9 models varies based on the choice of observational-based benchmark and how benchmarks are applied in model evaluation. This analysis uses global datasets integrated into ILAMB and new, regionally-specific observational products from the Arctic–Boreal Vulnerability Experiment. Our results cover the overall time period of 1979–2019 and show that model scores can vary substantially depending on the data product applied, with higher model scores indicating better model performance against observations. The lowest model scores occur when benchmarked against regional, compared to global, datasets. We also evaluate observed and modeled functional relationships between ecosystem respiration and air temperature and between gross primary production and precipitation. Here, we find that the magnitude and shape of the responses are strongly impacted by the choice of observational dataset and the approach used to construct the functional relationship benchmark. These results suggest that model evaluation studies could conclude a false sense of model skill if only using a single benchmark data product or if not applying regional data products when performing a regional model analysis. Collectively, our findings highlight the influence of benchmarking choices on model evaluation and point to the need for benchmarking guidelines when assessing model skill.

Poe, Jeralyn (ORCID:0000000318495278)↗

A mapped dataset of surface ocean acidification indicators in large marine ecosystems of the United States

Mapped monthly data products of surface ocean acidification indicators from 1998 to 2022 on a 0.25° by 0.25° spatial grid have been developed for eleven U.S. large marine ecosystems (LMEs). The data products were constructed using observations from the Surface Ocean CO 2 Atlas, co-located surface ocean properties, and two types of machine learning algorithms: Gaussian mixture models to organize LMEs into clusters of similar environmental variability and random forest regressions (RFRs) that were trained and applied within each cluster to spatiotemporally interpolate the observational data. The data products, called RFR-LMEs, have been averaged into regional timeseries to summarize the status of ocean acidification in U.S. coastal waters, showing a domain-wide carbon dioxide partial pressure increase of 1.4 ± 0.4 μatm yr -1 and pH decrease of 0.0014 ± 0.0004 yr -1 . RFR-LMEs have been evaluated via comparisons to discrete shipboard data, fixed timeseries, and other mapped surface ocean carbon chemistry data products. Regionally averaged timeseries of RFR-LME indicators are provided online through the NOAA National Marine Ecosystem Status web portal.

54 ENVIRONMENTAL SCIENCES↗

NGEE Arctic Integrated Modeling (IM3): Improved snow-vegetation interaction

This data product represents the integration of new code capability for arctic tundra snow-vegetation-terrain interactions into the Energy Exascale Earth System Model (E3SM), through the E3SM Land Model (ELM) component. This code integration is the result of collaborative effort between the NGEE Arctic project and the E3SM project. The NGEE Arctic project developed a total of six Integrated Modeling (IM) modules informed by observations and experiments. New ELM capability represented by this data product (IM3) falls into three categories: 1) Downscaling from gridcell to topographic unit level when working through the existing coupler bypass code. 2) Four new parameters (taper, stocking, bendresist, and vegshape) have been added to ELM to allow for flexible definition of snow-vegetation interactions. 3) Vegshape and bendresist parameters are used to calculate the fraction of leaf area and/or stem area buried by snow for a given snow depth. This data record consists of a single document (pdf format) that describes the theoretical basis for the snow-vegetation-terrain interactions added to ELM, and describes the modifications made to the ELM code. The Methods section of this metadata record includes a link to the public E3SM code repository where the exact code modifications as integrated in E3SM can be accessed. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Thornton, Peter E [ORNL] (ORCID:0000000247595158)↗

Tethered Balloon System Ozone Profiles during CoURAGE Summer Intensive Operational Period Field Campaign Report

During the summer IOP, the small ARM field campaign CRGTBSO3 collected measurements to gain an improved understanding of differences between the atmospheric composition in and just above the marine layer at the CoURAGE TBS site near the eastern shore of Chesapeake Bay. CRGTBSO3 included guest instrumentation with ozone (O 3 ) profile measurements on the TBS and surface O3 and meteorological measurements at the TBS site. An En-Sci 2Z electrochemical cell (ECC) ozonesonde (Komhyr 1969, 1986, Witte et al. 2018) was included on the TBS. The ozonesonde was connected to an InterMet iMet-4RSB radiosonde, and the overall data collected included vertical profiles of ozone, relative humidity, temperature, pressure, and altitude. The CRGTBSO3 iMet-4 radiosonde data is identical to that of the iMet in the Tethered Balloon System Merged Data Product (TBSMERGED; Gaustad and Dexheimer 2025). The TBSMERGED data product also includes meteorological data from a different sensor, the iMet XQ2. In some cases, the iMet XQ2 relative humidity (RH) data may be more accurate than the iMet-4, such as for some instances when the iMet-4 RH data stays at 100% for an extended period throughout a profile.

54 ENVIRONMENTAL SCIENCES↗

Size-resolved Eddy-Covariance Particle Flux Measurement during the TRACER Campaign (Final Report)

The main goal of the TRacking Aerosol Convection interactions ExpeRiment (TRACER) campaign was to study aerosol–cloud interactions during deep convection over the Houston area. This project deployed a suite of instrumentation with the aim to (1) quantify turbulent vertical particle fluxes during at DOE-ARM sites, including TRACER, (2) assess hygroscopic growth factors and hygroscopicity parameters of the material driving modal aerosol growth during new particle formation and growth events, (3) derive turbulent aerosol mass fluxes using co-located Doppler LIDAR measurements, and (4) create quality-controlled PI data products to support future research utilizing data collected during the TRACER campaign. This report summarized the main findings from the deployments at two DOE-ARM sites. Briefly, we found that new particle formation may occur aloft, in a residual layer, near the top of the boundary layer. Small grown particles appear later due to downward mixing with daytime turbulence. The species that are responsible for aerosol modal growth had hygroscopicity parameters varying between 0.05 and 0.34. These values systematically depended on the wind sector, suggesting that the chemical composition of the precursors differed. This work demonstrated that lidar retrievals of the elastic backscatter and Doppler velocity can be used to obtain surface number emissions of particles with a diameter greater than 0.53 µm. During TRACER, emission particle number fluxes peaked near ∼ 100 cm−2 s−1. Multiple quality-controlled PI data products that will support future TRACER related science were generated and made publically available.

54 ENVIRONMENTAL SCIENCES↗

CHESS 2025: Crown polygons and extracted reflectance for field sampling sites

This dataset contains (1) crown polygons for each tree, meadow, and shrub site sampled in the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign (in geojson format, .geojson) and (2) extracted reflectance, uncertainty, and shade estimates for each crown polygon from the 2018 National Ecological Observatory Network (NEON) and 2025 CHESS campaigns. (in CSV format, .csv). Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). Crown polygons were manually delineated for each site in the 2025 campaign using a combination of field-collected GPS data (doi:10.15485/3022418), RGB (red, green, blue) and false color reflectance mosaics (doi:10.15485/3013535), and LiDAR-derived (Light Detection and Ranging) canopy height (CHM) and digital surface (DSM) models (DOI and citation to be added upon publication). Where there was misalignment between the spectrometer- and LiDAR-derived data products, polygons prioritized alignment with the spectrometer-derived data products. Polygons were delineated conservatively to only select pixels representative of vegetation samples collected in the field. Crown polygons for 2018 are published at (doi:10.15485/1618130) and were developed using the same protocol. For each polygon, all pixels from all flightlines were extracted where the pixel centroid was contained within the polygon. For each pixel, we extracted the surface reflectance, uncertainty, and shade estimates. Details on the extracted datasets are available at (doi:10.15485/3013527, doi:10.15485/3013535). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

SITCOMTN-154: Initial studies of photometric redshifts with LSSTComCam from DP1

This technote holds reports based on the first analyses of the Data Preview 1 (DP1) data by the Science Unit for photometric redshifts. Although photometric redshifts are not an official DP1 data product, the "Photo-z Science Unit" generated photo-z estimates for every galaxy in DP1 using the available multi-band imaging on a best-effort basis. This work included developing training and test datasets by matching DP1 data to high-quality reference redshifts obtained with spectroscopy, Grism data, and multi-band photometry. The Science Unit used the RAIL software package to make photometric redshift estimates using eight different algorithms, developed simple scientific performance metrics, used those metrics to explore how the performance of the algorithms varied with configuration changes, derived more optimized configurations of the algorithms and tested the performance of those configurations. This work, the resulting data products and expected data distribution mechanism are all described there.

79 ASTRONOMY AND ASTROPHYSICS↗

Synthetic Streamflow Datasets to Support Emulation of Water Allocations via LSTM

This archive is the data companion to the bonney_et-al_2026_erc metarepo which generates synthetic data, trains an LSTM model, and generates performance metrics on the trained model. While the generation of the synthetic data is fully reprodicible, it is a computationally expensive process. This data archive contains the synthetic datasets needed for training and testing an LSTM model and reproduction of figures and tables. In addition, supplemenatary data products generating and visualizing results is also included, such as geospatial data for the basin. Contents There are two high level directories: `WRAP_archive/` and `repo_data/`. The `WRAP_archive` directory contains compressed intermediate dataproducts from the dataset generation workflow (marked as "I_Dataset_Generation" in the metarepo). These data products are not required by any scripts in the metarepo, but they are archived as they are expensive to generate and may have useful information for other analyses. The `repo_data` directory contains the necessary data for reproducing the workflow in the metarepo and should be decompressed and moved into the top level of the metarepo. Additional details are provided in README.md.

drought↗

RTN-011: Rubin Observatory Plans for an Early Science Program

This document outlines Rubin Observatory's plans for a dedicated \emph{Early Science Program} to enable high-impact science prior to the first annual data release of the Legacy Survey of Space and Time (LSST). Components of the Early Science Program include releasing science-grade commissioning data products via a series of ``Data Previews,'' ramping up of the transient alert stream during commissioning, implementing a program of incremental template generation to augment alert production in the early phases of the survey, and the first LSST Data Release, DR1, based on the first 6 months of data from the LSST. A detailed breakdown of which data products can be expected when is provided. The Rubin Operations team is working closely with the science community to optimize the Early Science Program for the time-domain and solar system science achievable in the first year of operations. This is a living document; both it and the Early Science Program will continue to evolve over the course of commissioning and pre-operations in response to the state of the as-built system and to community guidance.

79 ASTRONOMY AND ASTROPHYSICS↗

Advancing Our Understanding of System Availability through the PV Fleet Performance Data Initiative

The PV Fleet Performance Data Initiative partners with photovoltaic (PV) fleet owners to collect time-series data of PV production data and publishes aggregated anonymized results of system performance metrics. With an extensive dataset drawn from over 2,200 PV systems across the United States, comprising 8.5 GW and 24,000 separate inverter data channels, this initiative aims to ensure that systemic risks in the US PV fleet are detected. The current work explores system availability, revealing a pronounced dependence on time, especially within the initial 6 months of system performance. Following this start-up period, the average system availability stabilizes. Statistical analyses illustrate a median (P5O) monthly availability of 0.991 and a dependence on system size with a negative trend in availability with increasing system size. This finding indicates that larger systems experience lower availability compared to their smaller counterparts.

inverter availability↗

Site B - NREL ASSIST (SN11) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 11) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

Site G - NREL ASSIST (SN10) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 10) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

Site C1a - NREL ASSIST (SN12) Thermodynamic Retrievals TROPoe / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe v0.12 (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). The profiles are retrieved every 10 minutes from instantaneous observations from the NREL ASSIST-II (SN 12) infrared spectrometer. Observations are noise-filtered but not averaged in time to minimize errors due to non-uniform clouds. Additional input data in TROPoe are cloud base height (CBH), which is a combined data product that uses data from ceilometers at sites A1 and H and scanning lidars from ARM sites C1 and E37. The CBH is weighted inversely proportionally to the distance to the respective site to take into account the spatial variability of clouds (see https://github.com/StefanoWind/ASSIST_analysis/blob/main/awaken_processing/combine_cbh.py). The full pipeline for running the retrieval is available at https://github.com/StefanoWind/TROPoe_processor. Met data was not ingested. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at ARM SGP, OK.

17 WIND ENERGY↗

CRGTBSO3 TBS Ozone Data

The Tethered Balloon System (TBS) operated for two weeks during a summer IOP of the CoURAGE campaign. This dataset includes data collected from an En-Sci electrochemical cell (ECC) ozonesonde on the TBS. The ozonesonde was connected to an iMet-4RSB radiosonde, and the overall data collected included ozone, relative humidity, temperature, and altitude. The data from the iMet is the same as that found in the TBSMERGED data product. The TBS also had another instrument (iMet XQ2) that collected meteorological data, which may have more accurate relative humidity (RH) data. This ozonesonde data set is intended to complement the TBSMERGED data product, the CRGTBSO3 surface ozone measurements, and the CoURAGE SWARM ozone lidar (TOLNet) measurements from other locations.

Atmospheric relative humidity↗