Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data products”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

MINERvA open-data product

MINERvA is THE neutrino cross section experiment Scintillator tracker/calorimeter ran in the NuMI beam at Fermilab same beam as the MINOS and NOvA oscillation experiments With our data, we are solving systematic shortcomings in neutrino interaction rate/spectra that are the largest part of the systematic uncertainty in today s (and tomorrow s) measurements. Some aspects have NO equivalent in the neutrino program future. Scientific scope includes both particle and nuclear physics GeV scale cross sections, A dependence, MeV scale effects most published measurements for a neutrino experiments ever (well, tied with T2K) with 25% more papers in the pipeline.

Gran, Rik [Minnesota U., Duluth] (ORCID:0000000216

Atmospheric Radiation Measurement (ARM) airborne field campaign data products between 2013 and 2018

Airborne measurements are pivotal for providing detailed, spatiotemporally resolved information about atmospheric parameters and aerosol and cloud properties, thereby enhancing our understanding of dynamic atmospheric processes. For 30 years, the US Department of Energy (DOE) Office of Science supported an instrumented Gulfstream 1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) Data Center and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated dataset was recently developed covering the final 6 years of G-1 operations (2013 to 2018, https://doi.org/10.5439/1999133; Mei and Gaustad, 2024). The integrated dataset includes data collected from 236 flights (766.4 h), which covered the Arctic, the US Southern Great Plains (SGP), the US West Coast, the eastern North Atlantic (ENA), the Amazon Basin in Brazil, and the Sierras de Córdoba range in Argentina. These comprehensive data streams provide much-needed insight into spatiotemporal variability in the thermodynamic quantities and aerosol and cloud properties for addressing essential science questions in Earth system process studies. This paper describes the DOE ARM merged G-1 datasets, including information on the acquisition, data collection challenges and future potentials, and quality control processes. It further illustrates the usage of this merged dataset to evaluate the Energy Exascale Earth System Model (E3SM) with the Earth System Model Aerosol–Cloud Diagnostics (ESMAC Diags) package.

54 ENVIRONMENTAL SCIENCES

Image masks of global ship tracks for NASA MODIS data products

Ship tracks, long thin artificial cloud features formed from the pollutants in ship exhaust, are satellite-observable examples of aerosol-cloud interactions (ACI) that can lead to increased cloud albedo and thus increased solar reflectivity, phenomena of interest in solar radiation management. In addition to ship tracks being of interest to meteorologists and policy makers, their observed cloud perturbations provide benchmark evidence of ACI that remain poorly captured by climate models. To broadly analyze the effects of ship tracks, high-resolution satellite imagery data highlighting their presence are required. To support this, we provide a hand labelled dataset to serve as a benchmark for a variety of subsequent analyses. Established from a previous dataset that identified ship track presence using NASA’s MODIS Aqua satellite imager, our first-of-its-kind dataset is comprised of image masks: capturing full ship track regions, including their contours, emission points and dispersive patterns. In total, 300 images, or around 2,500 masked ship tracks, observed under varying conditions are provided, and may facilitate training of machine learning algorithms to automate extraction.

Atmospheric dynamics

Advanced Precipitation and Boundary Layer Data Products Derived from ARM Radar Wind Profilers

This research project was successful in delivering on four main objectives. First, software was developed to accurately calculate 915-MHz radar wind profiler (RWP) spectrum moments from the recorded Doppler velocity power spectra. Second, software was developed to calibrate the RWP reflectivity factor using collocated surface disdrometer observations. Third, the Python processing code was documented and given to the ARM Infrastructure to produce ARM ‘b level’ calibrated RWP products. Fourth, calibrated RWP products were uploaded to the ARM Archive as PI Products for 10 years of SGP RWP observations and for GoAmazon and TRACER field campaign RWP observations. In addition to working with RWP observations, this research project also worked with KAZR observations to distinguish insects from boundary layer clouds to help improve the ARSCL cloud mask product. The PI worked with senior and early career ARM funded scientists at BNL exploring how to include calibrated RWP moments into future versions of the ARSCL product.

54 ENVIRONMENTAL SCIENCES

Establishing robust correction schemes for improved and reliable ARM-AOS aerosol optical data products

Aerosol light absorption and scattering of solar radiation play an important role in the earth’s atmosphere in terms of direct and semi-direct radiative forcing. Optical parameters of importance to the US Department of Energy (DOE) climate models include absorption and scattering coefficients, single scattering albedo (SSA), absorption Angstrom exponents (AAE), and the asymmetry parameter (g). These parameters depend on aerosol size, shape and composition (refractive index), and are spectrally sensitive in the shortwave region. Additionally, these parameters have a complex dependency on the emission source, especially for carbonaceous aerosols. The DOE Atmospheric Radiation Measurement (ARM) user facility has deployed aerosol observing systems (AOS) containing several filter-based instruments to measure and constrain aerosol optical properties and related parameters at multiple sites worldwide. For measurement of aerosol light absorption, the AOS includes filter-based instruments (particle soot absorption photometer and tricolor absorption photometer) that infer particle-phase aerosol absorption coefficients at nominal red, green, and blue wavelength bands from the attenuation (ATN) of light passing through a particulate filter on which aerosols are deposited. Measurement of aerosol scattering is done in situ using nephelometers. By combining inferred absorption coefficients from filter-based ATN measurements and in situ scattering coefficients, value-added products (VAPs) such as SSA, AAE, and g are derived.

54 ENVIRONMENTAL SCIENCES

Vera C. Rubin Observatory Prompt Products

Data products produced by prompt and daily processing of images obtained in the Legacy Survey of Space and Time. These include realtime alerts sent to community alert brokers, newly-discovered Solar System Objects reported to the Minor Planet Center, processed visit and difference images, and source catalogs. Prompt Products are not a static single data release but continually grow throughout the ten-year LSST survey.

79 ASTRONOMY AND ASTROPHYSICS

Raw ERCOT 60-Day SCED Disclosure Reports

This dataset contains ERCOT plant-level wind power production data at a 15-minute time resolution. Data at ERCOT are called 60-Day SCED Disclosure Report; data product number NP3-965-ER. Raw data are in local time. DST is handled as "skip hour" (12a, 1a, 3a,...) in spring, "extra hour" (12a, 1a, 2a, 2a, 3a,...) in fall.

17 WIND ENERGY

Vera C. Rubin Observatory Prompt Products: alert packets data

Data products produced by prompt and daily processing of images obtained in the Legacy Survey of Space and Time. These include realtime alerts sent to community alert brokers, newly-discovered Solar System Objects reported to the Minor Planet Center, processed visit and difference images, and source catalogs. Prompt Products are not a static single data release but continually grow throughout the ten-year LSST survey. This dataset is a subset of the full data release consisting of a dataset named alert packets. This dataset contains measurements for 5-sigma sources detected in difference images that were issued to the community brokers.

79 ASTRONOMY AND ASTROPHYSICS

Machine Learned Empirical Numerical Integrator from Simulated Data

Recently, a number of state-of-the-art surrogate machine learning (ML) models have been designed for global weather and climate prediction, which have been trained using reanalysis data products. Reanalysis data products are constructed using numerical model simulations that combine numerical integration of partial differential equations and parameterization schemes. These products are typically only archived and made available using coarsened spatial and temporal resolutions. This study explores the impact of the numerical generation methods used to produce the training datasets and the temporal resolution of those datasets on machine learning surrogate models. Using the nonlinear vector autoregression (NVAR) machine as an explainable ML technique, simple dynamical systems are emulated with ML models trained on data produced by three classical numerical integration schemes. NVAR is validated as a skillful ML method, capable of producing accurate predictions and, more importantly, reconstructing both the underlying dynamics and the numerical integration scheme used to generate the training data. However, the machine fails to generalize predictions on unseen test data generated by different numerical integration schemes, despite the underlying dynamical system being the same. This result provides a word of caution for the growing field of machine learning emulation of weather and climate dynamics. Furthermore, we illustrate using NVAR that training on temporally coarsened data may increase the required complexity of ML models and potentially introduce new numerical challenges. Finally, we discover that empirical integration schemes with arbitrary time-stepping sizes can be constructed directly from the data, which implies a potential for the development of empirical numerical integration schemes.

54 ENVIRONMENTAL SCIENCES

Advancing Multiscale Simulation of Plasma-Surface Interfaces

We report the development of an atomistic-informed, surface-state-dependent predictive model for particle exchange in a carbon-tungsten plasma-surface interface. The predictive model uses machine learning (ML) techniques to learn the energy and angular distributions for particle exchange and rate functions for surface state evolution from molecular dynamics simulations of cumulative bombardment of tungsten by energetic carbon ions. Each predictive component is sensitive to the energy and trajectory of incident plasma species and the surface state. The surface state is represented by a set of surface state descriptors, which were derived from the atomistic surface state for each independent carbon bombardment event. These descriptors are representative of the composition and degree of amorphization of the outermost angstrom of surface material and were chosen to optimize predictive performance for particle exchange at the interface. The distributions for particle exchange (reflection/sputtering) are demonstrated to vary with each surface state descriptor, motivating the development of surface-state-dependent particle exchange models for plasma simulations. The performance of various ML methods was compared, including polynomial quantile regression, artificial neural networks, k-nearest neighbors, and random forest algorithms, with polynomial regression performing the best for interpolation and extrapolation of learned relationships. In addition to the particle exchange model, a neutral network was developed and used to identify data sufficiency throughout surface descriptor space, which will enable real-time feedback during future data production to ensure data is produced where it is most needed, and we provide commentary on improvements to the data production workflow for future endeavors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Comparisons of the v11.1 Orbiting Carbon Observatory‐2 (OCO‐2) X CO2 Measurements With GGG2020 TCCON

The Orbiting Carbon Observatory 2 (OCO-2) is NASA's first Earth observation satellite mission dedicated to studying the sources and sinks of carbon dioxide (CO 2 ) on a global scale. The observations of reflected sunlight are inverted in a retrieval algorithm to produce estimates of the dry air mole-fractions of CO 2 (X CO2 ). The OCO-2 Level 2 data release, version 11.1 (v11.1) retrievals from the Atmospheric Carbon Observations from Space (ACOS) algorithm, includes significant improvements in the X CO2 data product compared to older OCO-2 data versions. This work compares the v11.1 X CO2 from OCO-2 against X CO2 estimates collected from a global ground-based network known as the Total Carbon Column Observing Network (TCCON), OCO-2's primary validation source. The OCO-2 project provides a version of the Level 2 data product, called “lite” files that include calibrated and bias-corrected XCO2 values, accessible together with all OCO-2 data products through the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). This work shows that OCO-2 X CO2 observations made between September 2014 and December 2023, after quality filtering and the application of an averaging kernel correction, agree well with coincident TCCON data for all OCO-2 observational modes of land (nadir, glint, target) and ocean (glint). The aggregated, bias-corrected, and quality-filtered absolute average bias values are less than or equal to 0.20 parts per million (ppm) globally for all OCO-2 observation modes, where the biases do not indicate a statistically significant time dependence. The land nadir/glint mode has the lowest bias value of −0.03 ± 0.85 ppm.

54 ENVIRONMENTAL SCIENCES

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)

Impacts of benchmarking choices on inferred model skill of the Arctic–Boreal terrestrial carbon cycle

Abstract Land surface models require continuous validation against observations to improve and reduce simulation uncertainty. However, inferred model performance can be heavily influenced by subjective choices made in the selection and application of observational data products. A key area often misrepresented by models is the Arctic–Boreal region, which is a potential tipping point region in Earth’s climate system due to large permafrost carbon stocks that are vulnerable to release with climate warming. We use the International Land Model Benchmarking (ILAMB) framework to evaluate how the model skill of TRENDY-v9 models varies based on the choice of observational-based benchmark and how benchmarks are applied in model evaluation. This analysis uses global datasets integrated into ILAMB and new, regionally-specific observational products from the Arctic–Boreal Vulnerability Experiment. Our results cover the overall time period of 1979–2019 and show that model scores can vary substantially depending on the data product applied, with higher model scores indicating better model performance against observations. The lowest model scores occur when benchmarked against regional, compared to global, datasets. We also evaluate observed and modeled functional relationships between ecosystem respiration and air temperature and between gross primary production and precipitation. Here, we find that the magnitude and shape of the responses are strongly impacted by the choice of observational dataset and the approach used to construct the functional relationship benchmark. These results suggest that model evaluation studies could conclude a false sense of model skill if only using a single benchmark data product or if not applying regional data products when performing a regional model analysis. Collectively, our findings highlight the influence of benchmarking choices on model evaluation and point to the need for benchmarking guidelines when assessing model skill.

Poe, Jeralyn (ORCID:0000000318495278)

NGEE Arctic Integrated Modeling (IM3): Improved snow-vegetation interaction

This data product represents the integration of new code capability for arctic tundra snow-vegetation-terrain interactions into the Energy Exascale Earth System Model (E3SM), through the E3SM Land Model (ELM) component. This code integration is the result of collaborative effort between the NGEE Arctic project and the E3SM project. The NGEE Arctic project developed a total of six Integrated Modeling (IM) modules informed by observations and experiments. New ELM capability represented by this data product (IM3) falls into three categories: 1) Downscaling from gridcell to topographic unit level when working through the existing coupler bypass code. 2) Four new parameters (taper, stocking, bendresist, and vegshape) have been added to ELM to allow for flexible definition of snow-vegetation interactions. 3) Vegshape and bendresist parameters are used to calculate the fraction of leaf area and/or stem area buried by snow for a given snow depth. This data record consists of a single document (pdf format) that describes the theoretical basis for the snow-vegetation-terrain interactions added to ELM, and describes the modifications made to the ELM code. The Methods section of this metadata record includes a link to the public E3SM code repository where the exact code modifications as integrated in E3SM can be accessed. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Thornton, Peter E [ORNL] (ORCID:0000000247595158)