Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “production data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Raw ERCOT 60-Day SCED Disclosure Reports

This dataset contains ERCOT plant-level wind power production data at a 15-minute time resolution. Data at ERCOT are called 60-Day SCED Disclosure Report; data product number NP3-965-ER. Raw data are in local time. DST is handled as "skip hour" (12a, 1a, 3a,...) in spring, "extra hour" (12a, 1a, 2a, 2a, 3a,...) in fall.

17 WIND ENERGY↗

Global Carbon Budget 2021

Abstract. Accurate assessment of anthropogenic carbon dioxide (CO2) emissions and their redistribution among the atmosphere, ocean, and terrestrial biosphere in a changing climate is critical to better understand the global carbon cycle, support the development of climate policies, and project future climate change. Here we describe and synthesize datasets and methodology to quantify the five major components of the global carbon budget and their uncertainties. Fossil CO2 emissions (EFOS) are based on energy statistics and cement production data, while emissions from land-use change (ELUC), mainly deforestation, are based on land use and land-use change data and bookkeeping models. Atmospheric CO2 concentration is measured directly, and its growth rate (GATM) is computed from the annual changes in concentration. The ocean CO2 sink (SOCEAN) is estimated with global ocean biogeochemistry models and observation-based data products. The terrestrial CO2 sink (SLAND) is estimated with dynamic global vegetation models. The resulting carbon budget imbalance (BIM), the difference between the estimated total emissions and the estimated changes in the atmosphere, ocean, and terrestrial biosphere, is a measure of imperfect data and understanding of the contemporary carbon cycle. All uncertainties are reported as ±1σ. For the first time, an approach is shown to reconcile the difference in our ELUC estimate with the one from national greenhouse gas inventories, supporting the assessment of collective countries' climate progress. For the year 2020, EFOS declined by 5.4 % relative to 2019, with fossil emissions at 9.5 ± 0.5 GtC yr−1 (9.3 ± 0.5 GtC yr−1 when the cement carbonation sink is included), and ELUC was 0.9 ± 0.7 GtC yr−1, for a total anthropogenic CO2 emission of 10.2 ± 0.8 GtC yr−1 (37.4 ± 2.9 GtCO2). Also, for 2020, GATM was 5.0 ± 0.2 GtC yr−1 (2.4 ± 0.1 ppm yr−1), SOCEAN was 3.0 ± 0.4 GtC yr−1, and SLAND was 2.9 ± 1 GtC yr−1, with a BIM of −0.8 GtC yr−1. The global atmospheric CO2 concentration averaged over 2020 reached 412.45 ± 0.1 ppm. Preliminary data for 2021 suggest a rebound in EFOS relative to 2020 of +4.8 % (4.2 % to 5.4 %) globally. Overall, the mean and trend in the components of the global carbon budget are consistently estimated over the period 1959–2020, but discrepancies of up to 1 GtC yr−1 persist for the representation of annual to semi-decadal variability in CO2 fluxes. Comparison of estimates from multiple approaches and observations shows (1) a persistent large uncertainty in the estimate of land-use changes emissions, (2) a low agreement between the different methods on the magnitude of the land CO2 flux in the northern extra-tropics, and (3) a discrepancy between the different methods on the strength of the ocean sink over the last decade. This living data update documents changes in the methods and datasets used in this new global carbon budget and the progress in understanding of the global carbon cycle compared with previous publications of this dataset (Friedlingstein et al., 2020, 2019; Le Quéré et al., 2018b, a, 2016, 2015b, a, 2014, 2013). The data presented in this work are available at https://doi.org/10.18160/gcp-2021 (Friedlingstein et al., 2021).

Friedlingstein, Pierre (ORCID:0000000333094739)↗

Vera C. Rubin Observatory Prompt Products: alert packets data

Data products produced by prompt and daily processing of images obtained in the Legacy Survey of Space and Time. These include realtime alerts sent to community alert brokers, newly-discovered Solar System Objects reported to the Minor Planet Center, processed visit and difference images, and source catalogs. Prompt Products are not a static single data release but continually grow throughout the ten-year LSST survey. This dataset is a subset of the full data release consisting of a dataset named alert packets. This dataset contains measurements for 5-sigma sources detected in difference images that were issued to the community brokers.

79 ASTRONOMY AND ASTROPHYSICS↗

dbprocessing

dbprocessing is a python framework for automating data processing pipelines. The dbprocessing package uses a configuration file to define the input, intermediate, and output data products for a particular data type as well as the processes that connect those data products. In addition, dbprocessing uses python scripts called inspectors to determine if a particular file matches a configured data product. If a particular data product is found, it is ingested into the dbprocessing database (sqlite3 or postgres) which then triggers all of the chained processes to the final data product. While the framework is written in python, it can run software in any language but may require a “wrapper” to translate the dbprocessing command line arguments to the form expected by the software.

Walker, Andrew↗

Machine Learned Empirical Numerical Integrator from Simulated Data

Recently, a number of state-of-the-art surrogate machine learning (ML) models have been designed for global weather and climate prediction, which have been trained using reanalysis data products. Reanalysis data products are constructed using numerical model simulations that combine numerical integration of partial differential equations and parameterization schemes. These products are typically only archived and made available using coarsened spatial and temporal resolutions. This study explores the impact of the numerical generation methods used to produce the training datasets and the temporal resolution of those datasets on machine learning surrogate models. Using the nonlinear vector autoregression (NVAR) machine as an explainable ML technique, simple dynamical systems are emulated with ML models trained on data produced by three classical numerical integration schemes. NVAR is validated as a skillful ML method, capable of producing accurate predictions and, more importantly, reconstructing both the underlying dynamics and the numerical integration scheme used to generate the training data. However, the machine fails to generalize predictions on unseen test data generated by different numerical integration schemes, despite the underlying dynamical system being the same. This result provides a word of caution for the growing field of machine learning emulation of weather and climate dynamics. Furthermore, we illustrate using NVAR that training on temporally coarsened data may increase the required complexity of ML models and potentially introduce new numerical challenges. Finally, we discover that empirical integration schemes with arbitrary time-stepping sizes can be constructed directly from the data, which implies a potential for the development of empirical numerical integration schemes.

54 ENVIRONMENTAL SCIENCES↗

Understanding Solar Photovoltaic System Performance: An Assessment of 75 Federal Photovoltaic Systems

This report presents a performance analysis of 75 photovoltaic systems based on PV system production data collected as part of a FEMP Federal PV Performance Assessment project combined with co-incident insolation, and ambient temperature to analyze how actual performance compares with a performance model. FEMP collaborated with 17 Federal agencies and sub-agencies to collect the information required to analyze the performance of each system. The systems represent a total capacity of 30,714 kW and range in size from 1 kW to 4,043 kW, with an average size of 410 kW, and were installed between 2011 and 2020. The data is analyzed for Key Performance Indicators, Availability, Performance Ratio and Energy Ratio by comparing the measured production data to model production data. The System Advisor Model (SAM) combines a description of the system (such as inverter capacity, de-rating for temperature, balance-of-system efficiency) with environmental parameters (coincident solar and temperature data) to calculate predicted performance. The performance metrics are calculated by lining up the measured production data with the model estimate on an hour-by-hour, day-by-day, or month-by-month basis (depending on the interval resolution of the production data). A report with system description, photo of the system, special assumptions made for the site, graph of measured production and model production, table of key performance indicators, and links to O&M resources that might improve performance was produced and delivered to site and agency staff with a short on-line briefing.

14 SOLAR ENERGY↗

Statistical Validation of Multiple Related Data Sets—Case Study Using Interstellar Boundary Explorer Satellite Data

Abstract Space scientists often face the question of whether data collected by different instruments are measurements of the same source population. This paper proposes a statistical validation method for evaluating the agreement between such related data sets. It offers a detailed case study focused on validating a new data set from the Interstellar Boundary Explorer (IBEX) mission, which serves as a practical how-to guide for similar analyses. Since 2008, the IBEX satellite has been gathering data on heliospheric energetic neutral atoms (ENAs) while being exposed to various sources of background noise, such as cosmic rays and solar energetic particles. The IBEX mission initially released only a qualified triple-coincidence (qABC) data product, which was designed to provide observations of ENAs free of background contamination. Further measurements revealed that the qABC data were in fact susceptible to contamination, having relatively low ENA counts and high background rates. To mitigate this issue, the mission team recently considered releasing a certain qualified double-coincidence (qBC) data product, which has roughly twice the detection rate of the qABC data product. This paper presents a simulation-based validation of the new qBC data product against the already-released qABC data product. The results show that the qBCs can plausibly be said to be measuring the same source population as the qABCs up to an average absolute deviation of 3.6%. Visual diagnostics provide additional confirmation of source rate coherence across data products. The framework introduced here is general and can be applied to other validation problems both within and outside the field of space physics.

79 ASTRONOMY AND ASTROPHYSICS↗

Advancing Multiscale Simulation of Plasma-Surface Interfaces

We report the development of an atomistic-informed, surface-state-dependent predictive model for particle exchange in a carbon-tungsten plasma-surface interface. The predictive model uses machine learning (ML) techniques to learn the energy and angular distributions for particle exchange and rate functions for surface state evolution from molecular dynamics simulations of cumulative bombardment of tungsten by energetic carbon ions. Each predictive component is sensitive to the energy and trajectory of incident plasma species and the surface state. The surface state is represented by a set of surface state descriptors, which were derived from the atomistic surface state for each independent carbon bombardment event. These descriptors are representative of the composition and degree of amorphization of the outermost angstrom of surface material and were chosen to optimize predictive performance for particle exchange at the interface. The distributions for particle exchange (reflection/sputtering) are demonstrated to vary with each surface state descriptor, motivating the development of surface-state-dependent particle exchange models for plasma simulations. The performance of various ML methods was compared, including polynomial quantile regression, artificial neural networks, k-nearest neighbors, and random forest algorithms, with polynomial regression performing the best for interpolation and extrapolation of learned relationships. In addition to the particle exchange model, a neutral network was developed and used to identify data sufficiency throughout surface descriptor space, which will enable real-time feedback during future data production to ensure data is produced where it is most needed, and we provide commentary on improvements to the data production workflow for future endeavors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

AGR TRISO Fuel Fission Product Release Data Summary

Knowledge of fission product retention in and release from TRISO fuel under normal and off-normal conditions is needed for reactor safety analyses. This is important for coated-particle fuel used in high-temperature reactors relying on the functional containment strategy. Data on the release and retention of key fission products (e.g., Ag-110m, Cs-134, Eu-154, and Sr-90) in AGR UCO TRISO fuels have been summarized in this report, and empirical relationships with respect to time and temperature were developed. This included fission product accumulation in the OPyC and compact graphitic matrix during irradiation, release from compacts during irradiation, and release during post-irradiation safety testing at temperatures from 1600-1800°C. The frequencies of SiC failure and TRISO coating failure from irradiation and post-irradiation safety testing were also summarized as they have bearing on the quantities of Cs release from the fuel. In a forthcoming publication, a framework for combining and using these empirical relationships as part of a source term analysis will be presented.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Concurrent conditions access across validity intervals in CMSSW

The CMS software system, known as CMSSW, has a generalized conditions, calibration, and geometry data products system called the EventSetup. The EventSetup caches results of reading or calculating data products based on the 'interval of validity', IOV, which is based on the time period for which that data product is appropriate. With the original single threaded CMSSW framework, updating only on an IOV boundary meant we only required memory for a single data product of a given type at any time during the program execution. In 2016 CMS transitioned to using a multi-threaded framework as away to save on memory during processing. This was accomplished by amortizing the memory cost of EventSetup data products across multiple concurrent events. To initially accomplish that goal required synchronizing event processing across IOV boundaries, thereby decreasing the scalability of the system. In this presentation we will explain how we used 'limited concurrent task queues' to allow concurrent IOVs while still being able to limit the memory utilized.

Dagenhart, David↗

Comparisons of the v11.1 Orbiting Carbon Observatory‐2 (OCO‐2) X CO2 Measurements With GGG2020 TCCON

The Orbiting Carbon Observatory 2 (OCO-2) is NASA's first Earth observation satellite mission dedicated to studying the sources and sinks of carbon dioxide (CO 2 ) on a global scale. The observations of reflected sunlight are inverted in a retrieval algorithm to produce estimates of the dry air mole-fractions of CO 2 (X CO2 ). The OCO-2 Level 2 data release, version 11.1 (v11.1) retrievals from the Atmospheric Carbon Observations from Space (ACOS) algorithm, includes significant improvements in the X CO2 data product compared to older OCO-2 data versions. This work compares the v11.1 X CO2 from OCO-2 against X CO2 estimates collected from a global ground-based network known as the Total Carbon Column Observing Network (TCCON), OCO-2's primary validation source. The OCO-2 project provides a version of the Level 2 data product, called “lite” files that include calibrated and bias-corrected XCO2 values, accessible together with all OCO-2 data products through the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). This work shows that OCO-2 X CO2 observations made between September 2014 and December 2023, after quality filtering and the application of an averaging kernel correction, agree well with coincident TCCON data for all OCO-2 observational modes of land (nadir, glint, target) and ocean (glint). The aggregated, bias-corrected, and quality-filtered absolute average bias values are less than or equal to 0.20 parts per million (ppm) globally for all OCO-2 observation modes, where the biases do not indicate a statistically significant time dependence. The land nadir/glint mode has the lowest bias value of −0.03 ± 0.85 ppm.

54 ENVIRONMENTAL SCIENCES↗

ECOSTRESS: NASA's Next Generation Mission to Measure Evapotranspiration From the International Space Station

Abstract The ECOsystem Spaceborne Thermal Radiometer Experiment on Space Station (ECOSTRESS) was launched to the International Space Station on 29 June 2018 by the National Aeronautics and Space Administration (NASA). The primary science focus of ECOSTRESS is centered on evapotranspiration (ET), which is produced as Level‐3 (L3) latent heat flux ( LE ) data products. These data are generated from the Level‐2 land surface temperature and emissivity product (L2_LSTE), in conjunction with ancillary surface and atmospheric data. Here, we provide the first validation (Stage 1, preliminary) of the global ECOSTRESS clear‐sky ET product (L3_ET_PT‐JPL, Version 6.0) against LE measurements at 82 eddy covariance sites around the world. Overall, the ECOSTRESS ET product performs well against the site measurements (clear‐sky instantaneous/time of overpass: r 2 = 0.88; overall bias = 8%; normalized root‐mean‐square error, RMSE = 6%). ET uncertainty was generally consistent across climate zones, biome types, and times of day (ECOSTRESS samples the diurnal cycle), though temperate sites are overrepresented. The 70‐m‐high spatial resolution of ECOSTRESS improved correlations by 85%, and RMSE by 62%, relative to 1‐km pixels. This paper serves as a reference for the ECOSTRESS L3 ET accuracy and Stage 1 validation status for subsequent science that follows using these data.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of daily gridded climate products using in situ FLUXNET data and tree growth modeling

Gridded climate data products have facilitated research in climate and ecology by providing meteorological data continuously across large spatial scales. However, the sensitivity of scientific outcomes to dataset choice remains poorly understood, and evaluation using station-based records can favor datasets built heavily on weather stations. Here, we evaluate seven high-resolution daily gridded datasets covering the contiguous United States using independent meteorology from the FLUXNET2015 dataset, with a focus on the implications of dataset choice for process-based tree growth modeling. We find that gridded products tend to capture temperature accurately while consistently overestimating the magnitude and frequency of precipitation and its extremes. Moreover, datasets vary in how they define a ‘day,’ which significantly affects temporal alignment with FLUXNET2015 observations. Despite differences among the datasets, the interannual variability in tree ring simulations is insensitive to dataset choice, likely because daily-scale biases are averaged out through accumulated growth across several months. However, inaccuracies in temperature and precipitation can significantly bias modeled xylem cell production, with systematically higher annual precipitation in the gridded datasets leading to greater xylem production compared to simulations using in situ data. Our results suggest that model applications, especially those that integrate to time scales longer than one day, are likely insensitive to climate dataset choice, but applications that are sensitive to daily climate variations or to absolute climate values need to carefully consider biases in gridded climate products.

54 ENVIRONMENTAL SCIENCES↗

The disCO2ver Platform: Curating Data and Tools for Geologic Carbon Sequestration and Deep Subsurface Research Systems

The U.S. DOE National Energy Technology Laboratory has invested 12+ years of development into the data repository and digital laboratory, the Energy Data eXchange (EDX, edx.netl.doe.gov). Supporting a variety of research areas across the DOE Office of Fossil Energy and Carbon Management, the platform has successfully curated and preserved thousands of data products from DOE research. The Carbon Storage Program has successfully supported data curation, upload, and publishing of data products on EDX for many years, demonstrating a success story of how resources like EDX can effectively help with long term preservation and publishing of DOE data products. EDX continues to shift towards cloud-supported infrastructure, taking a hybrid approach combining on-premises compute and storage integrated with cloud-hosted services. The integration of cloud compute and hybrid architecture enables the development of EDX-hosted platforms that tailor the data and tools hosted on them to a specific community, enables implementation of machine learning tools for data discovery and filtering, and enables the hosting of virtual (online user interface) tools. Geologic carbon sequestration (GCS) research continues to scale up in response to the current administration goals to reduce greenhouse gas emissions and transition the energy economy. Over the last year, EDX’s disCO2ver platform has been developed in response to the need for access to data products and tools to support the scaling up of GCS research. disCO2ver provides access to data resources and tools, produced by DOE and outside authoritative external resources. The platform also provides a user-access control component for the virtualization and cloud hosting of tools. Tools that need to be virtualized, to eliminate the need for users to download the tool and use local compute resources, is essential to supporting big-data analysis and machine learning that is becoming common place in carbon storage modeling, risk analysis, and data publishing practices. This talk will review the EDX’s disCO2ver platform and the current work ongoing to curate data and tools to support GCS and deep subsurface systems research.

Morkner, Paige↗

Evaluation of distributed process-based hydrologic model performance using only a priori information to define model inputs

Fully distributed, integrated surface–subsurface hydrological models (ISSHMs) have seen renewed interest due to availability of better software, high performance computing facilities, and high-resolution, spatially extensive data products. ISSHMs are valuable as tools for advancing system understanding as they can resolve multiple processes defined on the plot scale including three-dimensional interaction of surface water and groundwater. Here, we evaluated the performance of an ISSHM, the Advanced Terrestrial Simulator (ATS), on seven diverse catchments across the continental US using widely available data products to define model inputs without calibration. We compare the ATS-simulated streamflow and evapotranspiration with gauge observations and MODIS-derived evapotranspiration, respectively. Using the Kling-Gupta Efficiency (KGE) as metric, ATS with default data products performed reasonably well at 6 of 7 catchments for streamflow. However, in one of those 6 catchments ATS had poor performance on baseflow and ATS’s overall performance was thus judged to be inadequate despite the acceptable KGE. ATS performance for evapotranspiration was good in all 7 catchments using default data products. In the two catchments where ATS streamflow performance using default data products was not acceptable, the performance was significantly improved by using local information on subsurface properties below the soil. We also compare the model-simulated streamflow and evapotranspiration with the Sacramento soil moisture accounting (SAC-SMA) model, a semi-distributed model that was calibrated on a catchment-by-catchment basis. Uncalibrated ATS performance is comparable to the calibrated SAC-SMA model in terms of streamflow while ATS performance is similar to or better (much better in certain catchments) in reproducing MODIS-derived evapotranspiration. Reasonably good performance of ATS without catchment-specific calibration provides new confidence in the ISSHM class of models and community data products as tools for advancing understanding of watershed function in a changing environment.

54 ENVIRONMENTAL SCIENCES↗

The Lick Observatory Supernova Search follow-up program: photometry data release of 70 SESNe

We present BVRI and unfiltered (Clear) light curves of 70 stripped-envelope supernovae (SESNe), observed between 2003 and 2020, from the Lick Observatory Supernova Search follow-up program. Our SESN sample consists of 19 spectroscopically normal SNe Ib, 2 peculiar SNe Ib, six SNe Ibn, 14 normal SNe Ic, 1 peculiar SN Ic, 10 SNe Ic-BL, 15 SNe IIb, 1 ambiguous SN IIb/Ib/c, and 2 superluminous SNe. Our follow-up photometry has (on a per-SN basis) a mean coverage of 81 photometric points (median of 58 points) and a mean cadence of 3.6 d (median of 1.2 d). From our full sample, a subset of 38 SNe have pre-maximum coverage in at least one passband, allowing for the peak brightness of each SN in this subset to be quantitatively determined. We describe our data collection and processing techniques, with emphasis toward our automated photometry pipeline, from which we derive publicly available data products to enable and encourage further study by the community. Using these data products, we derive host-galaxy extinction values through the empirical colour evolution relationship and, for the first time, produce accurate rise-time measurements for a large sample of SESNe in both optical and infrared passbands. By modelling multiband light curves, we find that SNe Ic tend to have lower ejecta masses and lower ejecta velocities than SNe Ib and IIb, but higher 56Ni masses.

Lick Observatory↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗

Global Temperature Responses to Large Tropical Volcanic Eruptions in Paleo Data Assimilation Products and Climate Model Simulations Over the Last Millennium

Abstract Large volcanic eruptions are one of the dominant perturbations to global and regional atmospheric temperatures on timescales of years to decades. Discrepancies remain, however, in the estimated magnitude and persistence of the surface temperature cooling caused by volcanic eruptions, as characterized by paleoclimatic proxies and climate models. We investigate these discrepancies in the context of large tropical eruptions over the Last Millennium using two state‐of‐the‐art data assimilation products, the Paleo Hydrodynamics Data Assimilation product (PHYDA) and the Last Millennium Reanalysis (LMR), and simulations from the National Center for Atmospheric Research Community Earth System Model‐Last Millennium Ensemble (NCAR CESM‐LME). We find that PHYDA and LMR estimate mean global and hemispheric cooling that is similar in magnitude and persistence once effects from eruptions occurring in short succession are removed. The estimates also compare well to Northern‐Hemisphere reconstructions based solely or partially on tree‐ring density, which have been proposed as the most accurate proxy estimates of surface cooling due to volcanism. All proxy‐based estimates also agree well with the magnitude of the mean cooling simulated by the CESM‐LME. Differences remain, however, in the spatial patterns of the temperature responses in the PHYDA, LMR, and the CESM‐LME. The duration of cooling anomalies also persists for several years longer in the PHYDA and LMR relative to the CESM‐LME. Our results demonstrate progress in resolving discrepancies between proxy‐ and model‐based estimates of temperature responses to volcanism, but also indicate these estimates must be further reconciled to better characterize the risks of future volcanic eruptions.

Tejedor, E.↗