Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Molecular simulation data for 'Data-guided Multi-Map variables for ensemble refinement of molecular movies'

These trajectories, scripts, and analysis performed on Summit underly the work published as 'Data-guided Multi-Map variables for ensemble refinement of molecular movies'. The trajectories include equilibrium and non-equilibrium sampling of ADK, CODH, and FLPP3, the scripts used to build the systems, and the scripts used to analyze the output. The directory structure is explained further in an internal README file.

59 BASIC BIOLOGICAL SCIENCES↗

Data-Driven Load Diversity and Variability Modeling for Quasi-Static Time-Series Simulation on Distribution Feeders

This paper presents a data-driven load modeling methodology for distribution system quasi-static time-series (QSTS) simulation considering both diversity and variability characteristics of distribution loads. Based on our previous work in [1]-[2], a variability library and diversity library have been established based on the realistic high-resolution data collected from actual utility feeders. Given the load profile for the start-of-circuit load of a feeder, the loads on the feeder nodes can be modeled with both diversity and variability instead of being directly scaled from the substation load profile according to the distribution allocation factors. With diversified load models, the load-induced impact on the feeder operation characteristics, such as voltage ramp and regulator operations, can be better considered in QSTS simulation. The proposed modeling methodology has been tested on both the IEEE 123-bus feeder and an actual utility feeder model, and the simulation results have demonstrated the merits of deploying the proposed load modeling methodology.

24 POWER TRANSMISSION AND DISTRIBUTION↗

First Sagittarius A* Event Horizon Telescope Results. IV. Variability, Morphology, and Black Hole Mass

In this paper we quantify the temporal variability and image morphology of the horizon-scale emission from Sgr A*, as observed by the EHT in 2017 April at a wavelength of 1.3 mm. We find that the Sgr A* data exhibit variability that exceeds what can be explained by the uncertainties in the data or by the effects of interstellar scattering. The magnitude of this variability can be a substantial fraction of the correlated flux density, reaching ~100% on some baselines. Through an exploration of simple geometric source models, we demonstrate that ring-like morphologies provide better fits to the Sgr A* data than do other morphologies with comparable complexity. We develop two strategies for fitting static geometric ring models to the time-variable Sgr A* data; one strategy fits models to short segments of data over which the source is static and averages these independent fits, while the other fits models to the full data set using a parametric model for the structural variability power spectrum around the average source structure. Both geometric modeling and image-domain feature extraction techniques determine the ring diameter to be 51.8 ± 2.3 μas (68% credible intervals), with the ring thickness constrained to have an FWHM between ~30% and 50% of the ring diameter. To bring the diameter measurements to a common physical scale, we calibrate them using synthetic data generated from GRMHD simulations. This calibration constrains the angular size of the gravitational radius to be ${4.8}_{-0.7}^{+1.4}$ μas, which we combine with an independent distance measurement from maser parallaxes to determine the mass of Sgr A* to be ${4.0}_{-0.6}^{+1.1}\times {10}^{6}$ M⊙.

79 ASTRONOMY AND ASTROPHYSICS↗

Subsurface redox potential and water level at the Elkhorn Slough NERR

This resource contains various hydrological, and geochemical data from Elkhorn Slough National Estuarine Research Reserve from the years 2020 and 2021. These data have been used to assess how continuous measurements of environmental variables. The data can be used to understand processes at timescales over which biochemical transformations can happen. Especially, the data were used to explain the local subsurface hydrology, and its implication, in an experimental transect in a coasta estuary. Water level and water temperature were measured in-situ with Solinst pressure transducer loggers (Ontario, Canada). Redox potential was collected using in-situ, redox sensors (Paleoterra, The Netherlands) connected to CR1000X Campbell data loggers (Logan, Utah). Meteorological data was gathered from the Elkhorn Slough meteorological station.

54 ENVIRONMENTAL SCIENCES↗

Front-of-Meter Model Results

These files contains aggregations of key variables from the NREL Distributed Wind Futures Study using full parcel level data. These variables describe total technical and economic potential for distributed wind turbine deployment. Aggregations are available at the (1) county, (2) zipcode (zip code tabulation area or zcta), and (3) US Census block group level. Each scenario is coded with the scenario name (e.g., baseline) and year (e.g., 2022). Those files postfixed with 'econpot' contain results for only those parcels that are economically viable while the files postfixed with 'techpot' include results for all parcels that are technically feasible. Hence these correspond to technoeconomic and technical potential respectively. The data are available as CSV or Geopackage. Columns in the files are as follows: * geoid: geographic identifier (FIPS code or similar) * min_techpot_sum_kw: technical potential for all parcels in kW using turbines downsized to demand when appropriate * max_techpot_sum_kw: technical potential for all parcels in kW without downsizing turbines * aep_sum_kwh: annual energy production estimate in kWh * cf_mean_ratio: mean capacity factor * lcoe_mean_cents_per_kwh: mean levelized cost of energy for parcels in geography in cents per kWh * lcoe_std_cents_per_kwh: standard deviation of the above * parcel_area_sum_acres: total area of viable parcels in acres * n_turbines: number of cited turbines (one per viable parcel currently) Note: These are preliminary results from the full-parcel 2024 update of the Distributed Wind Energy Futures study. Please take care when making use of the data, and feel free to contact the team with any questions. Full documentation in support of these data is in progress and will follow.

17 WIND ENERGY↗

Behind-the-Meter Model Results

These files contains aggregations of key variables from the NREL Distributed Wind Futures Study using full parcel level data. These variables describe total technical and economic potential for distributed wind turbine deployment. Aggregations are available at the (1) county, (2) zipcode (zip code tabulation area or zcta), and (3) US Census block group level. Each scenario is coded with the scenario name (e.g., baseline) and year (e.g., 2022). Those files postfixed with 'econpot' contain results for only those parcels that are economically viable while the files postfixed with 'techpot' include results for all parcels that are technically feasible. Hence these correspond to technoeconomic and technical potential respectively. The data are available as CSV or Geopackage. Columns in the files are as follows: * geoid: geographic identifier (FIPS code or similar) * min_techpot_sum_kw: technical potential for all parcels in kW using turbines downsized to demand when appropriate * max_techpot_sum_kw: technical potential for all parcels in kW without downsizing turbines * aep_sum_kwh: annual energy production estimate in kWh * cf_mean_ratio: mean capacity factor * lcoe_mean_cents_per_kwh: mean levelized cost of energy for parcels in geography in cents per kWh * lcoe_std_cents_per_kwh: standard deviation of the above * parcel_area_sum_acres: total area of viable parcels in acres * n_turbines: number of cited turbines (one per viable parcel currently)

17 WIND ENERGY↗

FunM2C: A Filter for Uncertainty Visualization of Multivariate Data on Multi-Core Devices

Uncertainty visualization is an emerging research topic in data visualization because neglecting uncertainty in visualization can lead to inaccurate assessments. In this paper, we study the propagation of multivariate data uncertainty in visualization. Although there have been a few advancements in probabilistic uncertainty visualization of multivariate data, three critical challenges remain to be addressed. First, the state-of-the-art probabilistic uncertainty visualization framework is limited to bivariate data (two variables). Second, existing uncertainty visualization algorithms use computationally intensive techniques and lack support for cross-platform portability. Third, as a consequence of the computational expense, integration into production visualization tools is impractical. In this work, we address all three issues and make a threefold contribution. First, we take a step to generalize the state-of-the-art probabilistic framework for bivariate data to multivariate data with an arbitrary number of variables. Second, through utilization of VTK-m’s shared-memory parallelism and cross-platform compatibility features, we demonstrate acceleration of multivariate uncertainty visualization on different many-core architectures, including OpenMP and AMD GPUs. Third, we demonstrate the integration of our algorithms with the ParaView software. We demonstrate the utility of our algorithms through experiments on multivariate simulation data with three and four variables.

Hari, Gautam↗

ESS-DIVE Reporting Format for Dataset Package Metadata

ESS-DIVE’s (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) dataset metadata reporting format is intended to compile information about a dataset (e.g., title, description, funding sources) that can enable reuse of data submitted to the ESS-DIVE data repository. The files contained in this dataset include instructions (dataset_metadata_guide.md and README.md) that can be used to understand the types of metadata ESS-DIVE collects. The data dictionary (dd.csv) follows ESS-DIVE’s file-level metadata reporting format and includes brief descriptions about each element of the dataset metadata reporting format. This dataset also includes a terminology crosswalk (dataset_metadata_crosswalk.csv) that shows how ESS-DIVE’s metadata reporting format maps onto other existing metadata standards and reporting formats.Data contributors to ESS-DIVE can provide this metadata by manual entry using a web form or programmatically via ESS-DIVE’s API (Application Programming Interface). A metadata template (dataset_metadata_template.docx or dataset_metadata_template.pdf) can be used to collaboratively compile metadata before providing it to ESS-DIVE.Since being incorporated into ESS-DIVE’s data submission user interface, ESS-DIVE’s dataset metadata reporting format, has enabled features like automated metadata quality checks, and dissemination of ESS-DIVE datasets onto other data platforms including Google Dataset Search and DataCite.

54 ENVIRONMENTAL SCIENCES↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 2 Sensor Data v2-1

This is the version v2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Please see v2-1 TEMPEST L2 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods. * Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021 * TEMPEST 1: June 22, 2022 * TEMPEST 2: June 6-7, 2023 * TEMPEST 3: June 11-13, 2024

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

COMPASS-FME Synoptic Sites Level 2 Sensor Data v2-1

This is the version 2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our synoptic field sites. COMPASS-FME is studying sites in two distinct regions, the Chesapeake Bay and the Western Lake Erie Basin. We established the network at seven "synoptic" (observational) sites along the Chesapeake Bay and Lake Erie coastlines, collectively generating over three million observations per month, to track and comprehend environmental changes where land and water intersect. Additionally, the two regions provide an interesting contrast of saltwater and freshwater coasts that allow us to differentiate the impacts of inundation and coastal water chemistries in two nationally important coastal systems. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. Please see v2-1 L2 Sensor Package QStart.pdf for detailed information on data package structure, temporal coverage, and versioning.

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

Evaluating the Nation's Pipeline Infrastructure with NETL's Advanced Infrastructure Integrity Model (AIIM)

This poster is a part of BIL-EDX4CCS Task 36: Advanced Infrastructure Integrity Modeling to Evaluate Existing Energy Infrastructure Reusability and Risk, the goal of which is to produce a smart tool that will assess existing energy infrastructure reusability and risk using the Advanced Infrastructure Integrity Model (AIIM). This model forecasts lifespan and potential risk using a multitude of factors such as incidents reports, structural characteristics, and the surrounding environment. The project aims to provide scientific insights for a better understanding of carbon storage (CS), potential to support CS stakeholder needs, national decarbonization, and mitigating climate change. AIIM will utilize an energy infrastructure database as its input, developed by acquiring publicly available data as well as NETL derived products. These resources include incidents, geohazards, and infrastructure variables. Soil data in the form of rasters and pipeline incident reports were processed and a script was developed to count the number of times features such as roads, railroads, and rivers intersected with pipeline segments which were then converted to points. Distance to oil and natural gas wells, petroleum ports, intermodal freight facilities, and geologic structures were also calculated. After data preparation and quality control was completed, the data was integrated into the pipeline points. Once models are complete, a smart tool will be created in the form of an online dashboard.

Malay, Caleb↗

Training data selection for event classification in a highly variable environment

A problem of interest for nuclear nonproliferation is monitoring activities at nuclear facilities, where proliferation events may only take place a few times and often under variable conditions. Machine learning has revolutionized data analytics by enabling the use of measurable signatures to generate predictive models of facility operations. However, traditional methods for training these models require large, reliable data sets with labeled observations, a challenge for nonproliferation. Highly variable conditions further complicate this as events from training data may have occurred in conditions quite different from the event of interest. Our hypothesis is that when events occur in a highly variable environment, careful training data selection for each test event could outperform the standard approach of using all available training data. We developed a method to optimize training data selection for the given test event and applied it to predicting the power level of the High Flux Isotope Reactor (HFIR) at Oak Ridge National Laboratory. In this study, the reactor startup exhibits variability between occurrences due to natural variability in environmental conditions and operational procedures. Using a combination of analysis techniques, a similitude assessment was performed on data collected from HFIR to isolate clusters that were optimal for training a predictive model. Concepts such as dynamic time warping and Jaccard similarity were used in conjunction with clustering analysis. In order to validate this approach, the model was trained on every combination of unique training events and the predictive performance was compared to the performance using a subset of the training data selected by isolated clusters found through the similitude assessment.

Iyer, A↗

A Nonstationary and Non-Gaussian Moving Average Model for Solar Irradiance

Historically, power has flowed from large power plants to customers. Increasing penetration of distributed energy resources such as solar power from rooftop photovoltaic has made the distribution network a two-way-street with power being generated at the customer level. The incorporation of renewables introduces additional uncertainty and variability into the power grid. Distribution network operation studies are being adapted to include renewables; however, such studies require high quality solar irradiance data that adequately reflect realistic meteorological variability. Data from satellite-based products are spatially complete, but temporally coarse, whereas solar irradiances exhibit high frequency variation at very fine timescales. We propose a new stochastic method for temporally downscaling global horizontal irradiance (GHI) to 1 min resolution, but we do not consider the spatial aspect due to limited availability of the in situ irradiance measurements. Solar irradiance's first and second-order structures vary diurnally and seasonally, and our model adapts to such nonstationarity. Empirical irradiance data exhibits highly non-Gaussian behavior; we develop a nonstationary and non-Gaussian moving average model that is shown to capture realistic solar variability at multiple timescales. We also propose a new estimation scheme based on Cholesky factors of empirical autocovariance matrices, bypassing difficult and inaccessible likelihood-based approaches. The model is demonstrated for a case study of three locations that are located in diverse climates through the United States. The model is compared against competitors from the literature and is shown to provide better uncertainty and variability quantification on testing data.

Cholesky factor↗

Unifying and benchmarking state-of-the-art quantum error mitigation techniques

Error mitigation is an essential component of achieving a practical quantum advantage in the near term, and a number of different approaches have been proposed. In this work, we recognize that many state-of-the-art error mitigation methods share a common feature: they are data-driven, employing classical data obtained from runs of different quantum circuits. For example, Zero-noise extrapolation (ZNE) uses variable noise data and Clifford-data regression (CDR) uses data from near-Clifford circuits. We show that Virtual Distillation (VD) can be viewed in a similar manner by considering classical data produced from different numbers of state preparations. Observing this fact allows us to unify these three methods under a general data-driven error mitigation framework that we call UNIfied Technique for Error mitigation with Data (UNITED). In certain situations, we find that our UNITED method can outperform the individual methods (i.e., the whole is better than the individual parts). Specifically, we employ a realistic noise model obtained from a trapped ion quantum computer to benchmark UNITED, as well as other state-of-the-art methods, in mitigating observables produced from random quantum circuits and the Quantum Alternating Operator Ansatz (QAOA) applied to Max-Cut problems with various numbers of qubits, circuit depths and total numbers of shots. We find that the performance of different techniques depends strongly on shot budgets, with more powerful methods requiring more shots for optimal performance. For our largest considered shot budget (10 10 ), we find that UNITED gives the most accurate mitigation. Hence, our work represents a benchmarking of current error mitigation methods and provides a guide for the regimes when certain methods are most useful.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Evaluation of Multi-Fidelity Soil Moisture Products Across the Continental United States

We have aggregated the most recent soil moisture datasets from a diverse range of sources, encompassing the Continental United States (CONUS). These sources encompass gridded data from remote sensing products, reanalysis products, machine learning-based projects, and land surface modeling products. Additionally, we have obtained and processed in-situ soil moisture observations from the International Soil Moisture Network. The collected datasets exhibit variations in both temporal and spatial resolutions. Among the 20 datasets, six are available at a spatial resolution of 0.25 degrees, while three are at a coarser spatial resolution of 25 km. To minimize spatial interpolation, we conducted data uncertainty evaluations at the 0.25-degree spatial resolution. For our data evaluations, we maintained a monthly temporal resolution, which effectively captures soil moisture seasonality and interannual variability. Our data processing strategy preserves the raw data and interpolated data at their original temporal resolutions. Datasets with higher temporal resolutions, including daily, three-hourly, and hourly datasets, are set aside for subsequent analyses. These analyses will delve into topics such as soil moisture changes and recovery during extreme weather events. Furthermore, we have processed auxiliary data to enhance our evaluation, leveraging tools such as Google Earth Engine. This includes incorporating topography data, land use land cover data, Köppen-Geiger climate classification, and more to provide a comprehensive assessment from multiple sources.

Li, Lingcheng↗

Estimating carrying capacity for juvenile salmon using quantile random forest models

Abstract Establishing robust methods and metrics to evaluate habitat quality is critical for the recovery of endangered Pacific salmonids ( Oncorhynchus spp.). A variety of modeling approaches are used for status and trend monitoring of anadromous species throughout the Pacific Northwest, USA, but current methods may fail to capture the complex relationship between fish and habitat and are often limited in predictive power beyond specific watersheds. Further, the focus on species distribution and abundance is not easily manipulated to predict carrying capacity and traditional stock‐recruitment analyses are reliant on long‐term data which are not always available. In this study, we developed a quantile random forest model to provide estimates of habitat carrying capacity for Chinook salmon ( O. tshawytscha ) parr during the summer months, at both the site and watershed scale. Quantile random forest models allow for the consideration of noisy data, correlated variables, and non‐linear relationships: common features in fish–habitat datasets. We leveraged Columbia Habitat Monitoring Program data to select habitat co‐variates and predict capacity at those sites. We also identified a set of globally available attributes to extrapolate capacity estimate predictions throughout wadeable streams within the Columbia River basin. Total capacity estimates for watersheds closely matched estimates from alternative fish productivity models. Carrying capacity estimates based on quantile random forest models, like those presented here, provide managers a framework to guide the identification, prioritization, and development of habitat rehabilitation actions to recover salmon populations.

See, Kevin E.↗

Using Cosmic Ray Muons to Assess Geological Characteristics in the Subsurface

Cosmic rays are energetic nuclei and elementary particles that originate from stars and intergalactic events. The interaction of these particles with the upper atmosphere produces a wide range of secondary particles that reach the surface of the earth, of which muons are the most prominent. With enough energy, muons can travel up to a few kilometers beneath the surface of the earth before being stopped completely. The terrestrial muon flux profile and associated zenith angle can be utilized to determine geological characteristics of a location (e.g., rock overburden and density) without having to use conventional methods such as boreholes. This work uses a low-power plastic scintillator-based muon detection system as a prototype for this non-destructive geological assay methodology. Four custom designed 102 cm x 51 cm x 5 cm plastic scintillation panels are used to realize two orthogonal detection planes. Optical photons from each scintillation panel are read using OnSemi J-Series 4x4 silicon photomultiplier (SiPM) arrays in conjunction with preamplifiers. Simultaneous triggers between detectors from two planes indicate a coincidence event which is recorded using the QuarkNet data acquisition system (DAQ) from Fermi National Accelerator Laboratory. A custom detector holder was designed to securely mount the detection system and rotate the panels along the zenith to collect data at variable angles. In order to quantify the systematic uncertainties associated with the detector, such as energy depositions and angular resolution of the detector design, a Monte Carlo (MC) simulation using Geant4 is being developed. Cosmic ray flux prediction will be included in the project by adding the CORSIKA MC code to the simulation toolchain. Simulated and experimental data will drive the development and validation of a reconstruction algorithm that, upon completion, is expected to predict average overburden and rock density. Extended detector exposure to muons can be used as a means to understand changes in the surrounding environment like rock porosity. On the experimental front, muons will initially be measured at the surface, establishing the baseline flux. This is followed by recording the muon flux at variable depths and zenith angles, where the data will be used by the reconstruction algorithm to predict the overburden. The result will be benchmarked against geological surveys. The measured flux data will also be used to benchmark independent and established models. Successful proof-of-concept demonstration of this technology can open doors for long term non-invasive geological monitoring. The detector design, experimental methodology, and the benchmarking efforts are detailed in this work.

Gadey, Harish Reddy↗