Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Comparison of Radiative Energy Flows in Observational Datasets and Climate Modeling

This study examines radiative flux distributions and local spread of values from three major observational datasets (CERES, ISCCP, and SRB) and compares them with results from climate modeling (CMIP3). Examinations of the spread and differences also differentiate among contributions from cloudy and clear-sky conditions. The spread among observational datasets is in large part caused by noncloud ancillary data. Average differences of at least 10Wm(exp -2) each for clear-sky downward solar, upward solar, and upward infrared fluxes at the surface demonstrate via spatial difference patterns major differences in assumptions for atmospheric aerosol, solar surface albedo and surface temperature, and/or emittance in observational datasets. At the top of the atmosphere (TOA), observational datasets are less influenced by the ancillary data errors than at the surface. Comparisons of spatial radiative flux distributions at the TOA between observations and climate modeling indicate large deficiencies in the strength and distribution of model-simulated cloud radiative effects. Differences are largest for lower-altitude clouds over low-latitude oceans. Global modeling simulates stronger cloud radiative effects (CRE) by +30Wmexp -2) over trade wind cumulus regions, yet smaller CRE by about -30Wm(exp -2) over (smaller in area) stratocumulus regions. At the surface, climate modeling simulates on average about 15Wm(exp -2) smaller radiative net flux imbalances, as if climate modeling underestimates latent heat release (and precipitation). Relative to observational datasets, simulated surface net fluxes are particularly lower over oceanic trade wind regions (where global modeling tends to overestimate the radiative impact of clouds). Still, with the uncertainty in noncloud ancillary data, observational data do not establish a reliable reference.

Raschke, Ehrhard↗

The Cumulus and Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD)

Low clouds continue to contribute greatly to the uncertainty in cloud feedback estimates. Depending on whether a region is dominated by cumulus (Cu) or stratocumulus (Sc) clouds, the interannual low-cloud feedback is somewhat different in both spaceborne and large-eddy simulation studies. Therefore, simulating the correct amount and variation of the Cu and Sc cloud distributions could be crucial to predict future cloud feedbacks. Here we document spatial distributions and profiles of Sc and Cu clouds derived from Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) and CloudSat measurements. For this purpose, we create a new dataset called the Cumulus And Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD), which identifies Sc, broken Sc, Cu under Sc, Cu with stratiform outflow and Cu. To separate the Cu from Sc, we design an original method based on the cloud height, horizontal extent, vertical variability and horizontal continuity, which is separately applied to both CALIPSO and combined CloudSat–CALIPSO observations. First, the choice of parameters used in the discrimination algorithm is investigated and validated in selected Cu, Sc and Sc–Cu transition case studies. Then, the global statistics are compared against those from existing passive- and active-sensor satellite observations. Our results indicate that the cloud optical thickness – as used in passive-sensor observations – is not a sufficient parameter to discriminate Cu from Sc clouds, in agreement with previous literature. Using clustering-derived datasets shows better results although one cannot completely separate cloud types with such an approach. On the contrary, classifying Cu and Sc clouds and the transition between them based on their geometrical shape and spatial heterogeneity leads to spatial distributions consistent with prior knowledge of these clouds, from ground-based, ship-based and field campaigns. Furthermore, we show that our method improves existing Sc–Cu classifications by using additional information on cloud height and vertical cloud fraction variation. Finally, the CASCCAD datasets provide a basis to evaluate shallow convection and stratocumulus clouds on a global scale in climate models and potentially improve our understanding of low-level cloud feedbacks. The CASCCAD dataset (Cesana, 2019, https://doi.org/10.5281/zenodo.2667637) is available on the Goddard Institute for Space Studies (GISS) website at https://data.giss.nasa.gov/clouds/casccad/ (last access: 5 November 2019) and on the zenodo website at https://zenodo.org/record/2667637 (last access: 5 November 2019).

Cesana, Gregory V.↗

Version 2 Ozone Monitoring Instrument SO2 Product (OMSO2 V2): New Anthropogenic SO2 Vertical Column Density Dataset

The Ozone Monitoring Instrument (OMI) has been providing global observations of SO2 pollution since 2004. Here we introduce the new anthropogenic SO2 vertical column density (VCD) dataset in the version 2 OMI SO2 product (OMSO2 V2). As with the previous version (OMSO2 V1.3), the new dataset is generated with an algorithm based on principal component analysis of OMI radiances, but features several updates. The most important among those is the use of expanded lookup tables and model a priori profiles to estimate SO2 Jacobians for individual OMI pixels, in order to better characterize pixel-to-pixel variations in SO2 sensitivity, including over snow and ice. Additionally, new data screening and spectral fitting schemes have been implemented to improve the quality of the spectral fit. As compared with the planetary boundary layer SO2 dataset in OMSO2 V1.3, the new dataset has substantially better data quality, especially over areas that are relatively clean or affected by the south Atlantic anomaly. The updated retrievals over snow/ice yield more realistic seasonal changes in SO2 at high latitudes and offer enhanced sensitivity to sources during wintertime. An error analysis has been conducted to assess uncertainties in SO2 VCDs from both the spectral fit and Jacobian calculations. The uncertainties from spectral fitting are reflected in SO2 slant column densities (SCDs) and largely depend on the signal-to-noise ratio of the measured radiances, as implied by the generally smaller SCD uncertainties over clouds or for smaller solar zenith angles. The SCD uncertainties for individual pixels are estimated to be~0.15-0.3 DU (Dobson Units) between ~40°S and ~40°N and to be~0.2-0.5 DU at higher latitudes. The uncertainties from the Jacobians are approximately ~50-100% over polluted areas, and primarily attributed to errors in SO2 a priori profiles and cloud pressures, as well as the lack of explicit treatment for aerosols. Finally, the daily mean and median SCDs over the presumably SO2-free equatorial East Pacific have increased by only~0.0035 DU and ~0.003 DU respectively over the entire 15-year OMI record; while the standard deviation of SCDs has grown by only~0.02 DU or ~10%. Such remarkable long-term stability makes the new dataset particularly suitable for detecting regional changes in SO2 pollution.

OMI, SO2, Remote Sensing↗

PERSIANN Dynamic Infrared–Rain Rate (PDIR-Now): A Near-Real-Time, Quasi-Global Satellite Precipitation Dataset

This study presents the Precipitation Estimation from Remotely Sensed Information Using Artificial Neural Networks–Dynamic Infrared Rain Rate (PDIR-Now) near-real-time precipitation dataset. This dataset provides hourly, quasi-global, infrared-based precipitation estimates at 0.04° × 0.04° spatial resolution with a short latency (15–60 min). It is intended to supersede the PERSIANN–Cloud Classification System (PERSIANN-CCS) dataset previously produced as the near-real-time product of the PERSIANN family. We first provide a brief description of the algorithm’s fundamentals and the input data used for deriving precipitation estimates. Second, we provide an extensive evaluation of the PDIR-Now dataset over annual, monthly, daily, and subdaily scales. Last, the article presents information on the dissemination of the dataset through the Center for Hydrometeorology and Remote Sensing (CHRS) web-based interfaces. The evaluation, conducted over the period 2017–18, demonstrates the utility of PDIR-Now and its improvement over PERSIANN-CCS at all temporal scales. Specifically, PDIR-Now improves the estimation of rain/no-rain days as demonstrated by a critical success index (CSI) of 0.53 compared to 0.47 of PERSIANN-CCS. In addition, PDIR-Now improves the estimation of seasonal and diurnal cycles of precipitation as well as regional precipitation patterns erroneously estimated by PERSIANN-CCS. Finally, an evaluation is carried out to examine the performance of PDIR-Now in capturing two extreme events, Hurricane Harvey and a cluster of summer thunderstorms that occurred over the Netherlands, where it is shown that PDIR-Now adequately represents spatial precipitation patterns as well as subdaily precipitation rates with a correlation coefficient (CORR) of 0.64 for Hurricane Harvey and 0.76 for the Netherlands thunderstorms.

Rainfall↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

AssistTaxi: A Comprehensive Dataset for Taxiway Analysis and Autonomous Operations

The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems. This poster presents AssistTaxi, which is a comprehensive novel dataset which is a collection of images for runway and taxiway analysis. The dataset comprises of more than 300,000 frames of diverse and carefully collected data, gathered from Melbourne (MLB) and Grant-Valkaria (X59) general aviation airports. The importance of AssistTaxi lies in its potential to advance autonomous operations, enabling researchers and developers to train and evaluate algorithms for efficient and safe taxiing. Researchers can utilize AssistTaxi to benchmark their algorithms, assess performance, and explore novel approaches for runway and taxiway analysis. Additionally, the dataset serves as a valuable resource for validating and enhancing existing algorithms as well as facilitating innovation in autonomous operations for aviation. We also propose an initial approach to label the dataset using a contour based detection and line extraction technique.

Data Collection↗

Power System Waveform Datasets for Machine Learning

The desire for increased visibility across the electricity grid will necessarily increase the deployment of sensing and measurement devices and associated data management needs to unprecedented levels. For the existing sensing and measurement infrastructure, there remains a great amount of “value” yet to be extracted through advanced data management and analytics. Availability of more data will not, by itself, lead to changes in grid visibility, security, and resiliency. To create the predictive and prescriptive environment required to enable new markets and transactions for customer revenue and a reliable grid, the data must be collected, organized, evaluated, and analyzed using sophisticated algorithms to provide actionable information allowing operators and customers to reliably manage an increasingly complex grid. Progress in artificial intelligence (AI) has been largely driven by large, publicly available datasets that can be used to train AI algorithms such as MNIST, a database of handwritten images of digits, and ImageNet, an image database of everyday objects. These types of publicly available databases of real-world training datasets have been largely credited for advancement of image processing, computer vision, and deep learning algorithms that these use cases deploy. However, in the power systems industry to date, there are few databases with proper event labeling, and data access to a publicly available collection of power system event waveforms that will allow users to interact with grid signature data. Publicly available datasets of power system event waveforms, such as the DOE/EPRI dataset, often lack critical metadata or contain limited examples of each event type, and data formats vary widely across these datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source↗

Data-driven key performance indicators and datasets for building energy flexibility: A review and perspectives

Energy flexibility, through short-term demand-side management (DSM) and energy storage technologies, is now seen as a major key to balancing the fluctuating supply in different energy grids with the energy demand of buildings. This is especially important when considering the intermittent nature of ever-growing renewable energy production, as well as the increasing dynamics of electricity demand in buildings. This paper provides a holistic review of (1) data-driven energy flexibility key performance indicators (KPIs) for buildings in the operational phase and (2) open datasets that can be used for testing energy flexibility KPIs. The review identifies a total of 48 data-driven energy flexibility KPIs from 87 recent and relevant publications. These KPIs were categorized and analyzed according to their type, complexity, scope, key stakeholders, data requirement, baseline requirement, resolution, and popularity. Moreover, 330 building datasets were collected and evaluated. Of those, 16 were deemed adequate to feature building performing demand response or building-to-grid (B2G) services. The DSM strategy, building scope, grid type, control strategy, needed data features, and usability of these selected 16 datasets were analyzed. This review reveals future opportunities to address limitations in the existing literature: (1) developing new data-driven methodologies to specifically evaluate different energy flexibility strategies and B2G services of existing buildings; (2) developing baseline-free KPIs that could be calculated from easily accessible building sensors and meter data; (3) devoting non-engineering efforts to promote building energy flexibility, standardizing data-driven energy flexibility quantification and verification processes; and (4) curating and analyzing datasets with proper description for energy flexibility assessm.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Plastic additives in the ocean: Use of a comprehensive dataset for meta-analysis and method development

In excess of 13,000 chemicals are added to plastics (‘additives’) to improve performance, durability, and production of plastic products. They are categorized into numerous chemical classes including flame retardants, light stabilizers, antioxidants, and plasticizers. While research on plastic additives in the marine environment has increased over the past decade, there is a lack of methodological standardization. To direct future measurement of plastic additives, we compiled a first-of-its-kind dataset of literature assessing plastic additives in marine environments, delineated by sample type (plastic debris, seawater, sediment, biota). Using this dataset, we performed a meta-analysis to summarize the state of the science. Currently, our dataset includes 217 publications published between 1978 and May 2023. The majority of publications analyzed plastic additives in biota collected from Europe and Asia. Analyses concentrated on plasticizers, brominated flame retardants, and bisphenols. Common sample preparation techniques included Solvent - Agitation extraction for plastic, sediment, and biota samples, and Solid Phase Extraction for seawater samples with dichloromethane and solvent mixtures including dichloromethane as the organic extraction solvent. Finally, most analyses were performed utilizing gas chromatography/mass spectrometry. There are a variety of data gaps illuminated by this meta-analysis, most notably the small number of compounds that have been targeted for detection compared to the large number of additives used in plastic production. The provided dataset facilitates future investigation of trends in plastic additive concentration data in the marine environment (allowing for comparison to toxicity thresholds) and acts as a starting point for optimizing and harmonizing plastic additive analytical methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Case Study: NREL Campus Chilled Water Storage Potential: Benchmark Datasets Development and Applications, Task 4 - Use Case Demonstration

The Benchmark Datasets Development and Applications project is a three-year collaboration between the National Renewable Energy Laboratory (NREL), Oak Ridge National Laboratory, Pacific Northwest National Laboratory, and Lawrence Berkeley National Laboratory. The project seeks to collect and curate high-resolution, well-calibrated time series of building operational and indoor/outdoor environmental data, which are crucial to understanding and optimizing building energy efficiency performance and demand flexibility capabilities as well as benchmarking energy algorithms. Project outcomes include approximately twelve high-fidelity building datasets, enhanced data representation tools, and four case studies to illustrate example applications. The goal of these case studies is to define and execute analyses that demonstrate how one or more datasets collected through this project can address a data gap or challenge historically faced by building stakeholders. This technical paper summarizes the findings of one of these case studies, in which we studied the operational efficiencies of the central cooling system at NREL. We looked at three years of data from the three chillers in the Field Test Laboratory Building (FTLB), from 2019 to 2021, to compare equipment operation and demand throughout the time period. Our analysis indicates that all three chillers are operating at or below the optimal loading conditions for most of the operation time, and thus there was no efficiency drop due to loading of the chillers at full capacity. Our recommendation is that no chiller capacity increase is needed; instead, the central plant could benefit from adopting advanced control logics for optimal sequencing of chillers during part load operations. Analysis of adding chilled water thermal storage to the central plant indicated 34% savings in demand cost and 24.5% savings in total cost (energy consumption and demand charge cost). The payback period is estimated to be 11-22 years with an assumed TES cost of $\$$100-$200 per ton. This case study shows how a selected dataset is used to solve a practical building problem - learning the operational status of its components, analyzing the effectiveness of a proposed new technique, and aiding decision-making for the building operations and maintenance team.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Preventing Failures By Dataset Shift Detection in Safety-Critical Graph Applications

Dataset shift refers to the problem where the input data distribution may change over time (e.g., between training and test stages). Since this can be a critical bottleneck in several safety-critical applications such as healthcare, drug-discovery, etc., dataset shift detection has become an important research issue in machine learning. Though several existing efforts have focused on image/video data, applications with graph-structured data have not received sufficient attention. Therefore, in this paper, we investigate the problem of detecting shifts in graph structured data through the lens of statistical hypothesis testing. Specifically, we propose a practical two-sample test based approach for shift detection in large-scale graph structured data. Our approach is very flexible in that it is suitable for both undirected and directed graphs, and eliminates the need for equal sample sizes. Using empirical studies, we demonstrate the effectiveness of the proposed test in detecting dataset shifts. We also corroborate these findings using real-world datasets, characterized by directed graphs and a large number of nodes.

97 MATHEMATICS AND COMPUTING↗

Predicting cutoff L-shells of solar protons using the GPPSn particle dataset

Solar energetic protons (SEPs) arriving at the Earth trigger severe radiation storms in the near-Earth space, directly impacting space missions operating at various altitudes. Therefore, monitoring SEP events and predicting the penetration depths of solar protons are critical for aerospace sectors. Building on previous efforts, here we demonstrate the feasibility of using proton measurements from the Global Prompt Proton Sensor network (GPPSn), enabled by Los Alamos National Laboratory developed combined X-ray dosimeters aboard GPS satellites, to characterize and predict the penetration of solar protons into the geomagnetic field. The inclined medium-Earth-orbits (MEOs) of the global GPS constellation offer a unique advantage of allowing simultaneous measurements of penetrating solar protons inside both open- and closed-field line regions. Therefore, the L-profiles of ∼10s–100 MeV solar protons and their associated cutoff L-shells can be determined from the GPPSn dataset, using predefined threshold proton flux values rather than traditional flux ratios. After examining a list of SEP event intervals across solar cycles 23, 24 and 25—including the 2024 Mother’s Day superstorm, we showcase how the latest GPPSn proton dataset (release v1.10), reprocessed and calibrated, can not only be used to monitor solar proton distributions inside the dynamic geomagnetic field for individual events, but also to derive a new empirical model linking cutoff L-shells with several key space weather parameters. This newly developed SEPCL-MEO model demonstrates high predictive performance; for example, predictions for > 30 MeV solar protons yield a correlation coefficient of 0.85 and performance efficiency of 0.67 when validated against GPPSn observations. Results from this pilot study underscores the scientific and operational value of the GPPSn dataset, and this dataset—when paired with machine-learning techniques—can play a critical role in observing and predicting the effects of future incoming SEP events, including extreme ones.

58 GEOSCIENCES↗

A Large Dataset of Fluvial Hydraulic and Geometry Attributes Derived From USGS Field Measurement Records

Accurate representation of river channel geometry is important for hydrologic and hydraulic modeling of fluvial systems. Often, channel geometry is estimated using simple rating curves that can be applied across various spatial scales. However, such methods are limited to power law relations that do not employ many potentially relevant catchment and river attributes. This paper introduce a new dataset, IFMHA (Inventory of Field Measurement of Hydraulic Attributes), to enable research studies on channel geometry and streamflow characteristics. IFMHA is derived from the National Water Information System (NWIS) site inventory for surface water field measurements and stream attributes from the National Hydrography Dataset (NHD). IFMHA includes 2,802,532 records from 10,050 sites (NWIS streamgaging stations). The dataset utility is demonstrated here by presenting a series of conceptual models for estimating channel geometry parameters (i.e., channel mean depth, channel maximum depth, wetted perimeter, and roughness) based on the available field attributes within IFMHA. Such a dataset and attributed channel geometry parameters can enhance the performance of operational flood forecasting frameworks (e.g. National Water Model) by providing more accurate initial conditions used in hydrologic and hydraulic routing models.

Hydrology↗

The Kimberlina synthetic multiphysics dataset for CO 2 monitoring investigations

Abstract We present a synthetic multi‐scale, multi‐physics dataset constructed from the Kimberlina 1.2 CO 2 reservoir model based on a potential CO 2 storage site in the Southern San Joaquin Basin of California. Among 300 models, one selected reservoir‐simulation scenario produces hydrologic‐state models at the onset and after 20 years of CO 2 injection. Subsequently, these models were transformed into geophysical properties, including P‐ and S‐wave seismic velocities, saturated density where the saturating fluid can be a combination of brine and supercritical CO 2 , and electrical resistivity using established empirical petrophysical relationships. From these 3D distributions of geophysical properties, we have generated synthetic time‐lapse seismic, gravity and electromagnetic responses with acquisition geometries that mimic realistic monitoring surveys and are achievable in actual field situations. We have also created a series of synthetic well logs of CO 2 saturation, acoustic velocity, density and induction resistivity in the injection well and three monitoring wells. These were constructed by combining the low‐frequency trend of the geophysical models with the high‐frequency variations of actual well logs collected at the potential storage site. In addition, to better calibrate our datasets, measurements of permeability and pore connectivity have been made on cores of Vedder Sandstone, which forms the primary reservoir unit. These measurements provide the range of scales in the otherwise synthetic dataset to be as close to a real‐world situation as possible. This dataset consisting of the reservoir models, geophysical models, simulated time‐lapse geophysical responses and well logs forms a multi‐scale, multi‐physics testbed for designing and testing geophysical CO 2 monitoring systems as well as for imaging and characterization algorithms. The suite of numerical models and data have been made publicly available for downloading on the National Energy Technology Laboratory's (NETL) Energy Data Exchange (EDX) website.

58 GEOSCIENCES↗

Resonant anomaly detection with multiple reference datasets

An important class of techniques for resonant anomaly detection in high energy physics builds models that can distinguish between reference and target datasets, where only the latter has appreciable signal. Such techniques, including Classification Without Labels (CWoLa) and Simulation Assisted Likelihood-free Anomaly Detection (SALAD) rely on a single reference dataset. They cannot take advantage of commonly available multiple datasets and thus cannot fully exploit available information. In this work, we propose generalizations of CWoLa and SALAD for settings where multiple reference datasets are available, building on weak supervision techniques. We demonstrate improved performance in a number of settings with realistic and synthetic data. As an added benefit, our generalizations enable us to provide finite-sample guarantees, improving on existing asymptotic analyses.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Estimation of soil classes and their relationship to grapevine vigor in a Bordeaux vineyard: advancing the practical joint use of electromagnetic induction (EMI) and NDVI datasets for precision viticulture

Working within a vineyard in the Pessac Léognan Appellation of Bordeaux, France, this study documents the potential of using simple statistical methods with spatially-resolved and increasingly available electromagnetic induction (EMI) geophysical and normalized difference vegetation index (NDVI) datasets to accurately estimate Bordeaux vineyard soil classes and to quantitatively explore the relationship between vineyard soil types and grapevine vigor. First, co-located electrical tomographic tomography (ERT) and EMI datasets were compared to gain confidence about how the EMI method averaged soil properties over the grapevine rooting depth. Then, EMI data were used with core soil texture and soil-pit based interpretations of Bordeaux soil types (Brunisol, Redoxisol, Colluviosol and Calcosol) to estimate the spatial distribution of geophysically-identified Bordeaux soil classes. A strong relationship (r = 0.75, p < 0.01) was revealed between the geophysically-identified Bordeaux soil classes and NDVI (both 2 m resolution), showing that the highest grapevine vigor was associated with the Bordeaux soil classes having the largest clay fraction. The results suggest that within-block variability of grapevine vigor was largely controlled by variability in soil classes, and that carefully collected EMI and NDVI datasets can be exceedingly helpful for providing quantitative estimates of vineyard soil and vigor variability, as well as their covariation. The method is expected to be transferable to other viticultural regions, providing an approach to use easy-to-acquire, high resolution datasets to guide viticultural practices, including routine management and replanting.

54 ENVIRONMENTAL SCIENCES↗

Estimating geographic origins of corn and soybean biomass for biofuel production: A detailed dataset

Sustainable fuel initiatives in the United States such as the Environmental Protection Agency’s Renewable Fuel Stan- dard and the Department of Energy’s Sustainable Aviation Fuel Grand Challenge have increased the production of corn ethanol and soybean biodiesel. However, the lack of precise information regarding biomass sourcing at a localized level has hindered accurate understanding of both biofuel costs and environmental impact of these production pathways. By harnessing the power of geospatial analysis and leveraging United States Department of Agriculture (USDA) crop cen- sus data, this dataset fills this critical knowledge gap. This dataset offers a novel estimation of geospatial biomass sourc- ing for biofuel production in the United States by synthe- sizing 2017 USDA crop census data, biorefinery data from the United States Energy Information Administration, and publicly available information about biomass sourcing for biofuel production. This dataset provides a detailed under- standing of biomass use for first generation biofuel pro- duction, enabling stakeholders to make informed decisions about resource allocation, investment strategies, and infras- tructure development. Furthermore, the county-level gran- ularity of the dataset allows for increased fidelity in the techno-economic assessments and life-cycle analyses of first- generation biofuels in the United States.

09 BIOMASS FUELS↗