Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE

This dataset is an open-source repository of spectral data measured at Pacific Northwest National Laboratory (PNNL). This database provides quantitative values for the complex index of refraction for seven polycyclic aromatic hydrocarbon (PAH) solids. A list of the chemicals is available in the readme file. These spectra consist of the optical constants, i.e., the real, n(ν), and imaginary, k(ν), refractive indices, over the spectral range from 7,800 to 400 cm-1 (1.28 – 25 μm). The conditions under which the individual data were acquired are described in the associated metadata files, and the user is strongly encouraged to read and understand this information to ensure the data are used appropriately for your application. Recommended Citation for Dataset Jessica M Salcido, Jeremy D. Erickson, Ashley M. Bradley, Russell G. Tonkyn, Timothy J. Johnson and Tanya L. Myers. 2026. PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE. [Data Set] PNNL DataHub. INSERT DOI License Information This work is marked with CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/. The authors do request that you appropriately cite the dataset when referencing or using the dataset.

Salcido, Jessica Marie Ortola↗

Evaluating the Benefits of Bayesian Hierarchical Methods for Analyzing Heterogeneous Environmental Datasets: A Case Study of Marine Organic Carbon Fluxes

Large compilations of heterogeneous environmental observations are increasingly available as public databases, allowing researchers to test hypotheses across datasets. Statistical complexities arise when analyzing compiled data due to unbalanced spatial sampling, variable environmental context, mixed measurement techniques, and other reasons. Hierarchical Bayesian modeling is increasingly used in environmental science to describe these complexities, however few studies explicitly compare the utility of hierarchical Bayesian models to simpler and more commonly applied methods. Here we demonstrate the utility of the hierarchical Bayesian approach with application to a large compiled environmental dataset consisting of 5,741 marine vertical organic carbon flux observations from 407 sampling locations spanning eight biomes across the global ocean. We fit a global scale Bayesian hierarchical model that describes the vertical profile of organic carbon flux with depth. Profile parameters within a particular biome are assumed to share a common deviation from the global mean profile. Individual station-level parameters are then modeled as deviations from the common biome-level profile. The hierarchical approach is shown to have several benefits over simpler and more common data aggregation methods. First, the hierarchical approach avoids statistical complexities introduced due to unbalanced sampling and allows for flexible incorporation of spatial heterogeneitites in model parameters. Second, the hierarchical approach uses the whole dataset simultaneously to fit the model parameters which shares information across datasets and reduces the uncertainty up to 95% in individual profiles. Third, the Bayesian approach incorporates prior scientific information about model parameters; for example, the non-negativity of chemical concentrations or mass-balance, which we apply here. We explicitly quantify each of these properties in turn. We emphasize the generality of the hierarchical Bayesian approach for diverse environmental applications and its increasing feasibility for large datasets due to recent developments in Markov Chain Monte Carlo algorithms and easy-to-use high-level software implementations.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty in Atlantic Multidecadal Oscillation derived from different observed datasets and their possible causes

As a leading mode of sea surface temperature (SST) variability over the North Atlantic in both observations and model simulations, the Atlantic Multidecadal Oscillation (AMO) can have a substantial influence on regional and global climate. By using Low-Frequency Component Analysis, we explore the uncertainties of the resulting AMO indices and the corresponding spatial patterns derived from three observational SST datasets. We found that the known coherent spatial pattern of the AMO at the basin scale over the North Atlantic appears in two out of the three datasets. Further analysis indicates that both the warming trend and the different techniques used to construct these observed gridded SSTs contribute to the AMO’s spatial coherence over the North Atlantic, especially during periods of sparse data sampling. The SST in the Extended Reconstructed SST dataset version 5 (ERSSTv5), changes from being systematically below the other datasets during the dense sampling periods on either side of the Second World War (WWII), to systematically above the other datasets during WWII, thereby introducing an artificial 10–20-year variability that affects the AMO’s spatial coherence. This coherence in the AMO’s spatial pattern is also affected by bias adjustment in ERSSTv5 at relative cool (i.e., non-summer) seasons, and by the heterogeneous North Atlantic warming pattern. The different AMO patterns can induce the different effects of wind, surface heat fluxes, and then drive ocean circulation and its heat transport convergence, especially for some seasons. For AMO indices, both the different detrending methods and different observational data result in uncertainty for the period 1935–1950. Such SST uncertainty is important to detect the relative role of the atmosphere and ocean in shaping the AMO.

54 ENVIRONMENTAL SCIENCES↗

BAWLD-CH 4 : a comprehensive dataset of methane fluxes from boreal and arctic ecosystems

Methane (CH 4 ) emissions from the boreal and arctic region are globally significant and highly sensitive to climate change. There is currently a wide range in estimates of high-latitude annual CH 4 fluxes, where estimates based on land cover inventories and empirical CH 4 flux data or process models (bottom-up approaches) generally are greater than atmospheric inversions (top-down approaches). A limitation of bottom-up approaches has been the lack of harmonization between inventories of site-level CH 4 flux data and the land cover classes present in high-latitude spatial datasets. Here we present a comprehensive dataset of small-scale, surface CH 4 flux data from 540 terrestrial sites (wetland and non-wetland) and 1247 aquatic sites (lakes and ponds), compiled from 189 studies. The Boreal–Arctic Wetland and Lake Methane Dataset (BAWLD-CH 4 ) was constructed in parallel with a compatible land cover dataset, sharing the same land cover classes to enable refined bottom-up assessments. BAWLD-CH 4 includes information on site-level CH 4 fluxes but also on study design (measurement method, timing, and frequency) and site characteristics (vegetation, climate, hydrology, soil, and sediment types, permafrost conditions, lake size and depth, and our determination of land cover class). The different land cover classes had distinct CH 4 fluxes, resulting from definitions that were either based on or co-varied with key environmental controls. Fluxes of CH 4 from terrestrial ecosystems were primarily influenced by water table position, soil temperature, and vegetation composition, while CH 4 fluxes from aquatic ecosystems were primarily influenced by water temperature, lake size, and lake genesis. Models could explain more of the between-site variability in CH 4 fluxes for terrestrial than aquatic ecosystems, likely due to both less precise assessments of lake CH 4 fluxes and fewer consistently reported lake site characteristics. Analysis of BAWLD-CH 4 identified both land cover classes and regions within the boreal and arctic domain, where future studies should be focused, alongside methodological approaches. Overall, BAWLD-CH 4 provides a comprehensive dataset of CH 4 emissions from high-latitude ecosystems that are useful for identifying research opportunities, for comparison against new field data, and model parameterization or validation.

54 ENVIRONMENTAL SCIENCES↗

Discrete global grid system-based flow routing datasets in the Amazon and Yukon basins

Abstract. Discrete global grid systems (DGGS) are emerging spatial data structures widely used to organize geospatial datasets across scales. While DGGS have found applications in various scientific disciplines, including atmospheric science and ecology, their integration into physically based hydrological models and Earth system models (ESMs) has been hindered by the lack of flow routing datasets based on DGGS. In response to this gap, this study pioneers the development of new flow routing datasets using icosahedral Snyder equal-area (ISEA) DGGS and a novel mesh-independent flow direction model. We present flow routing datasets for two large basins, the tropical Amazon River basin and the Arctic Yukon River basin. These datasets (1) facilitate the adoption of DGGS for hydrological models and (2) provide flow routing inputs for evaluation of DGGS-based flow routing in the Amazon and Yukon river basins. The data are available at https://doi.org/10.5281/zenodo.8377765 (Liao, 2023).

54 ENVIRONMENTAL SCIENCES↗

U-Surf: a global 1 km spatially continuous urban surface property dataset for kilometer-scale urban-resolving Earth system modeling

High-resolution urban climate modeling has faced substantial challenges due to the absence of a globally consistent, spatially continuous, and accurate dataset to represent the spatial heterogeneity of urban surfaces and their biophysical properties. This deficiency has long obstructed the development of urban-resolving Earth system models (ESMs) and ultra-high-resolution urban climate modeling, over large domains. Here, we present U-Surf, a first-of-its-kind 1 km resolution present-day (circa 2020) global continuous urban surface parameter dataset. Using the urban canopy model (UCM) in the Community Earth System Model as a base model for satisfying dataset requirements, U-Surf leverages the latest advances in remote sensing, machine learning, and cloud computing to provide the most relevant urban surface biophysical parameters, including radiative, morphological, and thermal properties, for UCMs at the facet and canopy level. Generated using a systematically unified workflow, U-Surf ensures internal consistency among key parameters, making it the first globally coherent urban canopy surface dataset. U-Surf significantly improves the representation of the urban land heterogeneity both within and across cities globally; provides essential, high-fidelity surface biophysical constraints to urban-resolving ESMs; enables detailed city-to-city comparisons across the globe; and supports next-generation kilometer-resolution Earth system modeling across scales. U-Surf parameters can be easily converted or adapted to various types of UCMs, such as those embedded in weather and regional climate models, as well as air quality models. The fundamental urban surface constraints provided by U-Surf can also be used as features for machine learning models and can have other broad-scale applications for socioeconomic, public health, and urban planning contexts. We expect U-Surf to advance the research frontier of urban system science, climate-sensitive urban design, and coupled human–Earth systems in the future. The dataset is publicly available at https://doi.org/10.5281/zenodo.11247598 (Cheng et al., 2024).

Cheng, Yifan [Univ. of Illinois at Urbana-Champaig↗

A 1 km soil moisture dataset over eastern CONUS generated by assimilating SMAP data into the Noah-MP land surface model

An improved fine-scale soil moisture (SM) dataset at 1 km grid spacing, covering much of the eastern continental US, was generated by assimilating 9 km Soil Moisture Active Passive (SMAP) SM data into the v4.0.1 Noah-MP land surface model. With 12 ensemble members, the assimilation was carried out using the ensemble Kalman filter algorithm within NASA's Land Information System. The SM analysis for 2016 was fully validated against in situ observations from four different networks and compared with four other existing datasets. Results indicate that this SM analysis surpasses other datasets in top-layer SM distribution, including a machine-learning-based product, despite all SM estimates being less heterogeneous than observed. The analysis of anomalous errors suggests that large similarity in intrinsic errors is likely due to overlapping data sources among the selected SM datasets. More detailed evaluations were performed over two geographic areas. The observations collected by the Atmospheric Radiation Measurement facility in Oklahoma suggest that soil temperature and surface heat fluxes are concurrently simulated with good accuracy. Investigation into the 2016 southeastern US drought response further indicates drier conditions and higher evapotranspiration estimates compared to GLEAMv4.1. Notably, large errors are associated with grids having clay soil textures, underscoring the need for refined model treatments for specific soil types to further improve SM estimates. The dataset is publicly available on Zenodo at https://doi.org/10.5281/zenodo.14370563 (Tai et al., 2024).

Tai, Sheng-Lun [Pacific Northwest National Laborat↗

Call to Action for Global Access to and Harmonization of Quality Information of Individual Earth Science Datasets

Knowledge about the quality of data and metadata is important to support informed decisions on the (re)use of individual datasets and is an essential part of the ecosystem that supports open science. Quality assessments reflect the reliability and usability of data. They need to be consistently curated, fully traceable, and adequately documented, as these are crucial for sound decision- and policy-making efforts that rely on data. Quality assessments also need to be consistently represented and readily integrated across systems and tools to allow for improved sharing of information on quality at the dataset level for individual quality attribute or dimension. Although the need for assessing the quality of data and associated information is well recognized, methodologies for an evaluation framework and presentation of resultant quality information to end users may not have been comprehensively addressed within and across disciplines. Global interdisciplinary domain experts have come together to systematically explore needs, challenges and impacts of consistently curating and representing quality information through the entire lifecycle of a dataset. This paper describes the findings of that effort, argues the importance of sharing dataset quality information, calls for community action to develop practical guidelines, and outlines community recommendations for developing such guidelines. Practical guidelines will allow for global access to and harmonization of quality information at the level of individual Earth science datasets, which in turn will support open science.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Legacy Survey of Space and Time Data Preview 1: calibrations dataset type

The Legacy Survey of Space and Time Data Preview 1 (DP1) is the first release of data from the NSF-DOE Vera C. Rubin Observatory. It consists of raw and calibrated single-epoch images, co-adds, difference images, detection catalogs, and other derived data products. DP1 is based on 1792 science-grade optical/near-infrared exposures acquired over 48 distinct nights by the Rubin Commissioning Camera, LSSTComCam, on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile during the first on-sky commissioning campaign in late 2024. DP1 covers a total of approximately 15 sq. deg. over seven roughly equally-sized non-contiguous fields, each independently observed in six broad photometric bands, ugrizy, spanning a range of stellar densities and latitudes and overlapping with external reference datasets. This dataset is a subset of the full data release consisting of the calibrations dataset type. These are a collection of calibration datasets such as biases, darks, and flats used to construct the data release. This release contains 496 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS↗

Reference Site Condition Datasets for Floating Wind Arrays in the United States

Floating offshore wind farm design is highly site-specific, requiring detailed information about the specific conditions of a project area for realistic design studies. Unfortunately, publicly available site condition data for potential floating offshore wind project sites in the United States is scarce. To support U.S. offshore wind research, we developed reference site condition datasets, including metocean and seabed information, for four potential floating wind project areas in the U.S.: Humboldt Bay, Morro Bay, the Gulf of Maine, and the Gulf of Mexico. These datasets were compiled using publicly available data. Our metocean analysis, covering wind, waves, and surface currents, utilized measurement data from 2000 to 2020. Sources included the National Renewable Energy Laboratory’s National Offshore Wind Dataset for wind data, National Data Buoy Center buoys for wave data, and the High Frequency Radar Network for surface currents. These data were integrated into hourly time series used to compute extreme return periods up to 500 years, monthly statistics, and joint probability clusters for fatigue analysis. Soil conditions were evaluated using the usSEABED database and bathymetry grids were interpolated from the NCEI Digital Elevation Model Global Mosaic. Further information on the datasets and how they were created can be found in: Biglu, M., M. Hall, E. Lozon, S. Housner. 2024. Reference Site Conditions for Floating Wind Arrays in the United States. Golden, CO: National Renewable Energy Laboratory (NREL). NREL/TP-5000-89897. The data are also available at: https://github.com/FloatingArrayDesign/SiteConditions The content of each dataset is as follows: _NOW23_wind.txt: Hourly NOW-23 wind data up to a height of 400 meter. _metocean_1hr.txt: Hourly time series including wind, wave, surface current and temperature data. _Summary.xlsx: Metocean data, including extreme values, joint probability distributions and monthly statistics. _usSEABED_soil.csv: Extract of the usSEABED database for this specific site. _bathymetry_200m.txt (and 500m, 1000m): Gridded seabed depth data.

16 TIDAL AND WAVE POWER↗

Power System Waveform Datasets for Machine Learning

The desire for increased visibility across the electricity grid will necessarily increase the deployment of sensing and measurement devices and associated data management needs to unprecedented levels. For the existing sensing and measurement infrastructure, there remains a great amount of “value” yet to be extracted through advanced data management and analytics. Availability of more data will not, by itself, lead to changes in grid visibility, security, and resiliency. To create the predictive and prescriptive environment required to enable new markets and transactions for customer revenue and a reliable grid, the data must be collected, organized, evaluated, and analyzed using sophisticated algorithms to provide actionable information allowing operators and customers to reliably manage an increasingly complex grid. Progress in artificial intelligence (AI) has been largely driven by large, publicly available datasets that can be used to train AI algorithms such as MNIST, a database of handwritten images of digits, and ImageNet, an image database of everyday objects. These types of publicly available databases of real-world training datasets have been largely credited for advancement of image processing, computer vision, and deep learning algorithms that these use cases deploy. However, in the power systems industry to date, there are few databases with proper event labeling, and data access to a publicly available collection of power system event waveforms that will allow users to interact with grid signature data. Publicly available datasets of power system event waveforms, such as the DOE/EPRI dataset, often lack critical metadata or contain limited examples of each event type, and data formats vary widely across these datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Data-driven key performance indicators and datasets for building energy flexibility: A review and perspectives

Energy flexibility, through short-term demand-side management (DSM) and energy storage technologies, is now seen as a major key to balancing the fluctuating supply in different energy grids with the energy demand of buildings. This is especially important when considering the intermittent nature of ever-growing renewable energy production, as well as the increasing dynamics of electricity demand in buildings. This paper provides a holistic review of (1) data-driven energy flexibility key performance indicators (KPIs) for buildings in the operational phase and (2) open datasets that can be used for testing energy flexibility KPIs. The review identifies a total of 48 data-driven energy flexibility KPIs from 87 recent and relevant publications. These KPIs were categorized and analyzed according to their type, complexity, scope, key stakeholders, data requirement, baseline requirement, resolution, and popularity. Moreover, 330 building datasets were collected and evaluated. Of those, 16 were deemed adequate to feature building performing demand response or building-to-grid (B2G) services. The DSM strategy, building scope, grid type, control strategy, needed data features, and usability of these selected 16 datasets were analyzed. This review reveals future opportunities to address limitations in the existing literature: (1) developing new data-driven methodologies to specifically evaluate different energy flexibility strategies and B2G services of existing buildings; (2) developing baseline-free KPIs that could be calculated from easily accessible building sensors and meter data; (3) devoting non-engineering efforts to promote building energy flexibility, standardizing data-driven energy flexibility quantification and verification processes; and (4) curating and analyzing datasets with proper description for energy flexibility assessm.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Plastic additives in the ocean: Use of a comprehensive dataset for meta-analysis and method development

In excess of 13,000 chemicals are added to plastics (‘additives’) to improve performance, durability, and production of plastic products. They are categorized into numerous chemical classes including flame retardants, light stabilizers, antioxidants, and plasticizers. While research on plastic additives in the marine environment has increased over the past decade, there is a lack of methodological standardization. To direct future measurement of plastic additives, we compiled a first-of-its-kind dataset of literature assessing plastic additives in marine environments, delineated by sample type (plastic debris, seawater, sediment, biota). Using this dataset, we performed a meta-analysis to summarize the state of the science. Currently, our dataset includes 217 publications published between 1978 and May 2023. The majority of publications analyzed plastic additives in biota collected from Europe and Asia. Analyses concentrated on plasticizers, brominated flame retardants, and bisphenols. Common sample preparation techniques included Solvent - Agitation extraction for plastic, sediment, and biota samples, and Solid Phase Extraction for seawater samples with dichloromethane and solvent mixtures including dichloromethane as the organic extraction solvent. Finally, most analyses were performed utilizing gas chromatography/mass spectrometry. There are a variety of data gaps illuminated by this meta-analysis, most notably the small number of compounds that have been targeted for detection compared to the large number of additives used in plastic production. The provided dataset facilitates future investigation of trends in plastic additive concentration data in the marine environment (allowing for comparison to toxicity thresholds) and acts as a starting point for optimizing and harmonizing plastic additive analytical methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Case Study: NREL Campus Chilled Water Storage Potential: Benchmark Datasets Development and Applications, Task 4 - Use Case Demonstration

The Benchmark Datasets Development and Applications project is a three-year collaboration between the National Renewable Energy Laboratory (NREL), Oak Ridge National Laboratory, Pacific Northwest National Laboratory, and Lawrence Berkeley National Laboratory. The project seeks to collect and curate high-resolution, well-calibrated time series of building operational and indoor/outdoor environmental data, which are crucial to understanding and optimizing building energy efficiency performance and demand flexibility capabilities as well as benchmarking energy algorithms. Project outcomes include approximately twelve high-fidelity building datasets, enhanced data representation tools, and four case studies to illustrate example applications. The goal of these case studies is to define and execute analyses that demonstrate how one or more datasets collected through this project can address a data gap or challenge historically faced by building stakeholders. This technical paper summarizes the findings of one of these case studies, in which we studied the operational efficiencies of the central cooling system at NREL. We looked at three years of data from the three chillers in the Field Test Laboratory Building (FTLB), from 2019 to 2021, to compare equipment operation and demand throughout the time period. Our analysis indicates that all three chillers are operating at or below the optimal loading conditions for most of the operation time, and thus there was no efficiency drop due to loading of the chillers at full capacity. Our recommendation is that no chiller capacity increase is needed; instead, the central plant could benefit from adopting advanced control logics for optimal sequencing of chillers during part load operations. Analysis of adding chilled water thermal storage to the central plant indicated 34% savings in demand cost and 24.5% savings in total cost (energy consumption and demand charge cost). The payback period is estimated to be 11-22 years with an assumed TES cost of $\$$100-$200 per ton. This case study shows how a selected dataset is used to solve a practical building problem - learning the operational status of its components, analyzing the effectiveness of a proposed new technique, and aiding decision-making for the building operations and maintenance team.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Preventing Failures By Dataset Shift Detection in Safety-Critical Graph Applications

Dataset shift refers to the problem where the input data distribution may change over time (e.g., between training and test stages). Since this can be a critical bottleneck in several safety-critical applications such as healthcare, drug-discovery, etc., dataset shift detection has become an important research issue in machine learning. Though several existing efforts have focused on image/video data, applications with graph-structured data have not received sufficient attention. Therefore, in this paper, we investigate the problem of detecting shifts in graph structured data through the lens of statistical hypothesis testing. Specifically, we propose a practical two-sample test based approach for shift detection in large-scale graph structured data. Our approach is very flexible in that it is suitable for both undirected and directed graphs, and eliminates the need for equal sample sizes. Using empirical studies, we demonstrate the effectiveness of the proposed test in detecting dataset shifts. We also corroborate these findings using real-world datasets, characterized by directed graphs and a large number of nodes.

97 MATHEMATICS AND COMPUTING↗

Predicting cutoff L-shells of solar protons using the GPPSn particle dataset

Solar energetic protons (SEPs) arriving at the Earth trigger severe radiation storms in the near-Earth space, directly impacting space missions operating at various altitudes. Therefore, monitoring SEP events and predicting the penetration depths of solar protons are critical for aerospace sectors. Building on previous efforts, here we demonstrate the feasibility of using proton measurements from the Global Prompt Proton Sensor network (GPPSn), enabled by Los Alamos National Laboratory developed combined X-ray dosimeters aboard GPS satellites, to characterize and predict the penetration of solar protons into the geomagnetic field. The inclined medium-Earth-orbits (MEOs) of the global GPS constellation offer a unique advantage of allowing simultaneous measurements of penetrating solar protons inside both open- and closed-field line regions. Therefore, the L-profiles of ∼10s–100 MeV solar protons and their associated cutoff L-shells can be determined from the GPPSn dataset, using predefined threshold proton flux values rather than traditional flux ratios. After examining a list of SEP event intervals across solar cycles 23, 24 and 25—including the 2024 Mother’s Day superstorm, we showcase how the latest GPPSn proton dataset (release v1.10), reprocessed and calibrated, can not only be used to monitor solar proton distributions inside the dynamic geomagnetic field for individual events, but also to derive a new empirical model linking cutoff L-shells with several key space weather parameters. This newly developed SEPCL-MEO model demonstrates high predictive performance; for example, predictions for > 30 MeV solar protons yield a correlation coefficient of 0.85 and performance efficiency of 0.67 when validated against GPPSn observations. Results from this pilot study underscores the scientific and operational value of the GPPSn dataset, and this dataset—when paired with machine-learning techniques—can play a critical role in observing and predicting the effects of future incoming SEP events, including extreme ones.

58 GEOSCIENCES↗

The Kimberlina synthetic multiphysics dataset for CO 2 monitoring investigations

Abstract We present a synthetic multi‐scale, multi‐physics dataset constructed from the Kimberlina 1.2 CO 2 reservoir model based on a potential CO 2 storage site in the Southern San Joaquin Basin of California. Among 300 models, one selected reservoir‐simulation scenario produces hydrologic‐state models at the onset and after 20 years of CO 2 injection. Subsequently, these models were transformed into geophysical properties, including P‐ and S‐wave seismic velocities, saturated density where the saturating fluid can be a combination of brine and supercritical CO 2 , and electrical resistivity using established empirical petrophysical relationships. From these 3D distributions of geophysical properties, we have generated synthetic time‐lapse seismic, gravity and electromagnetic responses with acquisition geometries that mimic realistic monitoring surveys and are achievable in actual field situations. We have also created a series of synthetic well logs of CO 2 saturation, acoustic velocity, density and induction resistivity in the injection well and three monitoring wells. These were constructed by combining the low‐frequency trend of the geophysical models with the high‐frequency variations of actual well logs collected at the potential storage site. In addition, to better calibrate our datasets, measurements of permeability and pore connectivity have been made on cores of Vedder Sandstone, which forms the primary reservoir unit. These measurements provide the range of scales in the otherwise synthetic dataset to be as close to a real‐world situation as possible. This dataset consisting of the reservoir models, geophysical models, simulated time‐lapse geophysical responses and well logs forms a multi‐scale, multi‐physics testbed for designing and testing geophysical CO 2 monitoring systems as well as for imaging and characterization algorithms. The suite of numerical models and data have been made publicly available for downloading on the National Energy Technology Laboratory's (NETL) Energy Data Exchange (EDX) website.

58 GEOSCIENCES↗

Resonant anomaly detection with multiple reference datasets

An important class of techniques for resonant anomaly detection in high energy physics builds models that can distinguish between reference and target datasets, where only the latter has appreciable signal. Such techniques, including Classification Without Labels (CWoLa) and Simulation Assisted Likelihood-free Anomaly Detection (SALAD) rely on a single reference dataset. They cannot take advantage of commonly available multiple datasets and thus cannot fully exploit available information. In this work, we propose generalizations of CWoLa and SALAD for settings where multiple reference datasets are available, building on weak supervision techniques. We demonstrate improved performance in a number of settings with realistic and synthetic data. As an added benefit, our generalizations enable us to provide finite-sample guarantees, improving on existing asymptotic analyses.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗