Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

A FAIR and AI-ready Higgs boson decay dataset

Abstract To enable the reusability of massive scientific datasets by humans and machines, researchers aim to adhere to the principles of findability, accessibility, interoperability, and reusability (FAIR) for data and artificial intelligence (AI) models. This article provides a domain-agnostic, step-by-step assessment guide to evaluate whether or not a given dataset meets these principles. We demonstrate how to use this guide to evaluate the FAIRness of an open simulated dataset produced by the CMS Collaboration at the CERN Large Hadron Collider. This dataset consists of Higgs boson decays and quark and gluon background, and is available through the CERN Open Data Portal. We use additional available tools to assess the FAIRness of this dataset, and incorporate feedback from members of the FAIR community to validate our results. This article is accompanied by a Jupyter notebook to visualize and explore this dataset. This study marks the first in a planned series of articles that will guide scientists in the creation of FAIR AI models and datasets in high energy particle physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SRBench++: Principled Benchmarking of Symbolic Regression With Domain-Expert Interpretation

Symbolic regression searches for analytic expressions that accurately describe studied phenomena. The main promise of this approach is that it may return an interpretable model that can be insightful to users, while maintaining high accuracy. The current standard for benchmarking these algorithms is SRBench, which evaluates methods on hundreds of datasets that are a mix of real-world and simulated processes spanning multiple domains. At present, the ability of SRBench to evaluate interpretability is limited to measuring the size of expressions on real-world data, and the exactness of model forms on synthetic data. In practice, model size is only one of many factors used by subject experts to determine how interpretable a model truly is. Furthermore, SRBench does not characterize algorithm performance on specific, challenging sub-tasks of regression such as feature selection and evasion of local minima. In this work, we propose and evaluate an approach to benchmarking SR algorithms that addresses these limitations of SRBench by 1) incorporating expert evaluations of interpretability on a domain-specific task, and 2) evaluating algorithms over distinct properties of data science tasks. We evaluate 12 modern symbolic regression algorithms on these benchmarks and present an in-depth analysis of the results, discuss current challenges of symbolic regression algorithms and highlight possible improvements for the benchmark itself.

97 MATHEMATICS AND COMPUTING↗

Evaluating a Commercial Dynamic Line Rating Software with the National PMU Dataset

To accelerate the development of data-driven applications for power systems, the Department of Energy (DOE) supported the collection and curation of a synchrophasor dataset spanning two years of observations from transmission utilities across the US. This National PMU Dataset (NPDS) was anonymized and distributed to awardees of a DOE research grant under nondisclosure agreements (NDAs) but has also been retained at PNNL to enable further research. Agreements with data contributors prevent the data from being shared outside the organization. However, establishing a blind research validation methodology is envisioned to maximize the value proposition of the NPDS. In this validation strategy, researchers may share algorithms/software (potentially as executables to protect intellectual property) with PNNL, and PNNL will share feedback about the software’s performance on subsets of the NPDS. Such a blind methodology ensures that sensitive information about critical infrastructure remains protected, but the value of the NPDS can be extended to research beyond PNNL. Through iterative feedback, the algorithms may be tweaked to address real-world artifacts. As the NPDS data is temporally and geographically diverse, it may capture features absent in smaller datasets used during the development of the algorithm under test. This report presents lessons learned from applying the blind validation methodology to LineID™, a synchrophasor-based dynamic line rating software developed by Topolonet Corporation. Improvements made to the software through iterative feedback, limitations of the validation methodology, as well as how the limitations of the NPDS affected the evaluation process are discussed. Observations indicate that the proposed validation methodology can be valuable for evaluating other tools in the future.

97 MATHEMATICS AND COMPUTING↗

High-Resolution Image Products Acquired from Mid-Sized Uncrewed Aerial Systems for Land–Atmosphere Studies

We assess the viability of deploying commercially available multispectral and thermal imagers designed for integration on small uncrewed aerial systems (sUASs, <25 kg) on a mid-size Group-3-classification UAS (weight: 25–600 kg, maximum altitude: 5486 m MSL, maximum speed: 128 m/s) for the purpose of collecting a higher spatial resolution dataset that can be used for evaluating the surface energy budget and effects of surface heterogeneity on atmospheric processes than those datasets traditionally collected by instrumentation deployed on satellites and eddy covariance towers. A MicaSense Altum multispectral imager was deployed on two very similar mid-sized UASs operated by the Atmospheric Radiation Measurement (ARM) Aviation Facility. This paper evaluates the effects of flight on imaging systems mounted on UASs flying at higher altitudes and faster speeds for extended durations. We assess optimal calibration methods, acquisition rates, and flight plans for maximizing land surface area measurements. We developed, in-house, an automated workflow to correct the raw image frames and produce final data products, which we assess against known spectral ground targets and independent sources. We intend this manuscript to be used as a reference for collecting similar datasets in the future and for the datasets described within this manuscript to be used as launching points for future research.

47 OTHER INSTRUMENTATION↗

Summarizing Multiple Aspects of Triple Collocation Analysis in a Single Diagram

With the ongoing expansion of global observation networks, it is expected that we shall routinely analyze records of geophysical variables such as temperature from multiple collocated instruments. Validating datasets in this situation is not a trivial task because every observing system has its own bias and noise. Triple collocation is a general statistical framework to estimate the error characteristics in three or more observational-based datasets. In a triple colocation analysis, several metrics are routinely reported but traditional multiple-panel plots are not the most effective way to display information. A new formula of error variance is derived for connecting the key terms in the triple collocation theory. A diagram based on this formula is devised to facilitate triple collocation analysis of any data from observations, as illustrated using three aerosol optical depth datasets from the recent Aerosol Cloud meTeorology Interactions oVer the western ATlantic Experiment (ACTIVATE). An observational-based skill score is also derived to evaluate the quality of three datasets by taking into account both error variance and correlation coefficient. Several applications are discussed and sample plotting routines are provided.

triple collocation↗

A Surface Radiation Balance Dataset from Siple Dome in West Antarctica for Atmospheric and Climate Model Evaluation

Abstract A field campaign at Siple Dome in West Antarctica during the austral summer 2019/20 offers an opportunity to evaluate climate model performance, particularly cloud microphysical simulation. Over Antarctic ice sheets and ice shelves, clouds are a major regulator of the surface energy balance, and in the warm season their presence occasionally induces surface melt that can gradually weaken an ice shelf structure. This dataset from Siple Dome, obtained using transportable and solar-powered equipment, includes surface energy balance measurements, meteorology, and cloud remote sensing. To demonstrate how these data can be used to evaluate model performance, comparisons are made with meteorological reanalysis known to give generally good performance over Antarctica (ERA5). Surface albedo measurements show expected variability with observed cloud amount, and can be used to evaluate a model’s snowpack parameterization. One case study discussed involves a squall with northerly winds, during which ERA5 fails to produce cloud cover throughout one of the days. A second case study illustrates how shortwave spectroradiometer measurements that encompass the 1.6- μ m atmospheric window reveal cloud phase transitions associated with cloud life cycle. Here, continuously precipitating mixed-phase clouds become mainly liquid water clouds from local morning through the afternoon, not reproduced by ERA5. We challenge researchers to run their various regional or global models in a manner that has the large-scale meteorology follow the conditions of this field campaign, compare cloud and radiation simulations with this Siple Dome dataset, and potentially investigate why cloud microphysical simulations or other model components might produce discrepancies with these observations. significance statement Antarctica is a critical region for understanding climate change and sea level rise, as the great ice sheets and the ice shelves are subject to increasing risk as global climate warms. Climate models have difficulties over Antarctica, particularly with simulation of cloud properties that regulate snow surface melting or refreezing. Atmospheric and climate-related field work has significant challenges in the Antarctic, due to the small number of research stations that can support state-of-the-art equipment. Here we present new data from a suite of transportable and solar-powered instruments that can be deployed to remote Antarctic sites, including regions where ice shelves are most at risk, and we demonstrate how key components of climate model simulations can be evaluated against these data.

54 ENVIRONMENTAL SCIENCES↗

TPCpp-10M: Simulated proton-proton collisions in a time projection chamber for AI foundation models

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community. This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability. The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of $\sqrt{s}$ = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory. This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.

Data Analysis, Statistics and Probability (physics↗

Evaluating the Uncertainty of Terrestrial Water Budget Components over High Mountain Asia

This study explores the uncertainties in terrestrial water budget estimation over High Mountain Asia (HMA) using a suite of uncoupled land surface model (LSM) simulations. The uncertainty in the water balance components of precipitation (P), evapotranspiration (ET), runoff (R), and terrestrial water storage (TWS) is significantly impacted by the uncertainty in the driving meteorology, with precipitation being the most important boundary condition. Ten gridded precipitation datasets along with a mix of model-, satellite-, and gauge-based products, are evaluated first to assess their suitability for LSM simulations over HMA. The datasets are evaluated by quantifying the systematic and random errors of these products as well as the temporal consistency of their trends. Though the broader spatial patterns of precipitation are generally well captured by the datasets, they differ significantly in their means and trends. In general, precipitation datasets that incorporate information from gauges are found to have higher accuracy with low Root Mean Square Errors and high correlation coefficient values. An ensemble of LSM simulations with selected subset of precipitation products is then used to produce the mean annual fluxes and their uncertainty over HMA in P, ET, and R to be 2.11 ± 0.45, 1.26 ± 0.11, and 0.85 ± 0.36 mm per day, respectively. The mean annual estimates of the surface mass (water) balance components from this model ensemble are comparable to global estimates from prior studies. However, the uncertainty/spread of P, ET, and R is significantly larger than the corresponding estimates from global studies. A comparison of ET, snow cover fraction, and changes in TWS estimates against remote sensing-based references confirms the significant role of the input meteorology in influencing the water budget characterization over HMA and points to the need for improving meteorological inputs.

Terrestrial water budget↗

Weather Sensitive High Spatio-Temporal Resolution Transportation Electric Load Profiles For Multiple Decarbonization Pathways

Electrification of transport compounded with climate change will transform hourly load profiles and their response to weather. We present a novel approach to generating hourly electric load profiles that considers charging strategies and evolving sensitivity to temperature. The approach consists of downscaling annual state-scale sectoral load projections from the multisectoral Global Change Analysis Model (GCAM) into hourly electric load profiles leveraging high resolution climate and population datasets. Profiles are developed and evaluated at the Balancing Authority scale, with a 5-year increment until 2050 over the Western U.S. Interconnect for multiple decarbonization pathways and climate scenarios. The datasets are readily available for production cost model analysis. Our open source approach is transferable to other regions.

Decarbonization, tranportation, Electric vehicle c↗

The Cumulus and Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD)

Low clouds continue to contribute greatly to the uncertainty in cloud feedback estimates. Depending on whether a region is dominated by cumulus (Cu) or stratocumulus (Sc) clouds, the interannual low-cloud feedback is somewhat different in both spaceborne and large-eddy simulation studies. Therefore, simulating the correct amount and variation of the Cu and Sc cloud distributions could be crucial to predict future cloud feedbacks. Here we document spatial distributions and profiles of Sc and Cu clouds derived from Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) and CloudSat measurements. For this purpose, we create a new dataset called the Cumulus And Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD), which identifies Sc, broken Sc, Cu under Sc, Cu with stratiform outflow and Cu. To separate the Cu from Sc, we design an original method based on the cloud height, horizontal extent, vertical variability and horizontal continuity, which is separately applied to both CALIPSO and combined CloudSat–CALIPSO observations. First, the choice of parameters used in the discrimination algorithm is investigated and validated in selected Cu, Sc and Sc–Cu transition case studies. Then, the global statistics are compared against those from existing passive- and active-sensor satellite observations. Our results indicate that the cloud optical thickness – as used in passive-sensor observations – is not a sufficient parameter to discriminate Cu from Sc clouds, in agreement with previous literature. Using clustering-derived datasets shows better results although one cannot completely separate cloud types with such an approach. On the contrary, classifying Cu and Sc clouds and the transition between them based on their geometrical shape and spatial heterogeneity leads to spatial distributions consistent with prior knowledge of these clouds, from ground-based, ship-based and field campaigns. Furthermore, we show that our method improves existing Sc–Cu classifications by using additional information on cloud height and vertical cloud fraction variation. Finally, the CASCCAD datasets provide a basis to evaluate shallow convection and stratocumulus clouds on a global scale in climate models and potentially improve our understanding of low-level cloud feedbacks. The CASCCAD dataset (Cesana, 2019, https://doi.org/10.5281/zenodo.2667637) is available on the Goddard Institute for Space Studies (GISS) website at https://data.giss.nasa.gov/clouds/casccad/ (last access: 5 November 2019) and on the zenodo website at https://zenodo.org/record/2667637 (last access: 5 November 2019).

Cesana, Gregory V.↗

Arctic Mixed-Phase Cloud Base Ice Precipitation Properties During the M-PACE Field Campaign

Cloud-climate feedbacks are still the greatest source of uncertainty in current climate projections. Arctic clouds, which are predominantly stratiform and supercooled, often long-lived, and nearly-continuously precipitate ice particles, contribute roughly 10% of the uncertainty attributed to the global cloud feedback. This Arctic cloud uncertainty is driven by incomplete observational and theoretical knowledge required to estimate and explain the state and active processes occurring in those clouds. A focus on ice precipitation properties at Arctic cloud base rather than the surface deconfounds the product of cloud condensate sink processes from the influence of the atmospheric thermodynamic state below cloud base, rendering cloud-base properties a more appealing target for inference and evaluation of model simulations. This dataset provides a set of 25 samples from the M-PACE field campaign, all of which were retrieved using the synthesis of ARM radar and lidar measurements. The retrieved ice precipitation variables in this dataset include, among others, the ice number concentration, water content, PSD parameters, precipitation rate, mass-weighted fall velocity, vertical air motion, and effective radius, all of which are highly valuable for model evaluation and a general understanding of polar cloud sink processes. Each variable sample includes its mean value and associated uncertainty. Additional variables based on ARM measurements (liquid layer statistics, etc.) are included in this dataset as well. The retrieval algorithm and analysis of this dataset are described in Silber (JGR, 2023, https://doi.org/10.1029/2022JD038202).

54 ENVIRONMENTAL SCIENCES↗

Applying a Multisector Scenario Framework to Evaluate Past and Future Public Surface Water Supply Infrastructure Strategies in Texas

Datasets supporting the index model and scenario analysis used in evaluating surface water supply strategies across different water system types in Texas. These data underpin the scenario development and application of five key indicators: Water Availability Index (WAI), Water Quality Index (WQI), Energy Requirement Index (ERI), Water Treatment Cost (WTC), and Water Infrastructure Cost (WIC). The datasets are organized by system type—stream reaches (flowlines), waterbodies, and reservoirs—and include both raw and standardized index values. The integrated datasets also provide scenario classifications (original and adjusted) based on infrastructure and planning priorities, enabling comparison across Shared Socioeconomic Pathways (SSPs). Additional strategy-level data are included to support evaluation of state-level new reservoir projects in relation to cost and availability tradeoffs. Please refer to the README file provided in Files for more details. Descriptions of the datasets are provided below. Dataset(s) Descriptions Folder: Index_model_database.zip Subfolder: Stream_reach.zip Fl_wf.csv, Fl_wq.csv, Fl_er.csv, Fl_wf_wtcUV.csv, Fl_wf_wtcnoUV.csv, Fl_allfac_wic1.csv, Fl_allfac_wic2.csvDatasets for computing WAI, WQI, ERI, WTC, and WIC for surface water systems classified as stream reaches (flowlines). Subfolder: Waterbody.zip Wb_wf.csv, Wb_wq.csv, Wb_er.csv, Wb_wf_wtcUV.csv, Wb_wf_wtcnoUV.csv, Wb_allfac_wic1.csv, Wb_allfac_wic2.csvEquivalent index model datasets for waterbodies, reflecting hydrologic and infrastructure attributes specific to impounded natural systems. Subfolder: Reservoir.zip Rs_wf.csv, Rs_wq.csv, Rs_er.csv, Rs_wf_wtcUV.csv, Rs_wf_wtcnoUV.csv, Rs_allfac_wic1.csv, Rs_allfac_wic2.csvIndex model datasets specific to regulated reservoir systems, incorporating both resource indicators and cost parameters. Folder: Integrated data.zip combined_merged_data.csv, combined_merged_data_scenario.csvDatasets integrating index model indicators (both raw and scaled) with scenario classifications, including adjustments reflecting SSP-aligned transitions and planning shifts. Folder: Additional data.zip wai_supplystrat_wic_merged.csvCurated dataset capturing proposed major reservoir-based municipal water supply strategies in Texas. Integrates site-level planning data with estimated capital infrastructure costs and water availability scores for comparative assessment.

geospatial↗

A New High-Impedance-Fault Detection Method to Prevent Power-Line-Induced Wildfires

High Impedance Faults (HIFs) occur when energized power lines come into contact with high impedance ground surfaces, such as tree branches and grassland. HIFs have the potential to cause arcing, leading to vegetation ignition and the initiation of wildfires. The challenge in detecting HIFs comes from the high impedance of the partially conductive materials in contact with the power lines. They create a fault current of low magnitude and traditional protective devices struggle to detect such faults. This paper proposes a novel HIF detection algorithm based upon the analyzed arcing signatures associated with HIFs. The algorithm is evaluated using the Australian Public Bushfire Safety Program (PBSP) dataset. For comparative analysis, a state-of-the-art commercial HIF detection product is also evaluated using the same dataset. The proposed algorithm demonstrates higher detection accuracy over the commercial products with fewer false flags and undetected faults.

grasslands↗

Model Data Archive for Manuscript Titled "Evaluation of a Coupled Surface–Subsurface Hydrologic Model Using Dense Water‑Level Sensors in a Mixed Urban–Rural Watershed"

This archive provides scripts, input files, and datasets used for the implementation and evaluation of a fully coupled surface–subsurface hydrologic model in the Neches River Basin, southeast Texas. The study uses the Advanced Terrestrial Simulator (ATS) to simulate coupled surface–subsurface hydrologic processes over a mixed urban–rural watershed and evaluates model performance using a dense network of 136 in situ water-level sensors, nine U.S. Geological Survey (USGS) stream gauges, and SSEBop-derived evapotranspiration estimates during the period October 2014–June 2024. The workflow is implemented primarily in Python 3 using the Watershed Workflow package. The Jupyter notebooks can be executed using open-source software such as Anaconda JupyterLab or Visual Studio Code. Other data files include TXT, CSV, XML, SHP, TIF, NetCDF, HDF5, and ExodusII files, which can be processed using the provided Python scripts. ATS input files are provided in XML format and can be edited using any commonly used text editor. This archive contains: *Scripts and input files used to generate the ATS model setup, including watershed discretization, mesh generation, parameter mapping, and model configuration. *Jupyter notebooks used for preprocessing observational data, evaluating streamflow, water levels, and evapotranspiration, computing performance metrics, and generating the figures presented in the manuscript. *ATS simulation outputs and processed observational datasets, including OneRain and DD6 water-level sensors, USGS streamflow observations, GIS data, and supporting spatial datasets used throughout the study.

Dense water-level sensor network↗

Evaluating lightweight unsupervised online IDS for masquerade attacks in CAN

Vehicular controller area networks (CANs) are susceptible to masquerade attacks by malicious adversaries. In masquerade attacks, adversaries silence a targeted ID and then send malicious frames with forged content at the expected timing of benign frames. As masquerade attacks could seriously harm vehicle functionality and are the stealthiest attacks to detect in CAN, recent work has devoted attention to compare frameworks for detecting masquerade attacks in CAN. However, most existing works report offline evaluations using CAN logs already collected using simulations that do not comply with the domain’s real-time constraints. Here we contribute to advance the state of the art by presenting a comparative evaluation of four different non-deep learning (DL)-based unsupervised online intrusion detection systems (IDS) for masquerade attacks in CAN. Our approach differs from existing comparative evaluations in that we analyze the effect of controlling streaming data conditions in a sliding window setting. In doing so, we use realistic masquerade attacks being replayed from the ROAD dataset. We show that although evaluated IDS are not effective at detecting every attack type, the method that relies on detecting changes in the hierarchical structure of clusters of time series produces the best results at the expense of higher computational overhead. We discuss limitations, open challenges, and how the evaluated methods can be used for practical unsupervised online CAN IDS for masquerade attacks.

Anomaly detection↗

Lake-Effect Snowstorm Events and Associated Snowfall Totals Integrated from NOAA Storm Reports, ERA5, and HRRR for the Laurentian Great Lakes (1997–2024)

Lake-effect snowstorms are localized, impactful winter weather phenomena that can generate substantial snowfall totals and pose significant challenges for forecasting, transportation, and regional infrastructure. To support the analysis and modeling of these events, this dataset compiles observational reports of lake-effect snowstorms alongside corresponding snowfall estimates derived from gridded atmospheric datasets. The observational component of the data originates from the National Weather Service (NWS) winter storm report, subset to lake-effect snow event type, covering 1997–2024. For each lake-effect snow event, this data provides the impacted county, event start and end datetimes at an hourly resolution, as well as relevant storm narratives. The complementary reanalysis-derived data is sourced from European Centre for Medium-Range Weather Forecasts (ECMWF) Reanalysis 5 (ERA5) and High-Resolution Rapid Refresh (HRRR) gridded data. For both gridded datasets, the maximum total snowfall (in units mm) was extracted, constrained by the county and datetimes specified by the observational report. ERA5 data covers the entire observational period (1997–2024), whereas HRRR data is only available from November 2016 – December 2024. Three CSV files are provided here: (1) the observational lake-effect snow event report, (2) ERA5 maximum snowfall detections for each event, and (3) HRRR maximum snowfall detections for each event. Relevant data from the observational files, such as impacted state and county, event datetimes, and event IDs, were included for convenience. Users can inspect and visualize the data using tools such as Microsoft Excel and Python pandas/matplotlib packages. This dataset may support a variety of applications, including climatological analyses of lake-effect snowfall, evaluation of snowfall representation in atmospheric datasets and numerical weather prediction models, and the development of machine learning approaches for detecting or predicting lake-effect snowfall events.

EARTH SCIENCE > ATMOSPHERE > PRECIPITATION > SOLID↗

A Data-Driven Method for Synthetic Extreme Weather Generation and Solar Impact Assessment: Preprint

High-resolution, high-fidelity weather datasets are essential for testing and evaluating the resilience of power systems, particularly under extreme weather conditions. However, existing extreme weather datasets are typically derived from historical events that are localized and may lack the spatial and temporal resolution or scenario diversity needed to test largescale power systems. In this work, we propose a synthetic extreme weather simulation approach capable of generating targeted extreme events, such as hurricanes, using publicly available data sources. Preliminary results demonstrate the impact of a simulated Category 1 hurricane on renewable generation and critical infrastructure in California. The work aims to provide a flexible approach for creating multiple types of extreme weather scenarios across different regions, enabling comprehensive system stress testing, training, and resilience assessment.

24 POWER TRANSMISSION AND DISTRIBUTION↗

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation

Retrieval-augmented generation (RAG) has emerged as a promising paradigm for improving factual accuracy in large language models (LLMs). We introduce a benchmark designed to evaluate RAG pipelines as a whole, evaluating a pipelines ability to ingest several modalities of information. We present (1) a curated dataset of 93 questions designed to evaluate a pipeline's ability to ingest textual data, tables, images, multimodal data, and cross-document multimodal data; (2) a phrase-level recall metric for correctness; (3) a nearest-neighbor embedding classifier in an attempt to classify pipeline hallucinations; (4) a comparative evaluation of 2 pipelines built with open-source retrieval mechanisms and 4 closed-source foundational models; and (5) a third-party human evaluation of the alignment of our correctness and hallucination metrics. We find that closed-source pipelines significantly outperform open-source pipelines in both the correctness and halucination metrics, with a wider performance gap in questions relying on multimodal and cross-document information. We also find after a human evaluation of our correctness and hallucination metric compared with our questions and pipeline responses, average agreement was 4.62 for correctness 4.53 for hallucination detection on a 1-5 Likert scale with 5 being strongly agree with our determination.

Hildebrand, Samuel [ORNL] (ORCID:0009000465963104)↗