DOE OSTI · 3030897
PRACtical LLM Evaluation using Performance, Response, and Context
Abstract
This project produced practical methods for evaluating Large Language Models (LLMs) based only on characteristics of LLM responses and without ground truth. Further development of these methods will enable the ability to rapidly identify and mitigate different failure types and supports the appropriate use of LLMs in mission applications.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wisniewski, Kyra Lynn [Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)] (ORCID:0009000445269115), Ting, Christina [Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)] (ORCID:0000000328718906), Field, Richard V. [Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)] (ORCID:0000000227657032), Datta, Esha [Sandia National Laboratories (SNL-NM), Albuquerque, NM (United States)]. 2026-04-01. PRACtical LLM Evaluation using Performance, Response, and Context. https://doi.org/10.2172/3030897
Cite the original work for its findings. Save a collection to share your selection of sources.