DOE OSTI · 3002371
Scalable workflow for evaluating and optimizing large language models
Abstract
This work describes the improved workflow for evaluating open-source large language models (LLMs) for trustworthiness. The workflow facilitates the acquisition of LLMs, the generation of LLM responses, and the evaluation of the responses for their trustworthiness. As a use case, the workflow is employed to evaluate dense, quantized, and pruned Meta Llama3.1 LLMs for their truthfulness. The outcome of the project could set the stage for understanding and developing trustworthy models in the future projects.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jin, Zheming [Oak Ridge National Laboratory (ORNL), Oak Ridge, TN (United States)]. 2025-06-01. Scalable workflow for evaluating and optimizing large language models. https://doi.org/10.2172/3002371
Cite the original work for its findings. Save a collection to share your selection of sources.