Engineering Papers⌕ Search

Engineering topics

Dodge, Haley Diane

Publications and source records attributed to Dodge, Haley Diane.

Predictive Indicators of the Performance of Large Language Models

In several mission contexts, it is desirable to estimate the performance of large language models (LLMs) on tasks that we cannot run directly. In light of published “scaling laws” our hypothesis is that some tasks should be consistently more challenging than others based on characteristics of the task. The goal of this project was to begin quantifying how much information about LLM performance can be gained from the features of a model and a task. Two of our statistical models struggled to converge. Pass/fail test results may provide limited information for inference beyond model quality and task difficulty, but we see no evidence at this time for significant feature interaction effect sizes, arguing for simple models. Future work extending the models to capitalize on perplexity of ground truth answers is suggested. This project also introduces “Depth of Knowledge Variant Testing” as a strategy for more finely assessing language models on open domain question and answer tasks. We developed sets of questions that ask a language model to produce similar information while demonstrating increasing depth of knowledge, and also relabeled existing Q&A test questions with their depth of knowledge. Our results suggest further consideration of Bloom’s taxonomy and further refinement of prompts to properly elicit information at varying depths. In the course of this work, we set up a basic infrastructure for standardizing tasks and testing many language models on these tasks. In addition to testing the predictive quality of model features and performance across test suites, with this project we have introduced two new task features to contextualize each test question: the Dewey Classification main category of information covered, and the Bloom’s taxonomy level that corresponds to the depth of knowledge probed by the question. Splits across these and other features produced over five hundred task subtypes with distinct feature vectors, which we tested on half a dozen models.

97 MATHEMATICS AND COMPUTING↗

Stingray Case Study

Improvised explosive devices (IEDs) have injured and killed numerous soldiers and civilians as a consequence of military operations in the Middle East. With gaps in existing technologies, the US military required devices for quickly addressing deadly IEDs without harming military personnel or inflicting severe damage to the environment. In response to this need, Sandia developed Stingray, a clear, plastic handheld device to quickly and safely disable threatening IEDs. Stingray is designed to be used in two configurations: a coherent water blade for cutting operations and as a water slug for general device disruption. Prior to disabling an IED, Explosive Ordnance Disposal (EOD) technicians will x-ray the target using tools such as Sandia's X-ray Toolkit (XTK) to determine which function operators should use to dismantle the IED. For example, if the operator knows exactly which wires to cut, they can use the precision water blade. If the operator wants to create a general disruption, they can use the water slug function.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗