Emergent Behaviors and the Limits of Large Language Model Generalization
Over the past couple years, large language models (LLMs) have rapidly grown in size and performance across wide numbers of tasks, leading some to herald them as machine learning models with truly “general” capabilities. We address some of the ways that the language used to describe these models can be both deceptive and a poor framing for making progress towards understanding their capabilities, and also review arguments on to what extent these models can represent and interact with the meaning of language. Emergent behaviors in these highly complex models are unpredictable and research into understanding them is still in early phases, which limits the strength of claims of generality about them. Moreover, while these models demonstrate surprising emergent reasoning capabilities on many tasks, there are still many limitations in generalization that overall high benchmark performance can hide. We review the literature showcasing limitations of LLMs as well as techniques to mitigate or overcome these challenges, while also highlighting how fundamental problems in model evaluation may prevent true claims of generality. We conclude with a practical section on using LLMs on real world problems while appropriately evaluating and validating these tasks to prevent or mitigate the downstream impacts of incorrect results.