A new study published on ArXiv (cs.AI) challenges the prevailing notion that massive scale is the only path toward AI models understanding human cognition. Researchers trained 14 models, ranging from 135 million to 14 billion parameters, across four architecture families using the Psych-101 dataset—a collection of 10.7 million trial-level choices from 160 psychological experiments.

The Scaling Paradox

The findings indicate that for in-distribution data, model scale barely impacts performance. Models with just 0.6B to 1B parameters were sufficient to match the performance of a 70B parameter baseline when tested on held-out participants. However, the importance of scale becomes evident in out-of-distribution (OOD) scenarios. When generalizing to entirely novel task structures, larger models demonstrated a markedly steeper scaling gradient, proving more capable of handling unfamiliar cognitive paradigms.

Diagnostic Insights

To determine what information these models prioritize, the researchers conducted diagnostic tests by stripping four prompt channels: task instructions, experimental stimuli, outcome feedback, and choice history. Masking stimuli and feedback content destroyed 75.7% of learned information, pushing model performance below chance levels. This suggests that these models do not rely solely on statistical shortcuts like choice history but actively process task-specific information.

While their scope remains bounded by the paradigms encountered during training, these small, cognitively fine-tuned models show promise as efficient estimators for psychological experiments, providing a sophisticated proxy for human behavioral data without the computational cost of massive LLMs.