Multimodal Large Language Models (MLLMs) have become a leading approach to visual reasoning, yet the field still lacks a quantitative synthesis that rigorously identifies the drivers of their performance. In this paper, we evaluate 21 models across 7 standardised benchmarks, yielding 102 model–benchmark observations, using a two-level mixed-effects design in which benchmark scores are nested within models. The multilevel model reveals substantial between-model variance (intra-class correlation coefficient, ICC = 0.801) and a significant positive association between model scale and standardised performance (β = 0.73, p = 0.016).
Training strategy also matters: relative to instruction tuning, the pretraining + alignment approach is associated with substantially lower performance (β = −1.88, p
Extract — continue reading at the source.