Meta’s new GAMUT benchmark scored the best AI model just 58.7% in delivering fully complete answers, exposing a gap few tests reveal.

Rather than measuring only accuracy or fluency, GAMUT, short for Grounded Assessment of Multimodal Factuality, focuses on whether AI responses include all necessary facts. The benchmark uses 1,813 expert-crafted yes-or-no questions, structured by relevance and category, to evaluate how completely an AI covers key information.

This evaluation set is built on real-world wearable imagery and spans ten diverse domains, requiring AI models to interpret visuals and generate full textual answers. When Meta tested 14 AI models against GAMUT, Gemini 3.1 Pro led, yet still managed only a 58.7% completeness score.

The benchmark’s hierarchical rubric design prevents the usual superficial grading, uncovering overlooked omissions. Meta has shared GAMUT publicly on Hugging Face, inviting broader testing and development.