arrow_backBack to Radio
News

Kimi Open-Sources PerceptionBench Visual Perception Benchmark, GPT-5.6-Sol Achieves Top Accuracy but No Model Exceeds 60%

en
July 29th News, Kimi announced the open-sourcing of its multimodal large model visual perception evaluation benchmark, PerceptionBench. The benchmark aims to decompose visual perception capabilities into 10 atomic-level abilities for independent assessment, covering dimensions such as visual relationships, counting, attributes, depth & 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection. The benchmark is built upon model failure cases from 42 existing evaluation datasets, containing a total of 3,000 manually verified questions. Each question examines a single visual capability without requiring reasoning or external knowledge. Evaluation results show that among 16 cutting-edge multimodal large models, no model achieved an overall accuracy exceeding 60%. GPT-5.6-Sol ranked first with 59.7% accuracy, followed by Kimi K3 (58.5%), Claude-Fable-5 (57.2%), Gemini-3.1-Pro (56.2%), and GPT-5.5 (55.8%) in the top five. The report points out that visual hallucination remains the weakest capability across all models, indicating significant room for improvement in overall perception abilities.
Share