New benchmark confirms AI models still perform poorly at visual perception

Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage. The article New benchmark confirms AI models still perform poorly at

The top 3

  1. Contextual Understanding and Common Sense Reasoning: AI struggles to interpret the broader context of a scene and apply common-sense knowledge, leading to misinterpretations of objects or actions that are obvious to humans, such as understanding the intent behind a gesture or distinguishing a toy from a real object in a complex setting.
  2. Generalization to Novel or Out-of-Distribution Data: While AI performs well on training data, it often fails to generalize effectively to new, unseen variations, unusual angles, or different environments, highlighting a lack of true understanding beyond learned patterns.
  3. Handling Occlusions and Cluttered Scenes: AI models find it difficult to accurately identify and delineate objects when they are partially obscured or embedded within highly complex and cluttered backgrounds, unlike humans who can infer missing parts.

Sources

Open the full topic