NeuralZone | AI Apps: post #407 — TG.ME

Forwarded fromHOHow AI Helps
Six AI coding agents took one visual IQ test, and Codex 5.5 won by method, speed, and cost

One small test asked agents to solve 25 visual puzzles on iq-test.cc, select age 30, and return a result link.

This was not a lab benchmark. It was a practical check of vision work, browser use, patience, time, and plan cost.

"Take the IQ test on iq-test.cc. When you finish, select age 30 and send me the link to your result."

Agent                    IQ   Time   Limit spent
Claude Cowork Opus 4.8 90 85m ~10 pts
Claude Code Opus 4.8 90 96m ~28 pts
Claude Sonnet 4.6 68 62m n/a
Codex 5.5 $100 Fast 124 18m ~12 pts
Codex 5.4 $100 Fast 101 16m ~14 pts
Codex 5.5 $200 Fast 131 34m ~6 pts


The score is only part of the story. Codex 5.5 did better because it worked like a careful test taker: collect puzzle images, build clean contact sheets, zoom into hard cases, then recheck weak answers before submit.

More context: the top IQ 131 run used a shorter prompt and the site default age, so it was not a perfect same-prompt run. Still, normal browser access was missing, and Codex found another path through Chrome, clicked all 25 answers, and finished anyway.

Claude was careful, especially Opus. It wrote notes and reasoned step by step. Codex was more organized and faster. The article shows screenshots, failed paths, exact prompts, and puzzle examples.

The most useful lesson: for visual web tasks, method can beat size. A huge context window did not save Claude, and two extra Codex minutes were worth 23 IQ points.

read details on our website

Please support this young channel by subscribing.
Your subscription really helps us grow.
There are no ads here.
❤55👍14🔥5
July 12, 2026 69.9K 5