The Bingo Board
Agentic coding, graded.
Every model and harness that matters, scored across five real-world dimensions. The facts sync themselves; the grades are mine.
synced 7d ago
› How we got here
-
2024-06
The autocomplete ceiling
Copilot-style completion was the state of the art. The model suggested; you typed.
-
2025-02
Agents enter the chat
First wave of agentic CLIs. Models started running commands, reading repos, and iterating.
-
2025-09
The harness wars begin
Codex, Claude Code, Cursor, and the open-source harnesses split the field. The harness became as important as the model.
-
2026-01
Reasoning tiers everywhere
Low to extra-high reasoning dials, ultracode, fast variants. One model became five depending on how hard you let it think.
-
2026-06
Frontier pack consolidates
Fable 5, Opus 4.8, GPT-5.5, and Composer 2.5 set the mid-year board. Million-token contexts and harness wars were the story.
-
2026-07
Sol, Flash, and the open wave
GPT-5.6 Sol takes the independent SWE-bench Verified lead. Gemini 3.6 Flash jumps computer-use as a cheap tier. Claude Opus 5 / Sonnet 5 and Grok 4.5 land as provisional grades while receipts catch up. Open-weight pressure graduates onto the board: Kimi K3 (Moonshot), DeepSeek V3, and GLM-4.5 (Zhipu), graded in the value/BYOK lane.
-
2026-08
Harnesses and workhorse Flash
Muse Code (Meta) ships as a new terminal/CI harness co-trained with Muse Spark 1.2. Grok 4.6 lands in Cursor and Grok Build for long-running agents. Gemini 3.7 Flash is the new cheap coding and agent workhorse. Qwen3.8-Max adds Max-class open-weight pressure (bespoke license, not Apache). SWE-Bench ProMax (170 multilingual refactor tasks, avg 11.4 files / 261.6 LOC; best model 41.2%) is the unsaturated yardstick after the SWE-bench Verified test-quality audit. August add-ons are watchlist rows. Facts synced. No letter grades invented.
Top of the board
Claude Fable 5
Best aggregate grade across all five dimensions.
Best value
Composer 2.5
Frontier-class coding for about seven cents a task.
What I run
Cursor Ultra + Codex
Daily seat. Cloud Auto (Composer 2.5 + Grok) for parallel work. Codex for knowledge, GTM, and ops. Not a letter grade.
Computer-use king
Gemini 3.6 Flash
83% OSWorld as a cheap Flash tier. Volume agent default.
The board
26 of 30 graded
| Model | |||||
|---|---|---|---|---|---|
GPT-5.5 OpenAI | A+ | A+ | A+ | A | A |
GPT-5.6 Sol OpenAI | S | S | A+ | A | A+ |
GPT-5.4 OpenAI | A | A | A | A | A+ |
GPT-5.3 Codex OpenAI | A | A | B+ | B+ | A |
GPT-5.3 Codex Spark OpenAI | B+ | B | B | B | B+ |
GPT-5.2 Codex OpenAI | B+ | B+ | B | B | B+ |
Claude Fable 5 Anthropic | S | S | S | A+ | A+ |
Claude Opus 4.8 Anthropic | A+ | S | A+ | A+ | S |
Claude Opus 4.7 Anthropic | A+ | A+ | A | A | A+ |
Claude Opus 4.6 Anthropic | B+ | B+ | A | A | A |
Claude Sonnet 4.6 Anthropic | B+ | B+ | A | A | A |
Claude Haiku 4.5 Anthropic | B | B | B+ | B+ | A |
Composer 2.5 Cursor | A | A | B | B | B+ |
Composer 2.5 Fast Cursor | A | B+ | B | B | B+ |
Gemini 3.5 Flash | A | B+ | A | A | A |
Gemini 3.6 Flash | A+ | A | A | A | S |
Gemini 3.5 Flash-Lite | B+ | B | B+ | B+ | A |
Gemini 3.5 Pro | A | A | A+ | A+ | A |
Grok 4.3 xAI | B+ | B | B+ | B+ | B+ |
Claude Opus 5 Anthropic | S | S | A+ | A+ | S |
Claude Sonnet 5 Anthropic | A | A | A+ | A+ | A+ |
Gemini 3.5 Flash Cyber | A | A | A | A | A+ |
Grok 4.5 xAI | A | B+ | A | A | A |
Kimi K3 Moonshot | A | A | A+ | A | A+ |
DeepSeek V3 DeepSeek | A | A | A | B+ | A |
GLM-4.5 Zhipu | A | B+ | A | B+ | A |
Qwen3.8-Max Alibabaungraded | ·· | ·· | ·· | ·· | ·· |
Muse Spark 1.2 Metaungraded | ·· | ·· | ·· | ·· | ·· |
Grok 4.6 xAIungraded | ·· | ·· | ·· | ·· | ·· |
Gemini 3.7 Flash Googleungraded | ·· | ·· | ·· | ·· | ·· |
Some harness cards include outbound links (UTM-tagged referrals; affiliate codes when available). Grades are independent of those links. I do not sell rank. See sponsorships for partner placement rules.