The Bingo Board

Agentic coding, graded.

Every model and harness that matters, scored across five real-world dimensions. The facts sync themselves; the grades are mine.

synced 7d ago

How we got here
  1. 2024-06

    The autocomplete ceiling

    Copilot-style completion was the state of the art. The model suggested; you typed.

  2. 2025-02

    Agents enter the chat

    First wave of agentic CLIs. Models started running commands, reading repos, and iterating.

  3. 2025-09

    The harness wars begin

    Codex, Claude Code, Cursor, and the open-source harnesses split the field. The harness became as important as the model.

  4. 2026-01

    Reasoning tiers everywhere

    Low to extra-high reasoning dials, ultracode, fast variants. One model became five depending on how hard you let it think.

  5. 2026-06

    Frontier pack consolidates

    Fable 5, Opus 4.8, GPT-5.5, and Composer 2.5 set the mid-year board. Million-token contexts and harness wars were the story.

  6. 2026-07

    Sol, Flash, and the open wave

    GPT-5.6 Sol takes the independent SWE-bench Verified lead. Gemini 3.6 Flash jumps computer-use as a cheap tier. Claude Opus 5 / Sonnet 5 and Grok 4.5 land as provisional grades while receipts catch up. Open-weight pressure graduates onto the board: Kimi K3 (Moonshot), DeepSeek V3, and GLM-4.5 (Zhipu), graded in the value/BYOK lane.

  7. 2026-08

    Harnesses and workhorse Flash

    Muse Code (Meta) ships as a new terminal/CI harness co-trained with Muse Spark 1.2. Grok 4.6 lands in Cursor and Grok Build for long-running agents. Gemini 3.7 Flash is the new cheap coding and agent workhorse. Qwen3.8-Max adds Max-class open-weight pressure (bespoke license, not Apache). SWE-Bench ProMax (170 multilingual refactor tasks, avg 11.4 files / 261.6 LOC; best model 41.2%) is the unsaturated yardstick after the SWE-bench Verified test-quality audit. August add-ons are watchlist rows. Facts synced. No letter grades invented.

Top of the board

Claude Fable 5

Best aggregate grade across all five dimensions.

Best value

Composer 2.5

Frontier-class coding for about seven cents a task.

What I run

Cursor Ultra + Codex

Daily seat. Cloud Auto (Composer 2.5 + Grok) for parallel work. Codex for knowledge, GTM, and ops. Not a letter grade.

Computer-use king

Gemini 3.6 Flash

83% OSWorld as a cheap Flash tier. Volume agent default.

The board

26 of 30 graded

Model

GPT-5.5

OpenAI

A+A+A+AA

GPT-5.6 Sol

OpenAI

SSA+AA+

GPT-5.4

OpenAI

AAAAA+

GPT-5.3 Codex

OpenAI

AAB+B+A

GPT-5.3 Codex Spark

OpenAI

B+BBBB+

GPT-5.2 Codex

OpenAI

B+B+BBB+

Claude Fable 5

Anthropic

SSSA+A+

Claude Opus 4.8

Anthropic

A+SA+A+S

Claude Opus 4.7

Anthropic

A+A+AAA+

Claude Opus 4.6

Anthropic

B+B+AAA

Claude Sonnet 4.6

Anthropic

B+B+AAA

Claude Haiku 4.5

Anthropic

BBB+B+A

Composer 2.5

Cursor

AABBB+

Composer 2.5 Fast

Cursor

AB+BBB+

Gemini 3.5 Flash

Google

AB+AAA

Gemini 3.6 Flash

Google

A+AAAS

Gemini 3.5 Flash-Lite

Google

B+BB+B+A

Gemini 3.5 Pro

Google

AAA+A+A

Grok 4.3

xAI

B+BB+B+B+

Claude Opus 5

Anthropic

SSA+A+S

Claude Sonnet 5

Anthropic

AAA+A+A+

Gemini 3.5 Flash Cyber

Google

AAAAA+

Grok 4.5

xAI

AB+AAA

Kimi K3

Moonshot

AAA+AA+

DeepSeek V3

DeepSeek

AAAB+A

GLM-4.5

Zhipu

AB+AB+A

Qwen3.8-Max

Alibabaungraded

··········

Muse Spark 1.2

Metaungraded

··········

Grok 4.6

xAIungraded

··········

Gemini 3.7 Flash

Googleungraded

··········

Some harness cards include outbound links (UTM-tagged referrals; affiliate codes when available). Grades are independent of those links. I do not sell rank. See sponsorships for partner placement rules.