July 2026 Bingo Board Refresh: Sol, Gemini 3.6 Flash, and What Actually Moved

GPT-5.6 Sol takes the independent SWE-bench Verified lead. Gemini 3.6 Flash jumps computer-use as a cheap Flash tier. Flash-Lite becomes the latency knife. Here is how the Bingo Board moved in July, and what I am actually running.

Chris Watkins 5 min read

Listen in my voice · AI narration (ElevenLabs clone)

Loading audio player…
On this page

The mid-year board was Fable 5, Opus 4.8, GPT-5.5, and Composer 2.5. That pack still matters. July did not wipe the board. It added a new coding ceiling, a new computer-use default, and a cheap latency knife.

I refreshed the Bingo Board live facts and human grades for three models that actually changed how I route work:

  1. GPT-5.6 Sol
  2. Gemini 3.6 Flash
  3. Gemini 3.5 Flash-Lite

This is the editorial note for that refresh. Facts sync from the machine layer. Grades stay human.

What moved

GPT-5.6 Sol took the OpenAI coding crown

Independent SWE-bench Verified at 96.2% puts Sol ahead of Fable 5’s 95% on the raw coding receipt. On Codex Pro, that is the new default when the ticket is hard and the harness is terminal-native.

That does not make Sol the top of my aggregate board. Fable still wins overall on knowledge depth and the five-dimension score. Opus still wins the just-talking-to-it feel and long multi-turn engineering sessions. Sol wins the “highest verified coding number right now” lane, and that lane is real.

If your muscle memory is GPT-5.5 inside Codex, keep 5.5 as the proven workhorse. Treat Sol as the ceiling you open when correctness is the whole game.

Gemini 3.6 Flash jumped computer-use as a Flash tier

83% OSWorld-Verified as a cheap Flash model is the story. Not “Pro finally shipped.” Not another unconfirmed preview. A volume tier that actually clicks through real desktops.

Coding on 3.6 Flash is strong enough for agentic workhorses. It is not Sol or Fable on one-shot hardness. For long-horizon computer-use and high-volume Antigravity steps, this is my new Google default. Gemini 3.5 Flash stays in rotation for compatibility and free-tier habits. New agent loops go to 3.6 unless I have a reason not to.

Gemini 3.5 Flash-Lite is the throughput knife

Google’s claim around 350 output tokens per second changes how search-style and document agents feel. You do not bring Flash-Lite to a hard SWE-bench fight. You bring it when latency and cost dominate, then hand off to 3.6 Flash when the loop needs a stronger closer.

Scout and closer. That pairing is the useful mental model.

What did not move (and why that matters)

Claude Opus 4.8 is still my daily driver. The package beat any single benchmark: SWE-bench strength, SWE-bench Pro toughness, and the conversation feel that keeps a long session coherent. Sol did not change that.

Composer 2.5 is still the value king at roughly seven cents a task. July did not invent a cheaper frontier-class editor model that I trust more inside Cursor.

Gemini 3.5 Pro is still provisional. Limited enterprise preview, thin public receipts. Do not bet production agent loops on unconfirmed Pro numbers when 3.6 Flash is already shipping.

Open-weight pressure (Kimi, DeepSeek, GLM) is real in the discourse and the value lane. Kimi K3 and DeepSeek V3 are now on the board as provisional value/BYOK grades. Independent coding receipts will move the letters. GLM stays watchlist until I have a row I can stand behind.

How I am routing July work

JobReach for
Hard correctness-critical codingFable 5 (aggregate top) or Sol on Codex (coding ceiling)
Daily multi-turn engineeringOpus 4.8
Terminal-native OpenAI loopsSol first, GPT-5.5 as workhorse
Computer use / cheap desktop agentsGemini 3.6 Flash
Latency and volume scoutingGemini 3.5 Flash-Lite → escalate to 3.6
High-volume Cursor editsComposer 2.5

Plans did not get less important. Codex Pro at $100 still looks like the best-value frontier coding plan I run, and Sol is why that plan got more interesting this month. Claude Max is still where Opus lives as flat-rate daily work. Antigravity is where Google’s Flash stack earns its keep.

Board vs hype

A few rules I am keeping:

  • Benchmarks are receipts, not religion. Sol winning independent SWE-bench Verified matters. It does not automatically win knowledge, content, or “do I want to talk to this for three hours.”
  • Harness fit beats model sheet music. Sol inside Codex is a different product than Sol in a random chat box. Opus inside Claude Code is why 4.8 stays sticky.
  • Cheap tiers can change defaults. 3.6 Flash is the clearest example this month: computer-use stopped being a premium-only story.
  • Grades stay human. The sync layer can update prices and scores. It cannot decide whether something is my daily driver.

Where to look

July’s lesson is simple. The frontier pack did not collapse. The routing table got sharper. Sol raised the coding ceiling. Gemini made computer-use cheap. Everything else is still earned at the keyboard.