Living in Color August · Episode 02

Living in Color W32: What Actually Moved

Week of Aug 10: Muse Code as the new harness, SWE-Bench ProMax as the unsaturated refactor yardstick, and Qwen3.8-Max as open-weight Max-class pressure.

Chris Watkins 4 min read

Listen in my voice · AI narration (ElevenLabs clone)

Loading audio player…
On this page

Last week was routing tables. This week is a new harness, a harder bench, and open-weight Max-class pressure. If W31 was “what I run,” W32 is “what showed up that I have not graded yet.”

Three things that mattered

1. Muse Code is the new harness to watch (Aug 5)

Meta shipped Muse Code (beta) with Muse Spark 1.2, a model co-trained with the harness. Persistent async background agents, git worktree fan-out, an append-only event log (replay-exact, restart-safe), bundled skills (/plan, /grill, /goal), and muse resume.

That is a real control surface, not a logo. It is on the Board as a watch row. No letter. I have not run it as a daily driver.

2. SWE-Bench ProMax is the unsaturated refactor yardstick (Aug 10)

arXiv:2608.09802. 170 multilingual refactor tasks, average 11.4 files and 261.6 LOC. Best model in the paper is 41.2% (they cite GPT-5.2). It exists because a SWE-bench Verified test-quality audit found a lot of unsolved instances had bad tests.

If your team is still arguing Sol vs Opus on Verified screenshots, this is the harder number. Saturated benches stop being a routing signal.

3. Qwen3.8-Max is open-weight Max-class pressure (Aug 3)

Alibaba’s 2.4T MoE (95B active), 1M context, pitched at coding and long-horizon cowork. API through Alibaba Cloud Model Studio / QwenCloud. Open weights were promised; Hugging Face Qwen/Qwen3.8-2.4T-A95B was reported Aug 8. The license is bespoke qwen3.8-max, not Apache.

Treat it as pressure on the closed Max lane, not a Board promotion. Ungraded on purpose.

One thing the discourse got wrong

“A new Max-class drop or a new harness means the Board flipped.”

No. Facts can sync without a grade event. I still have not run Muse Code or Qwen3.8-Max as daily tools. Watchlist is not a coronation.

Open weights also do not mean you skip governance. More agents in more places with less procurement friction makes inventory more important, not less.

Board / stack implication

JobReach for (week of Aug 10)
Daily multi-turn engineeringOpus 4.8 (pilot Opus 5 only with eyes open)
Hard coding ceilingSol on Codex / Fable 5
Computer useGemini 3.6 Flash
Cheap BYOK experimentsKimi K3 or DeepSeek V3 in opencode
Editor valueComposer 2.5
New harness to watchMuse Code (ungraded)
Open-weight Max-class pressureQwen3.8-Max (ungraded)

Gemini’s official 3.6 Flash model card is still the receipt behind the computer-use row. Muse Code and Qwen3.8-Max are on the live table as watch rows. Live table: Bingo Board.

One builder action for the week

If you try Muse Code this week, write down the control surface: background agents, worktree fan-out, event log, and who can muse resume a run. That is the harness question, not the model logo.

Watch next

  • Fri Aug 14: mid-month Bingo Board pulse (Grok 4.6 and Gemini 3.7 Flash land after this week’s cutoff)
  • Mon Aug 17: beginner deep dive, what a harness is
  • Big Mama diary Part 2 the following Friday

Kelpie Field Notes Part 2 (AI-BOM) still matters for regulated buyers. It is the closer, not the lead: if the “dependency” can send email or touch PHI, it belongs in the inventory. Open weights and new harnesses make that more true, not less.

Companion vlog: Living in Color: Week of Aug 10 (8 to 12 min).