Living in Color August · Episode 02
Living in Color W32: What Actually Moved
Week of Aug 10: Muse Code as the new harness, SWE-Bench ProMax as the unsaturated refactor yardstick, and Qwen3.8-Max as open-weight Max-class pressure.
Listen in my voice · AI narration (ElevenLabs clone)
On this page
- Three things that mattered
- 1. Muse Code is the new harness to watch (Aug 5)
- 2. SWE-Bench ProMax is the unsaturated refactor yardstick (Aug 10)
- 3. Qwen3.8-Max is open-weight Max-class pressure (Aug 3)
- One thing the discourse got wrong
- Board / stack implication
- One builder action for the week
- Watch next
Last week was routing tables. This week is a new harness, a harder bench, and open-weight Max-class pressure. If W31 was “what I run,” W32 is “what showed up that I have not graded yet.”
Three things that mattered
1. Muse Code is the new harness to watch (Aug 5)
Meta shipped Muse Code (beta) with Muse Spark 1.2, a model co-trained
with the harness. Persistent async background agents, git worktree fan-out, an
append-only event log (replay-exact, restart-safe), bundled skills (/plan,
/grill, /goal), and muse resume.
That is a real control surface, not a logo. It is on the Board as a watch row. No letter. I have not run it as a daily driver.
2. SWE-Bench ProMax is the unsaturated refactor yardstick (Aug 10)
arXiv:2608.09802. 170 multilingual refactor tasks, average 11.4 files and 261.6 LOC. Best model in the paper is 41.2% (they cite GPT-5.2). It exists because a SWE-bench Verified test-quality audit found a lot of unsolved instances had bad tests.
If your team is still arguing Sol vs Opus on Verified screenshots, this is the harder number. Saturated benches stop being a routing signal.
3. Qwen3.8-Max is open-weight Max-class pressure (Aug 3)
Alibaba’s 2.4T MoE (95B active), 1M context, pitched at coding and
long-horizon cowork. API through Alibaba Cloud Model Studio / QwenCloud. Open
weights were promised; Hugging Face Qwen/Qwen3.8-2.4T-A95B was reported Aug 8.
The license is bespoke qwen3.8-max, not Apache.
Treat it as pressure on the closed Max lane, not a Board promotion. Ungraded on purpose.
One thing the discourse got wrong
“A new Max-class drop or a new harness means the Board flipped.”
No. Facts can sync without a grade event. I still have not run Muse Code or Qwen3.8-Max as daily tools. Watchlist is not a coronation.
Open weights also do not mean you skip governance. More agents in more places with less procurement friction makes inventory more important, not less.
Board / stack implication
| Job | Reach for (week of Aug 10) |
|---|---|
| Daily multi-turn engineering | Opus 4.8 (pilot Opus 5 only with eyes open) |
| Hard coding ceiling | Sol on Codex / Fable 5 |
| Computer use | Gemini 3.6 Flash |
| Cheap BYOK experiments | Kimi K3 or DeepSeek V3 in opencode |
| Editor value | Composer 2.5 |
| New harness to watch | Muse Code (ungraded) |
| Open-weight Max-class pressure | Qwen3.8-Max (ungraded) |
Gemini’s official 3.6 Flash model card is still the receipt behind the computer-use row. Muse Code and Qwen3.8-Max are on the live table as watch rows. Live table: Bingo Board.
One builder action for the week
If you try Muse Code this week, write down the control surface: background
agents, worktree fan-out, event log, and who can muse resume a run. That is
the harness question, not the model logo.
Watch next
- Fri Aug 14: mid-month Bingo Board pulse (Grok 4.6 and Gemini 3.7 Flash land after this week’s cutoff)
- Mon Aug 17: beginner deep dive, what a harness is
- Big Mama diary Part 2 the following Friday
Kelpie Field Notes Part 2 (AI-BOM) still matters for regulated buyers. It is the closer, not the lead: if the “dependency” can send email or touch PHI, it belongs in the inventory. Open weights and new harnesses make that more true, not less.
Companion vlog: Living in Color: Week of Aug 10 (8 to 12 min).