Agent evaluation

Thirteen tasks, 11 models

Each model got the same thirteen briefs as independent tasks with a fresh context. Every task is scored out of 10.

Results

Scores come from the latest blind pass, which graded all eleven models together (2026-10-04, after MiniMax M3.1 Flash and MiniMax M3 were added). Grok 4.7's task 13 stopped when its usage balance ran out and awaits a re-run, so Grok is compared on its average per task over 12 tasks.

Avg / task Code Writing Ops Frontend Time Cost / usage
Sonnet 5.5 (extra) 9.31 9.50 9.25 9.50 9.00 ~28 min $15
Fable 5.1 (extra) 8.96 9.44 8.62 9.00 8.75 ~49 min $53
Opus 5.5 (high) 8.96 9.31 9.31 8.25 8.50 ~21 min $14
Grok 4.7 (xhigh), 12 tasks 8.67 9.12 7.69 9.38 9.00 ~80 min $6.74
GPT-6 Astra (xhigh) 8.60 9.19 8.69 9.12 7.33 ~35 min 3.6M in / 185k out tokens
GPT-6.1 Sol (xhigh) 8.52 9.00 8.50 8.88 7.67 ~48 min 5.8M in / 243k out tokens
muse (max) 7.21 6.88 7.25 8.12 7.00 ~20 min 6.2M in / 162k out tokens
GPT-6 Luna (max) 7.10 7.00 7.50 7.50 6.42 ~34 min 3.5M in / 220k out tokens
mimo 6.92 7.50 7.06 6.75 6.08 not recorded not recorded
MiniMax M3.1 Flash (xhigh, mcode) 6.81 7.56 5.94 6.38 7.25 ~42 min 0.58M in + 18.6M cache-read / 398k out tokens; free preview
MiniMax M3 (default, mcode) 5.52 5.62 5.19 5.50 5.83 ~9 min 0.55M in + 4.8M cache-read / 190k out tokens; 2% of a 5-hour window

Totals (of 130): Sonnet 121.0, Fable 116.5, Opus 116.5, Astra 111.75, Sol 110.75, muse 93.75, Luna 92.25, mimo 90.0, MiniMax M3.1 Flash 88.5, MiniMax M3 71.75; Grok 104.0 of 120.

What was observed

What is inferred

Recommendations by work type

Caveats

1

Sonnet 5.5

claude-sonnet-5-5

9.31 / 10 per task

total 121 of 130 (13 tasks)

Effort
extra
Time
~28 min
Cost
$15
Where
Claude Code cloud, account A
Code 9.50Writing 9.25Operations 9.50Frontend 9.00

One subagent per task, four at a time.

2=

Fable 5.1

claude-fable-5-1

8.96 / 10 per task

total 116.5 of 130 (13 tasks)

Effort
extra
Time
~49 min
Cost
$53
Where
Claude Code cloud, account B
Code 9.44Writing 8.63Operations 9.00Frontend 8.75

One subagent per task, four at a time.

2=

Opus 5.5

claude-opus-5-5

8.96 / 10 per task

total 116.5 of 130 (13 tasks)

Effort
high
Time
~21 min
Cost
$14
Where
Claude Code cloud, account A
Code 9.31Writing 9.31Operations 8.25Frontend 8.50

One subagent per task, four at a time. Ran on high, not extra: not like-for-like.

4

Grok 4.7

grok-4.7 (served as grok-4.7-build)

8.67 / 10 per task

total 104 of 120 (12 tasks)

Effort
xhigh
Time
~80 min for 12 tasks
Cost
$6.74 for 12 tasks (+$1.43 on the interrupted task 13)
Usage
2.00M input + 25.6M cache-read, 1.11M output (891k reasoning) tokens; 321 model turns
Where
Grok CLI on the Mac, grok.com account
Code 9.13Writing 7.69Operations 9.38Frontend 9.00

One grok -p per task, four at a time, with a clean HOME so no personal instructions or skills loaded; memory off, MCP denied, web tools off, workspace sandbox. Task 13 stopped when the Grok Build balance ran out and awaits a re-run; totals cover 12 tasks.

5

GPT-6 Astra

gpt-6-astra

8.60 / 10 per task

total 111.75 of 130 (13 tasks)

Effort
xhigh
Time
~35 min
Cost
tokens only
Usage
3.57M input (3.09M cached), 185k output, 42k reasoning tokens; tasks 01-09 used a full 5-hour window and 16% of the weekly quota on gmail-business
Where
Codex CLI on the Mac, gmail-business (01-09) and gmail-personal (10-13)
Code 9.19Writing 8.69Operations 9.13Frontend 7.33

One codex exec per task, four at a time; the accounts' personal AGENTS.md set aside. Tasks 10-13 were cut off when gmail-business ran out of credits and re-run from the template on gmail-personal.

6

GPT-6.1 Sol

gpt-6.1-sol

8.52 / 10 per task

total 110.75 of 130 (13 tasks)

Effort
xhigh
Time
~48 min
Cost
tokens only
Usage
5.85M input (5.28M cached), 243k output, 75k reasoning tokens; 48% of the account's 5-hour quota and 7% of its weekly quota
Where
Codex CLI on the Mac, gand-business account
Code 9.00Writing 8.50Operations 8.88Frontend 7.67

One codex exec per task, four at a time. The account's personal AGENTS.md was set aside for the run. An aborted duplicate launch preceded this clean run.

7

muse

muse-spark-1.3

7.21 / 10 per task

total 93.75 of 130 (13 tasks)

Effort
max
Time
~20 min
Cost
tokens only
Usage
6.20M input (5.54M cached), 162k output, 59.6k reasoning tokens; 172 model steps
Where
tincan on the Mac, 2026-09-30
Code 6.88Writing 7.25Operations 8.13Frontend 7.00

Local run through tincan: one stateless request per task, four at a time. Cost can't be measured, so tokens are recorded instead.

8

GPT-6 Luna

gpt-6-luna

7.10 / 10 per task

total 92.25 of 130 (13 tasks)

Effort
max
Time
~34 min
Cost
tokens only
Usage
3.48M input (3.05M cached), 220k output, 124k reasoning tokens; about 6% of the account's 5-hour quota and 1% of its weekly quota
Where
Codex CLI on the Mac, gand-business account
Code 7.00Writing 7.50Operations 7.50Frontend 6.42

One codex exec per task, four at a time. The account's personal AGENTS.md was set aside for the run. Task 12 left an earlier draft of its files at the repository root.

9

mimo

xiaomi/mimo-v2.6-pro

6.92 / 10 per task

total 90 of 130 (13 tasks)

Effort
default
Time
not recorded
Cost
not recorded
Where
tincan on cachyos, 2026-09-22/23
Code 7.50Writing 7.06Operations 6.75Frontend 6.08

Earlier run, one request per task (mimo run --pure). Timing and cost were not recorded.

10

MiniMax M3.1 Flash

minimax/MiniMax-M3.1-Flash-Preview

6.81 / 10 per task

total 88.5 of 130 (13 tasks)

Effort
xhigh
Time
~42 min
Cost
tokens only
Usage
582k input + 18.61M cache-read input, 398k output tokens; the Token Plan's coding-plan counters stayed at 0% (they don't track mcode's managed usage)
Where
mcode (MiniMax Code 0.6.2) on the Mac, 2026-10-04
Code 7.56Writing 5.94Operations 6.38Frontend 7.25

One mcode exec per task, four at a time, in the user's normal HOME (mcode refuses a clean one); no personal instruction files reached it, user skills did. Task 11's first attempt was killed ~15 min in by the orchestrator's job limit and re-run alone from the template; bench time = 28m38s first pass + 13m38s re-run.

11

MiniMax M3

minimax/MiniMax-M3

5.52 / 10 per task

total 71.75 of 130 (13 tasks)

Effort
default (not selectable)
Time
~9 min
Cost
tokens only
Usage
547k input + 4.78M cache-read input, 190k output tokens; 2% of the Token Plan's 5-hour window, weekly unchanged (1%)
Where
mcode (MiniMax Code 0.6.2) on the Mac, 2026-10-04
Code 5.63Writing 5.19Operations 5.50Frontend 5.83

One mcode exec per task, four at a time, same protocol as MiniMax M3.1 Flash; mcode rejects --effort for M3, so it ran at its fixed default. Usage was polled every 2 minutes with a hold at 85%, never reached.

Scores by task

TaskSonnet 5.5Fable 5.1Opus 5.5Grok 4.7GPT-6 AstraGPT-6.1 SolmuseGPT-6 LunamimoMiniMax M3.1 FlashMiniMax M3
01 Feature building9.75109.59.7598.757.7586.758.54.75
02 Bug fixing109.751010101098.759.58.258.75
03 Infra management98.58.2589.258.55.755.576.753.5
04 Project maintenance9.259.59.58.758.58.7555.756.756.755.5
05 Technical writing8.757.759.57.758.58.7567.756.54.254
06 Product planning9.59.259.258.58.758.58.58.255.758.756.5
07 Creative writing9.58.5988.758.256.7567.254.55.5
08 Social media9.2599.56.58.758.57.7588.756.254.75
09 Incident analysis9.258.258.59.2598.757.256.7566.755.25
10 General operations9.759.7589.59.25998.257.565.75
11 Frontend operations98.57.58.758.2586.566.757.56.25
12 Frontend editorial8.759.2599.256.75777.56.756.756.25
13 Frontend playful9.258.59–787.55.754.757.55
Code average9.509.449.319.139.199.006.887.007.507.565.63
Writing average9.258.639.317.698.698.507.257.507.065.945.19
Operations average9.509.008.259.389.138.888.137.506.756.385.50
Frontend average9.008.758.509.007.337.677.006.426.087.255.83
Average per task9.318.968.968.678.608.527.217.106.926.815.52
Total (tasks graded)121 (13)116.5 (13)116.5 (13)104 (12)111.75 (13)110.75 (13)93.75 (13)92.25 (13)90 (13)88.5 (13)71.75 (13)

Highlighted cell: the best score on that task (ties all highlighted).

How this was scored

13 of 13 tasks graded.