🔍 Read the full analysis: September 2026 AI Stack: My Build, Research, And Decision Flow on ThorstenMeyerAI.com
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A September 29, 2026, assessment of six AI models argues that task costs now vary far more than benchmark scores. Its author uses Opus 5.5 for building and GPT-6.1 Sol for details and review, while treating other models as task-specific alternatives. The rankings and cost figures come from Artificial Analysis and may not predict performance on an individual workload.
Thorsten Meyer’s September 29 assessment of six AI models says their benchmark scores are relatively close while their reported costs per task differ sharply, and sets out a workflow built around Claude Opus 5.5 for development and GPT-6.1 Sol for detailed review. The comparison matters to teams deciding whether higher model settings justify their added expense, though the figures reflect an index benchmark rather than every user’s workload.
Meyer bases the comparison primarily on the Artificial Analysis Intelligence Index v4.3.x. In its top settings, the cited scores range from 58 for Opus 5.5 to 37 for GPT-6 Luna. Reported cost per task ranges from $0.07 for Luna to $7.63 for Claude Fable 5.1. Meyer says GPT-6.1 Sol at xhigh scores 51 at $0.39 per task, while GPT-6 Astra scores 53 at max for $3.26 and Fable 5.1 scores 53 for $7.63.
The proposed allocation follows those figures. Meyer uses Opus 5.5 at high for routine development, citing a score of 54 and $1.82 per task, and at xhigh for harder work such as architecture and migrations, at 56 and $3.46. GPT-6.1 Sol at high or xhigh handles focused investigations and independent review. Astra or Fable serve as alternatives when results from Sol and Opus disagree; Sonnet 5.5 and Luna are assigned scoped subtasks and routine checks.
The report also compares effort settings, which affect both score and cost. For Opus 5.5, moving from medium to max raises the cited score from 51 to 58 while increasing cost per task from $1.34 to $5.98. For Sonnet 5.5, max costs $7.60 for a score of 56, compared with $2.74 and 52 at xhigh. Meyer consequently favors medium for everyday work and documents, and high or xhigh for development.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
How Effort Changes Model Costs
The comparison shifts purchasing decisions from choosing a single top-ranked model toward matching a model and effort setting to each task. If the reported estimates hold for a team’s own work, a lower-cost review pass could make it practical to check more changes with a second model family. That may help catch errors while keeping inference costs below the price of running a high-effort builder for every step.
Meyer also cautions that model cost is only one part of the expense. He writes that halving the model price saves 12.5% of the real cost in an illustrative example, and that an extra minute of human review can erase that saving. That statement is an example, not a measured result. His process calls for passing the failing case and evidence back to the builder when review finds a problem, and warns that passing tests alone does not authorize shipping.
From Rankings to Task Budgets
The report describes a change in how Meyer evaluates models: rather than treating a leaderboard position as a complete buying guide, it compares index performance with estimated cost per task. Its six-model table lists Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna. Release dates in the table range from September 1 for Fable 5.1 to September 29 for GPT-6.1 Sol.
Meyer says the Artificial Analysis index is a general capability measure, not a verdict on a specific workload, and recommends shadow-testing before switching systems. The report also notes that one index point may fall within measurement noise. Those caveats limit what can be concluded from small score differences, especially when comparing models that have not been tested on a team’s own tasks.
“The question from “which model is smartest?” to “which model clears my quality bar at the lowest cost per task?””
— Thorsten Meyer
What the Benchmark Cannot Settle
The cited figures do not establish how the models will perform on a particular organization’s code, documents or review standards. Meyer calls the index a general capability map and says teams should shadow-test before switching. He also says a one-point score difference is within the noise, so the small gaps between Sol, Astra and Fable do not by themselves establish a meaningful quality advantage.
The report says high and xhigh settings for GPT-6.1 Sol take 57 to 69 seconds to produce a first token in the index, which may matter for interactive use. It does not provide a complete comparison of latency or quality across real-world tasks. The source material also ends partway through an illustrative calculation about model prices and human review, leaving that example’s full assumptions and conclusion unavailable.
Test the Stack on Real Work
Meyer’s stated next step for anyone considering a switch is to shadow-test candidate models on their own workload before changing defaults. Such tests can reveal whether lower reported cost per task preserves the quality a team needs, and whether slower first-token times affect its workflow. The report does not announce a formal future benchmark date or a planned update.
Key Questions
Which model does Meyer use for development?
He uses Claude Opus 5.5 at high for regular development and xhigh for harder tasks such as architecture and migrations.
Why does he use GPT-6.1 Sol for review?
Meyer says its reported cost of $0.32 to $0.39 per task makes focused investigations and a second-model review affordable to run routinely. The report does not establish that it will be the best reviewer for every team.
Are the index scores proof that one model is better for every task?
No. The report describes the Artificial Analysis index as a general capability measure and recommends shadow-testing on the workload a team actually runs.
What remains unknown about the cost comparison?
The source does not show how the reported per-task costs translate to each team’s use, or provide complete real-world comparisons across tasks. It also presents its human-review cost example as illustrative rather than measured.
Source: ThorstenMeyerAI.com
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
