🔍 Read the full analysis: A Tour Of My September 2026 AI Stack on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A September 29, 2026 practitioner review finds six frontier AI models within roughly 20 index points of each other while cost per task differs by about 100×, shifting model selection from ‘which is smartest’ to ‘which clears a quality bar at lowest cost’. GPT-6.1 Sol launched the same day at $0.39 per task at xhigh, versus $5.98 for top-ranked Claude Opus 5.5.
Six frontier AI models now sit within about 20 index points of each other on general capability, while their cost per task differs by roughly 100×, according to a practitioner analysis published on 29 September 2026 by Thorsten Meyer. The same day brought the release of GPT-6.1 Sol, which scores 51 on the Artificial Analysis Intelligence Index v4.3.x at its xhigh setting while costing an estimated $0.39 per task — compared with $5.98 for the top-scoring model, Claude Opus 5.5 at 58. The result, Meyer argues, is that model selection is no longer about which system is smartest, but which model clears a given quality bar at the lowest cost per task.
The analysis, based on Artificial Analysis Intelligence Index v4.3.x scores and published per-task cost estimates, lays out a field where capability gaps have narrowed sharply. Claude Opus 5.5, released 22 September, holds the highest score at 58 at a cost of $5.98 per task. GPT-6.1 Sol, released 29 September, and GPT-6 Luna, released 22 September, anchor the cheap end: Sol at 51 points for $0.39 per task at xhigh, Luna at 37 points for $0.07. Between them sit Claude Sonnet 5.5 (56 points, $7.60), Claude Fable 5.1 (53 points, $7.63) and GPT-6 Astra (53 points, $3.26).
Three findings stand out in the data. First, Opus 5.5 outscores its more expensive sibling Fable 5.1 by 5 points while costing less per task. Second, Sonnet 5.5 at its max effort setting costs more per task than Opus at max for 2 fewer points, which Meyer says makes it hard to justify at that setting. Third, GPT-6.1 Sol costs roughly one-eighth of Astra and one-twentieth of Fable per task for a score only 1 to 2 points lower.
The analysis also identifies effort settings — not model choice — as the biggest cost lever. On Opus 5.5, moving from xhigh to max adds 2 index points while raising cost per task by 73%; moving from medium to max multiplies cost by 4.46× for 7 points. Sol has caveats: at its high and xhigh settings it takes 57 to 69 seconds to produce a first token, making it unsuitable for interactive use, and Artificial Analysis has not yet published low or max settings for it. Its high setting used 25 million output tokens on the index, against a stated median of 82 million for comparable models — unusually concise output.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Cost Per Task Now Drives Model Choice
The practical takeaway of the analysis is that the frontier has stopped being a leaderboard and become a price curve. When capability differences shrink to a few index points while costs diverge by two orders of magnitude, the economic question dominates: a review pass at $0.32 to $0.39 per task, as Sol offers, is cheap enough to run routinely on every meaningful code change, whereas a $5.98-per-task model cannot be used that way.
Meyer’s working stack reflects this: Opus 5.5 at high effort (54 points, $1.82 per task) as the main builder for features, APIs and multi-file refactors; Opus at xhigh (56 points, $3.46) for architecture, migrations and trust boundaries; and Sol at high or xhigh as a second model family reviewing Opus’s output, on the grounds that a different model family is a better check than a model reviewing itself. Astra, Fable, Sonnet 5.5 and Luna serve scoped roles rather than defaults.
The analysis also warns that cheaper tokens do not mean cheaper work. Meyer’s illustrative example: halving model price saves about 12.5% of real cost, and a single extra minute of human review erases that saving. He notes the example is illustrative, not measured.
Top picks for "tour september stack"
As an affiliate, we earn on qualifying purchases.
A Month of Back-to-Back Frontier Releases
September 2026 saw near-weekly frontier releases, according to the analysis: Claude Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Opus 5.5 and GPT-6 Luna on 22 September, Claude Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September. Sol launched at the same list price as its week-old predecessor GPT-6 Sol — $2 per million input tokens and $10 per million output tokens, versus $4/$20 for Opus 5.5, $10/$50 for Fable and Astra, and $0.10/$0.50 for Luna. Even Sol’s medium setting matches the earlier GPT-6 Sol’s score of 48 at one-fifth of that model’s $1.06 per-task cost, per the analysis. One notable measurement: Sonnet 5.5 at max effort writes about 193,000 output tokens per task, the most Artificial Analysis has measured, driving its cost from $2.74 at xhigh to $7.60 at max.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer, ThorstenMeyerAI.com
Benchmark Limits and Missing Settings
The analysis itself flags its limits. All capability scores come from the Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a map of general capability, not a verdict on any specific workload — he advises shadow-testing before switching models. At the top of the field, one index point is within measurement noise, meaning rankings among Opus, Sol, Astra and Fable at close scores should not be treated as decisive. Artificial Analysis has not yet published low or max settings for GPT-6.1 Sol, so its full cost-quality curve is incomplete. The cost-savings example comparing model price to human review time is labeled by the author as illustrative, not measured, and per-task costs are estimates tied to specific benchmark runs rather than guaranteed real-world prices.
What to Watch Through October
Watch for Artificial Analysis to publish low and max effort settings for GPT-6.1 Sol, which will complete its cost-quality picture and clarify whether it can stretch closer to the top tier. Expect continued pricing pressure: if a $0.39-per-task model sits 1 to 2 points behind $3-to-$8 models, mid-priced models such as Fable 5.1 and Sonnet 5.5 at max effort face the sharpest justification pressure. For practitioners, the next step the author recommends is shadow-testing candidate models against real workloads before any switch, and treating cross-family review passes — now affordable at Sol’s price point — as a routine part of the development loop.
Key Questions
What is GPT-6.1 Sol and when was it released?
GPT-6.1 Sol is a frontier model released on 29 September 2026, priced at $2 per million input tokens and $10 per million output tokens — the same as its week-old predecessor, GPT-6 Sol. It scores 51 on the Artificial Analysis Intelligence Index v4.3.x at its xhigh setting, at an estimated $0.39 per task.
Which model scores highest as of late September 2026?
According to the analysis, Claude Opus 5.5 holds the highest score at 58 points on the Artificial Analysis Intelligence Index v4.3.x, at an estimated $5.98 per task at its max setting. Its xhigh setting scores 56 for $3.46 per task.
Why does effort setting matter more than model choice for cost?
On Opus 5.5, going from xhigh to max adds 2 index points while raising cost per task by 73%, and moving from medium to max multiplies cost by 4.46× for 7 points. The analysis concludes that the effort dial often moves the bill more than the choice between models.
Is GPT-6.1 Sol suitable for interactive use?
Not at its higher settings, according to the analysis. At high and xhigh, Sol takes 57 to 69 seconds to produce a first token, making it better suited to batch review and deep-dive work than real-time interaction.
Are these benchmark scores a guarantee of real-world performance?
No. All scores come from the Artificial Analysis Intelligence Index v4.3.x, which the author describes as a map of general capability rather than a verdict on any specific workload. He recommends shadow-testing models against your own tasks before switching, and notes that one index point is within measurement noise.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
