Blog8 min read

Claude Opus 5.5 or GPT-6 Sol? 7 Facts for Choosing by the Job

Anthropic and OpenAI released Claude Opus 5.5 and GPT-6 Sol on the same day, September 22, 2026. 7 side-by-side facts for non-technical teams: where each one is available, Zapier's shared test (Opus 5.5 finished more tasks, GPT-6 Sol cost less per task), and what each company claims for writing, files and facts. No overall winner, with every source.

MB
Michael Bennett · AI marketing systems
Claude Opus 5.5 vs GPT-6 Sol, HEAD TO HEAD

Anthropic and OpenAI released their new work models on the same day, September 22, 2026. One outside test scored both the same way, and it did not crown either one. Opus 5.5 finished more of the tasks. GPT-6 Sol cost less per task.

What each one is, and who gets it. Opus 5.5 is Anthropic's newest model for the Claude app. It is the default model on the Pro, Max and Team plans, and it is available on Enterprise. GPT-6 Sol is the middle model in OpenAI's new GPT-6 family. On ChatGPT's Plus, Pro, Business, Enterprise and Edu plans it runs in Work mode, ChatGPT's mode for longer tasks, and in Codex, OpenAI's tool for writing code. OpenAI says the new models are “not yet available in Chat”, the regular chat mode most people use. Neither model is listed for free plans.

Below are 7 facts from both launches and from Zapier's test, in the same order as our carousel. Every claim is attributed to the company that made it. Where the reasoning is ours, it says "Our read".

Carousel slide: same launch day, same office test. Opus 5.5 finished more. GPT-6 Sol cost less per task. Neither won everything. Zapier's test, each model at its best setting.
Zapier AutomationBench 1.0.6, each model at its best setting.
Claude Opus 5.5 vs GPT-6 Sol, HEAD TO HEAD

1. Where you can use each one today

Date and source: September 22, 2026. OpenAI's launch post; Anthropic's product page, and Cat Wu of Anthropic on X.

Side by side. Anthropic's Cat Wu: “Claude Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans.” Anthropic's product page adds Enterprise. OpenAI: “GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users.” ChatGPT's Free plan gets GPT-6 Luna, OpenAI's lower-cost model, in the desktop app only.

What it means for your work. Our read: on Claude Pro, Max or Team, Opus 5.5 is already the default when you open the app. On a paid ChatGPT plan, Sol is in Work mode and Codex, not in the chat you may already use every day.

Try this. Check which plan your team pays for.

2. Neither is its company's top model

Date and source: September 2026. OpenAI's launch post and models page; Anthropic's launch post and model overview.

Side by side. OpenAI: “GPT‑6 Astra continues to be our best model across the board.” Sol is the middle tier of 3. Anthropic prices Fable 5.1 at $10 / $50 per million tokens, above Opus 5.5 at $4 / $20, and says Opus 5.5 “performs at the level of Claude Fable 5.1 on most work”. Anthropic also cautions that “at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences.”

What it means for your work. Our read: this is a fair fight between 2 everyday work models, not the flagships.

Try this. Compare them with each other, not with the top models you read about.

3. Opus 5.5 finished more of the tasks

Date and source: Zapier's AutomationBench 1.0.6 leaderboard, read September 23, 2026.

Side by side. Zapier's test gives each model real business errands across apps like email, CRM (the customer database a sales team works in) and spreadsheets. At each model's best setting, Opus 5.5 finished 40.0% of tasks and GPT-6 Sol 33.2%. Each company published its own model's score, and Zapier's leaderboard lists both. Anthropic's footnote says its result comes “from Zapier’s own evaluation during early access”, and that Opus 5.5 ran “without fallback models, so safeguard interventions were considered failures”, which Anthropic says lowers its score. Neither launch post compares against the other's new model: on this test, Anthropic's post lists GPT-5.6 Sol, OpenAI's older model, and GPT-6 Astra, not GPT-6 Sol.

What it means for your work. Both still fail most of these tasks. 40% means 6 in 10 errands were not completed. Our read: finishing more is not the same as finishing reliably.

Try this. Give both your 3 most common errands, and check each result before anything goes out.

Carousel slide: on Zapier's test of business errands across email, CRM and spreadsheets, at each model's best setting, Opus 5.5 finished 40.0% of tasks and GPT-6 Sol 33.2%. Both still fail most of these tasks.
Zapier AutomationBench 1.0.6. No overall winner: see fact 4.

4. GPT-6 Sol cost less per task

Date and source: Zapier's AutomationBench 1.0.6 leaderboard, read September 23, 2026.

Side by side. On the same test, GPT-6 Sol cost about $0.27 per task at its best setting, and Opus 5.5 about $1.28. At equal scores, the gap narrows: when both finished 32.0% of tasks, Sol cost $0.34 per task and Opus 5.5 $0.65, about half.

What it means for your work. This is API cost: what a company pays to run a model inside its own software, billed by use. It is not your subscription price, and it says nothing about what ChatGPT Plus or Claude Pro costs. Our read: it matters most to teams sending many tasks through their own tools.

Try this. If you run tasks at volume, price them on your own work first.

5. Writing: both claim it, no shared test

Date and source: September 22, 2026. Both launch posts.

Side by side. Anthropic on Opus 5.5: “It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it.” OpenAI on Sol: “Expect to see more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance.” OpenAI adds that it thinks this “will be especially noticeable in technical and coding conversations.” No test we found scored both on writing the same way.

What it means for your work. Our read: this one is yours to judge, with your own style rules.

Try this. Give both the same non-confidential email thread and compare the summaries.

6. Files you send: only Opus 5.5 has a claim

Date and source: Anthropic's prompting guide for Opus 5.5, read September 23, 2026; September 3, 2026, ChatGPT release notes.

Side by side. Anthropic: “The spreadsheets, slides, and documents it produces need less editing before you share them.” OpenAI's claim about files is for GPT-6 Astra, not for Sol: “Astra can create documents, spreadsheets, and presentations that follow your templates and instructions, and adapt when you add requirements or change direction.” OpenAI makes no such claim for Sol.

What it means for your work. Our read: for decks and spreadsheets, Opus 5.5 is the one pitched for the job. That is a claim, not a test. A missing claim is also not evidence that Sol does worse.

Try this. Ask both for the same 5-slide deck and count the fixes each one needs.

Carousel slide: Anthropic says Opus 5.5's files need less editing before you share them. OpenAI's claim about files is for GPT-6 Astra, not Sol. A claim, not a test.
A vendor claim on one side, no claim on the other. Not a test result.

7. Facts: 2 claims, 2 different tests

Date and source: September 22, 2026. Both launch posts.

Side by side. Anthropic ran a research test in which an automated grader checked every figure and quote: “Across different effort settings, 16 out of 18 of Opus 5.5’s reports cleared our quality bar, where any invented figure or quote would have failed.” OpenAI ran its own test on conversations where users had flagged mistakes, and says Sol “makes about half as many mistakes as its predecessor”, the older GPT-5.6 Sol. OpenAI adds: “These error-inducing conversations are not representative of typical usage, where factual errors are more rare.”

What it means for your work. Different tests with different baselines, so they cannot be compared. Neither company claims zero mistakes. In Anthropic's own test, 2 of 18 reports failed.

Try this. Whichever you use, check every number before it leaves your hands.

Pick by the job

Our read, not either company's:

  1. Building files you will send: Opus 5.5 is the one pitched for that job.
  2. Running tasks at volume on a budget: on Zapier's test, Sol cost about half as much at the same score.
  3. Writing and research: try both on your own work, and check every number.

Same launch day. Both fail most of Zapier's tasks. Test them on your own work this week, on a task you would otherwise do by hand, and keep the one that needs fewer fixes.

What neither company said

That one model won overall. Only 1 outside test of office work scored both the same way, and each model won something different on it. Neither launch post names the other's new model. OpenAI compared Sol with Claude Opus 5 and Fable 5 / 5.1. Anthropic compared Opus 5.5 with GPT-5.6 Sol and GPT-6 Astra.

That there is a shared writing test. Both claim clearer writing. Neither tested the other.

That their internal tests can be compared. Anthropic's 16 of 18 and OpenAI's "half as many mistakes" measure different things against different baselines. Neither is independent.

That either is its company's top model. OpenAI says Astra is its best. Anthropic prices Fable 5.1 above Opus 5.5.


Sources: verified September 23, 2026 against OpenAI's launch post of September 22, 2026, "Introducing GPT‑6 Sol and Luna"; OpenAI's API changelog and models page; OpenAI's ChatGPT release notes of September 3, 2026; Anthropic's launch post of September 22, 2026, its Opus product page, its prompting guide and model overview for Opus 5.5; Cat Wu's post on X; and Zapier's AutomationBench 1.0.6 leaderboard. Cost per task is at API list prices, not subscription prices. Quotations are exact. Lines marked Our read are MichaelBennett.co's reasoning, not either company's. Every source, with its URL, is listed in SOURCES.md.

MB
Michael Bennett
I build AI marketing systems that acquire, convert & retain customers.

Working out where AI actually fits in your marketing?

I write these while building the systems behind them: measurement, creative pipelines, and agents that do real work. Connect on LinkedIn and tell me what you are working on. That is where these conversations start.

Connect on LinkedIn