Blog11 min read

ChatGPT Images 2.5: Same Price, Half The Wait, And One Number Everyone Will Misread

ChatGPT Images 2.5 emits about a quarter of the tokens of the model it replaces, so high quality now costs what medium used to and arrives six times faster. Measured against the live API, with the images it made.

MB
Michael Bennett · AI marketing systems
A woman photographs a ceramic mug on a sheet of white paper at her kitchen table, using a phone clamped to a small tripod and a clip-on lamp, in daylight from the window behind her.

OpenAI shipped ChatGPT Images 2.5 today. The coverage so far says sharper, faster, better at editing, and all of that is true. It is also the least interesting thing about the release.

Two facts are sitting in OpenAI's own documents that nobody has picked up, and both change what you should actually do about this.

The first is in the pricing table. The second is in a footnote in the system card.

ChatGPT Images 2.5, IMAGE GENERATION

What actually shipped

The model is better at working from your reference photos, so a person or a product stays recognizable when you move it into a new setting. It edits only what you asked it to edit and leaves the rest alone. It holds its work together across a long back-and-forth instead of degrading a little with every turn. It handles transparent backgrounds and more complex layouts.

Generation latency is down up to 50% against Images 2.0. Note the hedge: "up to" is doing work in that sentence.

Four product features came with it, and these are the ones you can use this afternoon:

  • Sketch. Draw your idea directly in ChatGPT and use it as the layout guide. Type @Sketch. Rough is fine; the point is to specify composition without writing a paragraph describing it.
  • Templates. Named starting points for common formats, "Poster" and "Merch" among them, instead of a blank prompt box.
  • Comments on images. Point at the part you want changed rather than describing its location in words.
  • Shareable prompts. Share an image with the prompt attached, so someone else can run it with their own photos.

It is live for all ChatGPT, ChatGPT Work, and Codex users, on every tier, across desktop, mobile and web. No staged rollout by plan, which is not how most of these land.

Developers get two API models: GPT-Image-2.5 Flare, the default, and GPT-Image-2.5 Sunburst, built for precision at the cost of longer generation times.

A picture costs about a penny

Image pricing is published as dollars per million tokens, which is why almost nobody quotes a figure you can use. So I generated images against the live API and read the cost off each response. These are measured, not estimated:

QualityPer image100 imagesTime
Low0.62 cents$0.628-11s
Medium1.35 cents$1.358-11s
High5.30 cents$5.3016-21s

A standard 1024x1024 image. Landscape at high quality is cheaper still, 4.14 cents.

These figures are exact rather than sampled: output token counts turned out to be deterministic for a given model, size and quality. Two completely different prompts returned identical counts, 196, 439 and 1756 tokens, so the same request always costs the same.

The published table would have misled you, and it misled me

OpenAI publishes a per-image cost table, and it does not cover the 2.5 models. It covers GPT Image 2, at 0.6, 5.3 and 21.1 cents. Since 2.5 bills at identical per-token rates, the reasonable inference is that per-image costs match.

That inference is wrong, and I published it before I checked. 2.5 emits roughly a quarter of the output tokens for the same quality setting, so the same nominal quality costs about a quarter as much:

Same prompt, same sizeGPT Image 2ChatGPT Images 2.5
Medium5.30 cents, 40s1.35 cents, 8-11s
High21.10 cents, 121s5.30 cents, 16-21s

High quality on 2.5 costs exactly what medium cost on GPT Image 2, and arrives six times faster.

The same prompt measured on both models at 1024x1024. GPT Image 2: medium 5.30 cents in 40 seconds, high 21.10 cents in 121 seconds. ChatGPT Images 2.5: medium 1.35 cents in 8 to 11 seconds, high 5.30 cents in 16 to 21 seconds.
High quality on 2.5 costs what medium cost on the model it replaces.

Identical per-token pricing does not imply identical per-image cost, which is the trap in reading a rate card instead of a bill.

What that means for your budget

A hundred product images at high quality is $5.30. A thousand is $53. Drafts are two thirds of a cent, so experimenting is effectively free.

The one comparison worth running before you migrate: gpt-image-1-mini measured 3.36 cents at high and 0.87 at medium. Still the cheapest option, but the gap has narrowed enough that it is no longer the obvious choice it looks like on the published tables, and 2.5 at medium now undercuts Mini at high.

What it actually made

Everything below was generated for this article, against the live API, at the prices shown.

Two photographs of the same plain kraft coffee bag with a white cup and loose beans. On the left it stands on a pale marble kitchen counter; on the right the identical bag and cup are on a rustic wooden cafe table with a blurred cafe interior behind. The bag's shape, folds and color are unchanged.
One good product photograph, moved anywhere.

You have one good photograph and need four: on white for the store, on a table for the ad, in a lifestyle shot for social, seasonal for the campaign. That was a reshoot, a designer, or an afternoon in Photoshop.

Four versions of the same coffee bag photograph in different settings: a plain white store listing background, a festive table with pine and fairy lights, an outdoor summer picnic table in dappled sun, and a modern office kitchen counter.
Four seasonal variants from one source photograph, six cents each.

The second thing that changed is text. Words inside a generated image used to come out as plausible-looking nonsense. Both sets of copy below were typed into the prompt and came back correct, correctly spelled and sensibly laid out.

Two generated graphics. A portrait sale poster reading SPRING SALE, 20% OFF ALL SHIRTS, THIS WEEK ONLY, with a flat illustration of two folded shirts on a cream background. A landscape three-step diagram with circular icons labeled COLLECT, REVIEW and SEND joined by arrows.
Promo graphics and deck diagrams stop being things you brief.

You can also skip the description entirely and draw. In ChatGPT this is Sketch: type @Sketch and draw the layout, however roughly.

On the left, a crude line drawing: a rectangle with a folded top, a circle with a handle, a horizontal line and a few small ovals. On the right, a photorealistic product photograph following that exact layout, with a kraft coffee bag, a white cup of coffee, loose beans and a marble counter edge.
A box and a circle specify a composition faster than a paragraph.

And it is not locked to one house style. The same subject and the same sentence, with only the style instruction changed:

The same road bicycle rendered eight ways: a photograph against a stone wall, a flat vector illustration in navy and orange, a pastel isometric render, a mid-century screenprint poster, a loose watercolor, a soft clay 3D render, a single-weight line drawing, and a white-on-blue technical blueprint with exploded components.
Eleven cents for all eight.

The per-token table, and what it tells you

ModelImage inCachedOutput
gpt-image-2.5-sunburst$8.00$2.00$30.00
gpt-image-2.5-flare$8.00$2.00$30.00
gpt-image-2$8.00$2.00$30.00
gpt-image-1-mini$2.50$0.25$8.00

Three things follow, and all three are arithmetic rather than opinion.

The new model costs exactly what the old one costs. Identical on every line. A meaningful quality jump and up to half the latency, at no price increase. That is not the usual shape of a frontier model release, and if you are already building on GPT-Image-2 the migration question answers itself.

The premium model is not a premium tier. Flare and Sunburst are priced the same, to the cent. OpenAI positions Sunburst for "premium visual workflows," which reads like it should cost more, and it does not. So the choice between them is purely latency against precision. If you are running a campaign asset through six rounds of edits, there is no budget argument for staying on the fast one.

There is still no 2.5 mini. gpt-image-1-mini outputs at $8.00 against $30.00 per million tokens. Measured per image it is 3.36 cents at high and 0.87 at medium, so it remains the cheapest option. But the gap is far narrower than the per-token rates suggest, and 2.5 at medium (1.35 cents) now undercuts Mini at high (3.36 cents). Price them against each other before assuming the old model wins on cost.

Two cautions. Pricing is per token, not per image, so cost per image depends on size and quality settings; measure your own workload rather than trusting a quoted figure, including mine. And no batch price is published for either 2.5 model, while GPT-Image-2 has one at half the standard rate. If your pipeline depends on the batch discount, check before you migrate.

About that speed number

OpenAI says up to 50% lower latency. Fifty percent lower latency is two times faster.

Manus, a launch partner quoted on OpenAI's own page, says Flare delivers "high-quality images at two to four times the speed of GPT-Image-2." That top end is double what OpenAI claims for its own model.

It may well be right. But it is a partner evaluation printed in a launch announcement, not an independent benchmark, and no third-party numbers exist yet. Plan against OpenAI's figure. Treat anything above it as upside.

The number everyone will misread

OpenAI published a system card, and it contains a safety table that looks like a clean win.

Against an adversarial prompt set, the share of violative images that slipped through and reached the user:

Unsafe shown
GPT-Image-2.5-Sunburst1.09%
GPT-Image-2.5-Flare1.41%
ChatGPT Images 2.01.64%

Both new models beat the old one. You will see that framed as a safety improvement this week.

Now read the footnote. OpenAI marks statistically significant results with an asterisk, and then states, in its own words:

"No unsafe-shown difference meets this threshold."

The headline safety improvement is not statistically significant, and OpenAI is the one telling you so. Nobody is hiding it. It is simply in a footnote under a table, which is where things go when they are true and inconvenient.

Break it down by category and it gets more interesting. Two genuine wins: extremism went from 2.56% to zero for both models, and sexual content dropped by roughly half to three quarters. Those are real.

But on nonviolent and violent wrongdoing, both new models are worse than the model they replace: 2.86% for Sunburst and 3.27% for Flare, against 2.45% for Images 2.0. And Flare, the default, matches rather than beats the old model on violence and gore (1.93%) and on abuse (2.54%). Sunburst is the safer of the two in seven of nine categories.

The safer model is not the default. If you are deploying this in a product where a violative output is your problem rather than OpenAI's, that is worth knowing before you accept the default.

Three caveats travel with all of these numbers, and they cut in the model's favor. The prompts were built specifically to force violations, and OpenAI states the results are "not representative of how often such prompts arise in production traffic." Automated policy labels can be wrong. And per-category sample sizes are not published, so a gap of a few tenths of a percent may rest on a handful of examples. This is not evidence that the model is dangerous. It is evidence that "safer" is a more complicated claim than a single row of a table suggests.

The deepfake line, and a new watermark

OpenAI names the risk directly rather than letting someone else name it:

"ChatGPT Images 2.5 allows for heightened realism that could, absent safeguards, allow more convincing deepfakes, including political, sexual, or otherwise sensitive imagery of real people, places or events."

Alongside that, a change that did not make the launch post and only appears in the system card: OpenAI has added Google DeepMind's SynthID invisible watermarking, across ChatGPT, Codex and the API, layered on top of its existing C2PA metadata.

An AI lab adopting a direct competitor's provenance technology is not a small thing. It is probably the most substantive item in the entire release, and it was published in an appendix. OpenAI is careful not to oversell it. Its own framing is that "there is no single solution to provenance," and a watermark you can strip is not a guarantee.

What to do this week

If you use ChatGPT for images, try Sketch on the next thing where you catch yourself writing a paragraph to describe a layout. Drawing a box takes four seconds and specifies more.

If you build on the API, the migration from GPT-Image-2 is free in the literal sense. Check batch pricing first if you rely on it, and default to Sunburst rather than Flare where a bad output would be expensive, because it costs the same.

If you generate at high volume, measure gpt-image-1-mini against 2.5 rather than reading the rate card. Mini is still cheapest at 3.36 cents high, but 2.5 at medium is 1.35 and may be good enough.

If you publish what you generate, the deepfake sentence and the SynthID addition are the two lines from today worth carrying into whatever policy you have. Provenance is now layered, and still not a guarantee.

The genuinely good news is the simplest thing here. Image generation got materially faster and better, and the price did not move. That is worth more than the number in the headline.


Facts verified 8 September 2026 against OpenAI's launch post, system card and published API pricing. Safety percentages are from OpenAI's adversarial evaluation set and are explicitly not production rates.

MB
Michael Bennett
I build AI marketing systems that acquire, convert & retain customers.

Working out where AI actually fits in your marketing?

I write these while building the systems behind them: measurement, creative pipelines, and agents that do real work. Connect on LinkedIn and tell me what you are working on. That is where these conversations start.

Connect on LinkedIn