OpenAI shipped ChatGPT Images 2.5 today. The coverage so far says sharper, faster, better at editing, and all of that is true. It is also the least interesting thing about the release.
Two facts are sitting in OpenAI's own documents that nobody has picked up, and both change what you should actually do about this.
The first is in the pricing table. The second is in a footnote in the system card.

What actually shipped
The model is better at working from your reference photos, so a person or a product stays recognizable when you move it into a new setting. It edits only what you asked it to edit and leaves the rest alone. It holds its work together across a long back-and-forth instead of degrading a little with every turn. It handles transparent backgrounds and more complex layouts.
Generation latency is down up to 50% against Images 2.0. Note the hedge: "up to" is doing work in that sentence.
Four product features came with it, and these are the ones you can use this afternoon:
- Sketch. Draw your idea directly in ChatGPT and use it as the layout guide. Type
@Sketch. Rough is fine; the point is to specify composition without writing a paragraph describing it. - Templates. Named starting points for common formats, "Poster" and "Merch" among them, instead of a blank prompt box.
- Comments on images. Point at the part you want changed rather than describing its location in words.
- Shareable prompts. Share an image with the prompt attached, so someone else can run it with their own photos.
It is live for all ChatGPT, ChatGPT Work, and Codex users, on every tier, across desktop, mobile and web. No staged rollout by plan, which is not how most of these land.
Developers get two API models: GPT-Image-2.5 Flare, the default, and GPT-Image-2.5 Sunburst, built for precision at the cost of longer generation times.
A picture costs about a penny
Image pricing is published as dollars per million tokens, which is why almost nobody quotes a figure you can use. So I generated images against the live API and read the cost off each response. These are measured, not estimated:
| Quality | Per image | 100 images | Time |
|---|---|---|---|
| Low | 0.62 cents | $0.62 | 8-11s |
| Medium | 1.35 cents | $1.35 | 8-11s |
| High | 5.30 cents | $5.30 | 16-21s |
A standard 1024x1024 image. Landscape at high quality is cheaper still, 4.14 cents.
These figures are exact rather than sampled: output token counts turned out to be deterministic for a given model, size and quality. Two completely different prompts returned identical counts, 196, 439 and 1756 tokens, so the same request always costs the same.
The published table would have misled you, and it misled me
OpenAI publishes a per-image cost table, and it does not cover the 2.5 models. It covers GPT Image 2, at 0.6, 5.3 and 21.1 cents. Since 2.5 bills at identical per-token rates, the reasonable inference is that per-image costs match.
That inference is wrong, and I published it before I checked. 2.5 emits roughly a quarter of the output tokens for the same quality setting, so the same nominal quality costs about a quarter as much:
| Same prompt, same size | GPT Image 2 | ChatGPT Images 2.5 |
|---|---|---|
| Medium | 5.30 cents, 40s | 1.35 cents, 8-11s |
| High | 21.10 cents, 121s | 5.30 cents, 16-21s |
High quality on 2.5 costs exactly what medium cost on GPT Image 2, and arrives six times faster.

Identical per-token pricing does not imply identical per-image cost, which is the trap in reading a rate card instead of a bill.
What that means for your budget
A hundred product images at high quality is $5.30. A thousand is $53. Drafts are two thirds of a cent, so experimenting is effectively free.
The one comparison worth running before you migrate: gpt-image-1-mini measured 3.36 cents at high and 0.87 at medium. Still the cheapest option, but the gap has narrowed enough that it is no longer the obvious choice it looks like on the published tables, and 2.5 at medium now undercuts Mini at high.
What it actually made
Everything below was generated for this article, against the live API, at the prices shown.

You have one good photograph and need four: on white for the store, on a table for the ad, in a lifestyle shot for social, seasonal for the campaign. That was a reshoot, a designer, or an afternoon in Photoshop.

The second thing that changed is text. Words inside a generated image used to come out as plausible-looking nonsense. Both sets of copy below were typed into the prompt and came back correct, correctly spelled and sensibly laid out.

You can also skip the description entirely and draw. In ChatGPT this is Sketch: type @Sketch and draw the layout, however roughly.

And it is not locked to one house style. The same subject and the same sentence, with only the style instruction changed:

The per-token table, and what it tells you
| Model | Image in | Cached | Output |
|---|---|---|---|
gpt-image-2.5-sunburst | $8.00 | $2.00 | $30.00 |
gpt-image-2.5-flare | $8.00 | $2.00 | $30.00 |
gpt-image-2 | $8.00 | $2.00 | $30.00 |
gpt-image-1-mini | $2.50 | $0.25 | $8.00 |
Three things follow, and all three are arithmetic rather than opinion.
The new model costs exactly what the old one costs. Identical on every line. A meaningful quality jump and up to half the latency, at no price increase. That is not the usual shape of a frontier model release, and if you are already building on GPT-Image-2 the migration question answers itself.
The premium model is not a premium tier. Flare and Sunburst are priced the same, to the cent. OpenAI positions Sunburst for "premium visual workflows," which reads like it should cost more, and it does not. So the choice between them is purely latency against precision. If you are running a campaign asset through six rounds of edits, there is no budget argument for staying on the fast one.
There is still no 2.5 mini. gpt-image-1-mini outputs at $8.00 against $30.00 per million tokens. Measured per image it is 3.36 cents at high and 0.87 at medium, so it remains the cheapest option. But the gap is far narrower than the per-token rates suggest, and 2.5 at medium (1.35 cents) now undercuts Mini at high (3.36 cents). Price them against each other before assuming the old model wins on cost.
Two cautions. Pricing is per token, not per image, so cost per image depends on size and quality settings; measure your own workload rather than trusting a quoted figure, including mine. And no batch price is published for either 2.5 model, while GPT-Image-2 has one at half the standard rate. If your pipeline depends on the batch discount, check before you migrate.
About that speed number
OpenAI says up to 50% lower latency. Fifty percent lower latency is two times faster.
Manus, a launch partner quoted on OpenAI's own page, says Flare delivers "high-quality images at two to four times the speed of GPT-Image-2." That top end is double what OpenAI claims for its own model.
It may well be right. But it is a partner evaluation printed in a launch announcement, not an independent benchmark, and no third-party numbers exist yet. Plan against OpenAI's figure. Treat anything above it as upside.
The number everyone will misread
OpenAI published a system card, and it contains a safety table that looks like a clean win.
Against an adversarial prompt set, the share of violative images that slipped through and reached the user:
| Unsafe shown | |
|---|---|
| GPT-Image-2.5-Sunburst | 1.09% |
| GPT-Image-2.5-Flare | 1.41% |
| ChatGPT Images 2.0 | 1.64% |
Both new models beat the old one. You will see that framed as a safety improvement this week.
Now read the footnote. OpenAI marks statistically significant results with an asterisk, and then states, in its own words:
"No unsafe-shown difference meets this threshold."
The headline safety improvement is not statistically significant, and OpenAI is the one telling you so. Nobody is hiding it. It is simply in a footnote under a table, which is where things go when they are true and inconvenient.
Break it down by category and it gets more interesting. Two genuine wins: extremism went from 2.56% to zero for both models, and sexual content dropped by roughly half to three quarters. Those are real.
But on nonviolent and violent wrongdoing, both new models are worse than the model they replace: 2.86% for Sunburst and 3.27% for Flare, against 2.45% for Images 2.0. And Flare, the default, matches rather than beats the old model on violence and gore (1.93%) and on abuse (2.54%). Sunburst is the safer of the two in seven of nine categories.
The safer model is not the default. If you are deploying this in a product where a violative output is your problem rather than OpenAI's, that is worth knowing before you accept the default.
Three caveats travel with all of these numbers, and they cut in the model's favor. The prompts were built specifically to force violations, and OpenAI states the results are "not representative of how often such prompts arise in production traffic." Automated policy labels can be wrong. And per-category sample sizes are not published, so a gap of a few tenths of a percent may rest on a handful of examples. This is not evidence that the model is dangerous. It is evidence that "safer" is a more complicated claim than a single row of a table suggests.
The deepfake line, and a new watermark
OpenAI names the risk directly rather than letting someone else name it:
"ChatGPT Images 2.5 allows for heightened realism that could, absent safeguards, allow more convincing deepfakes, including political, sexual, or otherwise sensitive imagery of real people, places or events."
Alongside that, a change that did not make the launch post and only appears in the system card: OpenAI has added Google DeepMind's SynthID invisible watermarking, across ChatGPT, Codex and the API, layered on top of its existing C2PA metadata.
An AI lab adopting a direct competitor's provenance technology is not a small thing. It is probably the most substantive item in the entire release, and it was published in an appendix. OpenAI is careful not to oversell it. Its own framing is that "there is no single solution to provenance," and a watermark you can strip is not a guarantee.
What to do this week
If you use ChatGPT for images, try Sketch on the next thing where you catch yourself writing a paragraph to describe a layout. Drawing a box takes four seconds and specifies more.
If you build on the API, the migration from GPT-Image-2 is free in the literal sense. Check batch pricing first if you rely on it, and default to Sunburst rather than Flare where a bad output would be expensive, because it costs the same.
If you generate at high volume, measure gpt-image-1-mini against 2.5 rather than reading the rate card. Mini is still cheapest at 3.36 cents high, but 2.5 at medium is 1.35 and may be good enough.
If you publish what you generate, the deepfake sentence and the SynthID addition are the two lines from today worth carrying into whatever policy you have. Provenance is now layered, and still not a guarantee.
The genuinely good news is the simplest thing here. Image generation got materially faster and better, and the price did not move. That is worth more than the number in the headline.
Facts verified 8 September 2026 against OpenAI's launch post, system card and published API pricing. Safety percentages are from OpenAI's adversarial evaluation set and are explicitly not production rates.
Working out where AI actually fits in your marketing?
I write these while building the systems behind them: measurement, creative pipelines, and agents that do real work. Connect on LinkedIn and tell me what you are working on. That is where these conversations start.
Connect on LinkedIn