Skip to content
Table of contents6 sections · tap to jump
  1. What OpenAI published
  2. Where the second price starts
  3. The multiplier stack
  4. The rate limit nobody mentions
  5. What to do about it
  6. One caveat on that last point
Wooden typesetting box with metal movable type characters arranged in compartments, one character in focus in the center

Newsbusiness6 min read

GPT-6 Astra has a million-token window and a price cliff at 272,000

Ahmad JSep 5, 2026

Signalsolid1independent source

OpenAI's API pricing page lists 37 models. Eight of them have two prices.

GPT-6 Astra, published today, is the newest of the eight and the most expensive model on the page. Its headline rate is $10.00 per million input tokens and $50.00 per million output tokens. Its second rate is $20.00 and $75.00.

The line between the two sits at 272,000 input tokens. The model's context window is 1,050,000.

So the advertised window is a little under four times the size of the cheap half of it, and almost everything most people would actually buy a million-token window for lands on the far side of the line.

What OpenAI published#

From the model page, in OpenAI's own words: "GPT-6 Astra is our most capable model, built for the hardest end-to-end work. Use it for complex reasoning, coding, computer use, research, and document creation."

SpecValue
Model IDgpt-6-astra
Context window1,050,000 tokens
Maximum input tokens922,000
Maximum output tokens128,000
Knowledge cutoffApr 30, 2026
Input modalitiesText, image
Output modalitiesText
Reasoning effort settingslow, medium, high, xhigh, max

Worth noting that the context window and the maximum input are different numbers. You cannot fill 1,050,000 tokens with prompt. The input ceiling is 922,000, and the remainder is headroom for the model's own output and reasoning tokens.

Where the second price starts#

The rule is one sentence on the model page:

"Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request."

The last four words are the expensive part.

This is not a marginal rate. Nothing is billed in bands. Cross 272,000 input tokens by a single token and every token in the request, including the 272,000 that were under the line a moment ago, reprices.

Two requests, identical except for one token of input:

272,000 in, 10,000 out272,001 in, 10,000 out
Input0.272M at $10.00 = $2.720.272M at $20.00 = $5.44
Output0.010M at $50.00 = $0.500.010M at $75.00 = $0.75
Total$3.22$6.19

One token adds $2.97. The request costs 92% more than the one before it.

That is what a cliff means, and it is the single most important thing to know about running this model. A slope you can optimise gradually. A cliff you either stay behind or you do not, and there is no partial credit for being close.

The multiplier stack#

The 272K rule is one of four things that move the number. All are published, none are hidden, and they compose.

Service tierInput, under 272KOutput, under 272KInput, over 272KOutput, over 272K
Batch$5.00$25.00$10.00$37.50
Flex$5.00$25.00$10.00$37.50
Standard$10.00$50.00$20.00$75.00
Fast mode$20.00$100.00$40.00$150.00

Batch and Flex are 50% of Standard. Fast mode, which OpenAI renamed from Priority processing on July 30, 2026, is 2x. On top of any of these, regional processing endpoints for data residency carry a 10% uplift on models released on or after March 5, 2026, which includes this one.

Run the extremes on a single maximum-size request, 922,000 tokens in and 128,000 out:

  • Batch, over the line: $9.22 + $4.80 = $14.02
  • Standard, over the line: $18.44 + $9.60 = $28.04
  • Fast mode, over the line: $36.88 + $19.20 = $56.08

One call. The gap between the cheapest and dearest way to send exactly the same tokens is four times.

The rate limit nobody mentions#

There is a second ceiling that has nothing to do with money.

At usage Tier 1, GPT-6 Astra allows 500,000 tokens per minute. The maximum input is 922,000 tokens. So a new account on Tier 1 cannot send one maximum-length prompt at all, regardless of budget. Tier 2 raises it to 1,000,000 tokens per minute, which fits roughly one such request per minute. Tier 3 doubles that again to 2,000,000.

If your plan for this model involves very large documents, the tier you are on is a harder constraint than the price, and it moves only as you spend.

What to do about it#

Four things, in the order they are worth doing.

Treat 272,000 as a hard budget line. This is the whole game. Measure your real prompts, not your estimated ones, and if they land anywhere near the line, cut them until they clear it with room to spare. A retrieval step that trims context is now worth roughly double what it was worth yesterday, because it is not saving you tokens at $10.00, it is saving you the reprice of every other token in the request. This is the concrete version of a general point about why a bigger context window is not automatically better.

Cache, and know when caching starts paying. Cached input is $1.00 per million against $10.00 uncached, a 90% saving. But cache writes are billed at $12.50, which is 1.25x the uncached input rate. So a cached prefix costs more on the first call and less on every call after it. The break-even is early, but it is not zero, and caching only pays on a stable prefix you actually reuse.

Use Batch or Flex unless you need the latency. Half price, same tokens, same model. Most document processing, evaluation runs and offline analysis do not need a synchronous answer. This is the least clever item on the list and usually the largest saving.

Check whether you need Astra. GPT-5.6 Sol is $4.00 in and $20.00 out, which is 40% of Astra's rate. GPT-5.6 Terra is $2.00 and $12.00. GPT-5.6 Luna is $0.20 and $1.20, a fiftieth of Astra's input price. All three carry the same 272K structure, so the cliff is not the thing that distinguishes them. The framework for that decision is the same one you would use for any model going into production: find the cheapest model that passes your evaluation, then stop.

One caveat on that last point#

The obvious advice, use Sol instead, comes with a date attached.

OpenAI's pricing page carries a footnote: "GPT-5.6 Sol's promotional pricing is available at least through November 21, 2026." So Sol's $4.00 and $20.00 are a promotion with a published floor date and no published successor rate. Budgeting a production workload on it is the same mistake as budgeting on Gemini Flash's current half price, and the fix is the same: know which of your numbers is a promotion before you build a margin on it.

For contrast, the previous flagship is still listed and still cheaper than the new one. GPT-5.5 sits at $5.00 in and $30.00 out with no promotional footnote, and its long-context rates are $10.00 and $45.00. Against Astra's $10.00 and $50.00, the new model is exactly double the input price of the one it succeeds.

That is not a complaint. Astra is a bigger model with a much larger window, and OpenAI is entitled to charge for it. It does mean the upgrade is a pricing decision as much as a capability one, and the only way to see that is to read the rate card rather than the announcement. Everything above is on two pages OpenAI publishes for free, and the sentence that matters most is twelve words about how tokens are counted rather than anything in a launch post.

Sources

  1. GPT-6 Astra model page, OpenAI API documentationdevelopers.openai.com
  2. API pricing, OpenAI API documentationdevelopers.openai.com

Ask about this article

Answered only from this piece — the AI never invents.

React
ShareXLinkedInBluesky

More in businessMore in business

Discussion