Skip to content
Skip to content
AI consultancy

OpenAI API rate limits drop to three tiers, and a cheaper way to decide arrives

Summarise with

OpenAI trimmed its API ladder from five rungs to three and, on the same day, put a new endpoint into beta that answers yes, pick or score for a tenth of a dollar per million input tokens. The sticker is the easy part.

Quick answerOpenAI API rate limits now sit on three tiers, Build, Launch and Grow, with the top tier needing $500 in credit purchases for $200,000 a month. The new Decisions API adds cheap typed answers, untested for accuracy.
Section 01

The new OpenAI API rate limits in one ladder#

A fox on the prowl does not climb five stairs when three will do. On 6 October OpenAI's changelog recorded that it had simplified API usage tiers from five to three, and the same day it put a new endpoint into beta. The tier change is the dull half of the news and the half most likely to touch your budget, because OpenAI API rate limits decide how many requests a team can send before the gate swings shut.

Here is the new ladder of OpenAI API rate limits as OpenAI's own rate limits guide states it. Build needs $5 in total credit purchases and carries a $500 monthly usage limit. Launch needs $100 and allows $5,000 a month. Grow needs $500 and allows $200,000 a month. Organisations move up automatically as their credit purchases cross each line, so nobody applies for a tier.

Each rung of the OpenAI API usage tiers is a different order of magnitude
Staircase chart on a log scale of OpenAI API monthly usage limits: Free 100 dollars, Build 500, Launch 5,000 and Grow 200,000Monthly usage limit by tier, log scale$100$1000$10000$100000$1000000Free: $100$100FreeBuild: $500$500BuildLaunch: $5000$5000LaunchGrow: $200000$200000Grow
Staircase chart on a log scale of OpenAI API monthly usage limits: Free 100 dollars, Build 500, Launch 5,000 and Grow 200,000
ItemValue
Free$100
Build$500
Launch$5000
Grow$200000
Monthly usage limits from OpenAI's rate limits guide, 7 October 2026. The free tier is $100 a month; the top paid tier is two thousand times that.

The viral summary of the OpenAI API rate limits got one detail slightly wrong, and the correction is the useful part. One widely shared post said OpenAI had cut the price of its biggest API limits in half. What halved is the entry ticket, not the ceiling. An archived copy of the same page from 1 October shows the old top tier, Tier 5, needed $1,000 in payments for the same $200,000 a month limit that Grow now offers for $500.

@trivediprateek
openai just cut the price of its biggest api limits in half. openai's top api tier used to need $1,000 in payments. now it's $500. five paid tiers just became three: build, launch, grow.
7 October 2026View on X
Credit purchases needed to reach two OpenAI API ceilings, in US dollars
Credit purchases needed to reach two OpenAI API ceilings, in US dollarsDumbbell chart of OpenAI API credit purchases required: the top tier fell from 1,000 to 500 dollars and the 5,000 dollar a month tier fell from 250 to 100 dollarsBefore 6 OctAfter$200k a month tier: 1000 to 500$200k a month tier500$5k a month tier: 250 to 100$5k a month tier100
Dumbbell chart of OpenAI API credit purchases required: the top tier fell from 1,000 to 500 dollars and the 5,000 dollar a month tier fell from 250 to 100 dollars
ItemValue
$200k a month tier1000 to 500
$5k a month tier250 to 100
The $200,000 a month ceiling now needs $500 of credit purchases, down from $1,000; the $5,000 a month ceiling needs $100, down from $250. The ceilings themselves did not move.

The old thresholds come from the archived page. In the old scheme the $5,000 a month rung was Tier 4 at $250 paid, and in the new one it is Launch at $100. Easier to reach, then, but not larger. Anyone whose volume was capped by the old ladder still meets the same ceiling.

Section 02

How the OpenAI API usage tiers compare with two rivals#

A ladder is only useful next to other ladders, so here is a neutral sniff along the trail, dated 7 October and attributed to each vendor's own page. Anthropic names its tiers Start, Build and Scale, with monthly spend caps of $500, $1,000 and $200,000, and it says organisations are placed on a tier automatically based on usage history and account standing. Google's Gemini API ladder is different again: its third paid tier needs $1,000 paid plus 30 days from the first successful payment, and its limits apply per project, not per API key.

Three vendors, three philosophies: OpenAI API rate limits now qualify on cumulative credit purchases, Anthropic on usage history and standing, Google on spend plus waiting days. None of these is better in the abstract. They matter when you plan a launch, because the question is not what the tier is called but how many days it takes to reach the ceiling your campaign needs.

  • Build

    Entry paid tier

    Qualifies at
    $5 total credit purchases
    Monthly usage limit
    $500
    Luna limits
    5,000 RPM, 2M TPM
    Flagship limits
    5,000 RPM, 1M TPM

    Best for

    • Prototypes and pilots
  • Launch

    Working volume

    Qualifies at
    $100 total credit purchases
    Monthly usage limit
    $5,000
    Luna limits
    10,000 RPM, 10M TPM
    Flagship limits
    10,000 RPM, 4M TPM

    Best for

    • Production features
  • Grow

    Top paid tier

    Qualifies at
    $500 total credit purchases
    Monthly usage limit
    $200,000
    Luna limits
    30,000 RPM, 180M TPM
    Flagship limits
    15,000 RPM, 40M TPM

    Best for

    • High-volume pipelines

One quiet detail in the OpenAI API rate limits sits in that card. The small model gets a much wider lane than the flagships: at Grow, 180 million tokens a minute for Luna against 40 million for the flagship models. If a job can be done by the small model, the rate limits quietly say so. That sets up the second half of the news.

Section 03

What the Decisions API actually returns#

The new endpoint is not a chat model. According to OpenAI's Decisions guide, it evaluates text, images or both and hands back one of three answer types. A predicate returns a probability that a condition is true. A choice picks one option from a list you supply. A score rates the input against ordered levels. Nothing is written; something is decided.

It is a public beta, so hold your plans around the OpenAI API rate limits loosely: OpenAI says it expects general availability in the coming weeks, and only the small GPT-6 Luna model is supported. Images must be sent inline as encoded data, not as links. The guide also tells builders to include a fallback option such as an "other" choice, which is the sort of humility every classifier needs.

  • 10x

    OpenAI's own speed claim against the Responses API, no baseline stated

  • $0.10

    per 1M input tokens with Luna, no output charge

  • 3

    answer types: predicate, choice and score

  • 0

    accuracy or calibration figures in the guide

The price needs two readings. The guide says input costs $0.10 per 1M tokens and that there are no cache or output-token charges. Then it adds the usual small print: regional processing premiums and long-context multipliers apply The pricing page itself does not list the Decisions API, so the $0.10 lives in one place only.

watercolour fox holding three brass keys of rising size, illustrating the three OpenAI API usage tiers and openai api rate limits
Three keys where there were five.
Section 04

Where a cheap, typed decision earns its place#

Marketing teams are full of small judgements that nobody wants to pay a flagship model to make: is this lead worth a call, does this comment need a human, which of five queues should this email land in, does this ad creative break the brand rubric. Each is a yes, a pick or a score, which is precisely the shape of the three answer types. A quiet change in the OpenAI API rate limits and a new endpoint may turn those judgements from a manual chore into a pipeline.

Jobs that fit a predicate, a choice or a score

  • Lead triage
  • Comment moderation
  • Support routing
  • Creative rubric scoring
  • Search query intent

The guide's own advice on thresholds is the part to copy. OpenAI says to use labelled examples from your application and to choose thresholds by the cost of false positives and false negatives. That is the whole method in two sentences. A false positive on lead scoring wastes a sales call; a false negative on a complaint router loses a customer. They cost different amounts, so they deserve different cut-offs.

01

Lead triage, predicate

Flag leads worth a human call · 15 minutes to draft, a week to calibrate

Question type: predicate
Instruction: Does this enquiry describe a regulated business with a budget over a stated minimum and a deadline inside ninety days?
Then: send anything above your chosen threshold to a person, and log every answer.
02

Complaint routing, choice

Route to the right team · 20 minutes to draft

Question type: choice
Options: billing, technical, delivery, press, other
Instruction: Which team should handle this message?
Then: always keep an other option, and send low-confidence answers to a general queue.

Treat those cards as templates for your own labelled set, not as tested recipes. The point of a fast, cheap decision lane is that you can afford to run it on every item, which makes the evidence you gather on a hundred real examples worth more than any vendor table. We help clients design exactly that kind of test inside our AI consultancy work, and the lead-triage pattern connects naturally to paid search programmes where lead quality, not click volume, pays the bills.

Section 05

Why chatgpt api pricing is not your running cost#

Search any week and you will find people asking why is openai api so expensive, and others quietly running it for pennies. Both are right, because the headline rate is the start of the sum and the pricing thicket grows around it, whatever the OpenAI API rate limits say. OpenAI's own model page for GPT-6 Luna says prompts above 272,000 input tokens are priced at double the input rate and one and a half times the output rate for the whole request, and that Batch and Flex are half price while Fast mode costs double.

Standard is $0.10 input and $0.50 output; the other rows apply the multipliers OpenAI states on the model page, which its pricing page also lists. Check the pricing page before you budget.
ModeInputOutput
Standard, up to 272K input tokens$0.10$0.50
Long context, above 272K$0.20$0.75
Batch or Flex$0.05$0.25
Fast$0.20$1.00
  • Standard, up to 272K input tokens$0.10$0.50
  • Long context, above 272K$0.20$0.75
  • Batch or Flex$0.05$0.25
  • Fast$0.20$1.00

The flagship comparison is just as stark. The changelog gives GPT-6.1 Sol standard pricing as $2 for input and $10 for output, which is twenty times Luna on both counts at list price. That ratio is the reason to ask which of your jobs need the flagship at all. Our earlier notes from the same undergrowth, on GPT-6.1 Sol and its allowance and on what an AI subscription is really worth make the same point from the consumer side: the meter, not the sticker, is the story.

Please also keep the balance rule in mind. These are the vendor's published prices on one date, and capability claims in this field go stale within a quarter. Re-read the pricing page the day you commit, and put the date in your budget note.

Section 06

What nobody has published yet: calibration, residency and the law#

Every vendor speed claim about OpenAI API rate limits or latency is marketing until someone tests it. The 10x figure appears in the changelog and the guide with no baseline model, payload size or method. The guide also says that choice and score answers return a probability distribution and a separate confidence field, but it does not say how that confidence is derived, and it gives no accuracy or calibration figure. The nearest independent evaluator, Artificial Analysis, scores the Responses route at maximum reasoning, which is a different thing from a Decisions call.

Why should a buyer care? Because probabilities from language models are not automatically trustworthy. Researchers found that larger models can be well calibrated in the right format, yet a later study reported that models verbalising their confidence tend to be overconfident. One group found that verbalised confidences were often better calibrated than token probabilities, and an older neural network paper concluded that modern networks are poorly calibrated. These studies predate Luna and concern other tasks, so they prove nothing about Decisions. They justify one habit: test the probability you are handed before you build a workflow on it.

The legal and data questions deserve a sentence each. OpenAI's data controls page lists the endpoint under both US and EU residency, which matters if your leads are European. If a decision about a person is made solely by automated processing, GDPR Article 22 gives that person the right to human intervention. And where a system is used to filter job applications or assess creditworthiness, Annex III of the EU AI Act treats it as high-risk. The Commission's own page says those obligations have been extended to 2 December 2027, although the original text said 2 August 2026, so check your use case with counsel rather than a blog post.

The practical close is gentle. Whatever your tier, and whatever the OpenAI API rate limits allow, move one small, cheap, reversible decision onto the new endpoint. Label a hundred real examples. Choose thresholds by what each error costs you. Keep a person in the loop for anything about a human being. If it pays, widen the lane and revisit the OpenAI API rate limits for your tier; if it does not, you have spent an afternoon and learned the shape of your own data. That is how a patient fox crosses a new field, quarry in mind and den behind: one quiet step, one careful look, and a way back to the hedgerow.

Questions

Frequently asked questions#

What are the OpenAI API rate limits after the October 2026 change?

OpenAI now has three paid usage tiers. Build needs $5 in total credit purchases and allows $500 a month, Launch needs $100 and allows $5,000, and Grow needs $500 and allows $200,000. Per-model requests and tokens per minute rise with each tier, which is how OpenAI API rate limits are set, per OpenAI's rate limits guide.

What are the OpenAI API usage tiers called now?

The paid OpenAI API usage tiers are Build, Launch and Grow, replacing the previous five numbered tiers. Organisations upgrade automatically as their total credit purchases reach each threshold, so the OpenAI API rate limits follow your spending with no application step, and the free tier remains at $100 a month.

How much does the Decisions API cost?

OpenAI's Decisions guide says input costs $0.10 per 1M tokens with gpt-6-luna and there are no cache or output-token charges. It also says regional processing premiums and long-context multipliers apply. The pricing page does not list the endpoint, so confirm the rate in the guide on the day you commit.

Why is openai api so expensive?

Often it is not the rate but the route, and the OpenAI API rate limits are rarely the culprit. Long prompts above 272,000 tokens are charged at higher multiples, Fast mode costs double, and flagship models list at twenty times the small model's price. Matching each job to the cheapest model that passes your own test usually cuts the bill more than any discount.

Is the Decisions API really ten times faster?

OpenAI says it returns typed answers about 10x faster than the Responses API, but it gives no baseline model, payload size or method. Treat it as a vendor claim and time it on your own inputs. The guide also publishes no accuracy or calibration figures for the confidence it returns.

Does gpt api pricing differ between Luna and Sol?

Yes. At standard processing GPT-6 Luna lists at $0.10 input and $0.50 output per 1M tokens, while GPT-6.1 Sol lists at $2 input and $10 output, twenty times higher on both. Long-context, Batch, Flex and Fast modes then change each rate.

Keep reading

Read more on this topic#

Not sure which jobs belong on which model?

folkfox helps teams test cheap, typed model decisions on their own data before they commit budget: labelled sets, error-cost thresholds and a person kept in the loop.

Want folkfox in your Google results and AI answers? Set folkfox as a preferred source.

The den

Where to?

Pricing

Choose a section. Enter opens it, Escape continues reading.

Cookie preferences

folkfox uses data the way we use strategy: only when it earns its place.