GPT-5.4 Pro uses more compute to think harder and provide consistently better answers.
GPT-5.4 Pro is available in the Responses API only to enable support for multi-turn model interactions before responding to API requests, and other advanced API features in the future. Since GPT-5.4 Pro is designed to tackle tough problems, some requests may take several minutes to finish. To avoid timeouts, try using background mode. Reasoning.effort supports: medium (default), high and xhigh.
For models with a 1.05M context window (GPT-5.4 and GPT-5.4 Pro), prompts with >272K input tokens are priced at 2x input and 1.5x output for the full session for standard, batch, and flex.
Regional processing (data residency) endpoints are charged a 10% uplift for GPT-5.4 and GPT-5.4 Pro.
| Tier | RPM | TPM | Batch queue limit |
|---|---|---|---|
| Free | Not supported | ||
| Tier 1 | 50 | 50,000 | 900,000 |
| Tier 2 | 500 | 100,000 | 1,350,000 |
| Tier 3 | 500 | 200,000 | 100,000,000 |
| Tier 4 | 1,000 | 400,000 | 200,000,000 |
| Tier 5 | 1,500 | 4,000,000 | 15,000,000,000 |