NVIDIA announces GPT-6 Astra Ultrafast on October 2, 2026, saying the mode is available through the OpenAI API and to eligible ChatGPT Work and Codex users. NVIDIA says it runs on NVIDIA Blackwell GPUs and claims token generation can be up to 8× faster than Astra Standard.
That figure is NVIDIA’s maximum claim, not a promise of the same speed for every task. For developers, the practical starting point is the API’s exact model and service-tier settings, plus the rate limits tied to each usage tier.
For background on the model’s broader rollout, NeoTeo previously covered GPT-6 Astra.
Configure GPT-6 Astra Ultrafast through the API
Set the API model to gpt-6-astra and the service tier to ultrafast. OpenAI supports both HTTP and WebSocket requests. TPM means tokens per minute: it is a request limit, not a measure of how quickly a particular response will arrive.
OpenAI lists these default Ultrafast limits by API usage tier:
| API usage tier | Default limit |
| Tiers 1–3 | 500,000 tokens per minute |
| Tier 4 | 1,000,000 tokens per minute |
| Tier 5 | 5,000,000 tokens per minute |
For agentic applications that make frequent tool calls, OpenAI recommends persistent WebSocket connections because network overhead can eat into the latency benefit.
Work and Codex eligibility and processing options
OpenAI lists Ultrafast for one Pro tier and eligible Enterprise and Edu workspaces in ChatGPT Work and Codex. Workspace access uses workspace credits and depends on the necessary permissions. Plus, Business and other Pro tiers are not listed for Ultrafast at launch; ordinary Astra access in Work and Codex is separate from Ultrafast access.
For API processing, OpenAI lists U.S. data residency and global processing. The guide does not offer EU or other non-U.S. regional processing endpoints; that is a limit on regional processing choices.
NVIDIA’s Blackwell and speed claims
NVIDIA attributes the acceleration to inference optimization on its GPUs and their programmability. OpenAI says it used internal models to optimize inference on NVIDIA GPUs. NVIDIA’s announcement identifies Blackwell GPUs and compares Ultrafast with Astra Standard, but its stated figure is specifically about token-generation speed.
A Blender workflow, not a speed comparison
A sponsored demonstration shows GPT-6 Astra Ultrafast working through Blender to create a 3D dragon model. Its narration reports a little over 10 minutes for the initial version, followed by an additional 35 minutes for refinement. Those timings describe one workflow, not a controlled comparison with Astra Standard.