- Sources: OpenAI model documentation, HN discussion
- Summary: The model page states 4 dollars per million input tokens and 20 dollars per million output tokens, a 20 percent reduction on input and 33 percent on output, and gives the pricing as promotional at least through 2026-11-21. Cached input stays at 0.40 dollars per million and cache writes bill at 1.25x the uncached input rate. A prompt above 272K input tokens prices the whole request at 2x input and 1.5x output.
- Why it matters: The realised saving follows cache behaviour and prompt shape rather than the headline rate, because cache writes bill above uncached input and one prompt over 272K tokens reprices the entire request.
- Follow-up: Track whether the promotional rate holds past 2026-11-21.
send feedback on this story