- Sources: primary, discussion
- Summary: Cerebras and OpenAI previewed an Ultrafast tier for GPT-5.6 Sol at up to 750 output tokens per second, as a limited preview for selected customers with no published price. The Humanity's Last Exam and GDP-Val figures are Cerebras' own benchmarking, with run dates in July, including a Humanity's Last Exam run of 11 hours 11 minutes against 78 hours 27 minutes for Claude Fable 5. The post attributes its output-speed multiples to Artificial Analysis.
- Why it matters: A frontier model at that output rate moves agent work off the wait-and-context-switch pattern, but the tier is a limited preview for selected customers and carries no published price.
send feedback on this story