Accelerating GPT-5.6 Sol Ultrafast
Ranked #2 on Hacker News with 271 points and 107 comments.
Today, Cerebras and OpenAI are sharing an early look at Ultrafast Mode , a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast is available initially to a select group of customers, with access expanding over time. Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second and without any quality compromise, allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical work.
Frontier Intelligence at Unprecedented Speed
AI builders have always needed to choose between speed and intelligence. As models scale up in size and intelligence, they incur higher computational and data movement costs, slowing down response times. Users often need to wait for high-quality results or accept inferior results within a shorter timeframe.
GPT-5.6 Sol Ultrafast resolves this tradeoff, bringing frontier intelligence to products and workflows where every second matters. Compared with output speeds reported by Artificial Analysis GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.
At Cerebras, we put Ultrafast to the test by running it head-to-head with popular models on Humanity's Last Exam. HLE is a challenging model benchmark that consists of 2,500 questions typically answerable only by those holding PhDs in fields such as chemistry, economics, and literature.
In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.