Ranked #2 on Hacker News with 563 points and 251 comments.
For the past few weeks, I've been playing with gpt-5.6-luna . It is shockingly capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base.
Of course, the biggest thing with luna is the cost . I've tried running some fairly complicated research threads, and it's pretty tough to run up a large bill. Even having it search across thousands of emails, I end up with an API cost in the tens of cents.
With GLM 5.3, we even have a new option at the Pareto frontier.
When doing coding work, I almost always reach for the most expensive and capable models (Fable 5, 5.6 Sol). So it's been easy to miss the progress the small fast models have made.
One thing a few investors I've talked with have mentioned: "It's weird we're not seeing more consumer AI companies. Why is that?"
There's a straightforward answer: token costs.