Post

AO
Ahead of AI

GPT-6 Astra, Looped Transformers, and Hidden Reasoning

A lot has happened in the last few weeks. I am sure that OpenAI’s GPT-6 Astra is top of mind for everyone right now. In particular, thoughts on its performance, the looped transformer/recurrent depth aspects, and rumors that Astra is “hiding” its reasoning trace (i.e., chain of thought). So, in this article, I want to start with some brief impressions of Astra and some thoughts on where all this is headed. Then, I will discuss, in detail, what “looped transformers” are, and how (or rather, if) this relates to hiding chains of thought. Lastly, after covering the basics of the looped transformer, I wanted to highlight some new insights from recent research papers on the topic. 1. GPT-6 Astra impressions First things first. Before getting into the architecture rumors and related research literature, let me briefly summarize some GPT-6 Astra observations and tidbits. Last week, OpenAI’s new GPT-6 Astra was released with a big fanfare. I used it over the past couple of days, and it’s a

By Sebastian Raschka, PhD
Tweet media