🤖Is OpenAI's Astra a Looped Transformer? Raschka Says No
TL;DR
Sebastian Raschka pushes back on claims that OpenAI's Astra runs recurrent depth, then explains what looped transformers actually do. Reusing a 22-layer stack twice doubles effective depth at no extra storage, costs roughly 2x compute, and keeps about 75% of standard token efficiency.
Sebastian Raschka pushes back on claims that OpenAI's Astra runs recurrent depth, then explains what looped transformers actually do. Reusing a 22-layer stack twice doubles effective depth at no extra storage, costs roughly 2x compute, and keeps about 75% of standard token efficiency.

Key Points
Posted Sept 2, 2026, one day before the Astra rollout coverage peaked
Nanbeige 4.2-3B pretrains on 28T tokens with a 22-layer stack reused twice
Two passes was the best trade-off; more passes barely helped and slowed training
Retains roughly 75% of the token efficiency of a standard architecture
Traces the idea to the NeurIPS Mixture-of-Recursions paper and its per-token router
Why It Matters
Layer reuse buys depth with memory you already have and compute you do not. Worth knowing if you serve models on constrained hardware and can trade latency for quality.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,484 builders reading daily.