Skip to content
daily-hour-news·

🤖Is OpenAI's Astra a Looped Transformer? Raschka Says No

TL;DR

Sebastian Raschka pushes back on claims that OpenAI's Astra runs recurrent depth, then explains what looped transformers actually do. Reusing a 22-layer stack twice doubles effective depth at no extra storage, costs roughly 2x compute, and keeps about 75% of standard token efficiency.

Sebastian Raschka pushes back on claims that OpenAI's Astra runs recurrent depth, then explains what looped transformers actually do. Reusing a 22-layer stack twice doubles effective depth at no extra storage, costs roughly 2x compute, and keeps about 75% of standard token efficiency.

Is OpenAI's Astra a Looped Transformer? Raschka Says No — daily-hour-news

Key Points

1

Posted Sept 2, 2026, one day before the Astra rollout coverage peaked

2

Nanbeige 4.2-3B pretrains on 28T tokens with a 22-layer stack reused twice

3

Two passes was the best trade-off; more passes barely helped and slowed training

4

Retains roughly 75% of the token efficiency of a standard architecture

5

Traces the idea to the NeurIPS Mixture-of-Recursions paper and its per-token router

Why It Matters

Layer reuse buys depth with memory you already have and compute you do not. Worth knowing if you serve models on constrained hardware and can trade latency for quality.

Quick Facts

Sebastian RaschkaOpenAI Astralooped transformersmodel architectureNanbeigeinference

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,484 builders reading daily.

Also get