🚀Modal Supports 1M Concurrent Sandboxes, Redefining Scale
Modal's Scalability Breakthrough for AI Workloads
TL;DR
Modal rebuilt its sandbox infrastructure to support 1M concurrent sandboxes and 100K+ workers, pushing Kubernetes' limits. This is a game-changer for teams running large-scale AI workloads.
Modal just rebuilt its sandbox infrastructure to support 1 million concurrent sandboxes and tens of thousands of sandbox creations per second. This is a big deal for anyone running large-scale AI workloads, as it pushes the limits of traditional container orchestration systems like Kubernetes. Modal's approach, which involves making everything horizontally scalable by default and using a fleet of scheduling servers, allows them to handle the load without a centralized coordination system. In their benchmark, they created 1 million sandboxes in under a minute, with median startup-to-code time under 0.5 seconds. This architecture is a significant step forward in scaling AI workloads.

Key Points
Modal's new architecture supports 1 million concurrent sandboxes and tens of thousands of sandbox creations per second.
Load testing shows the system remains viable with over 100,000 workers.
Median sandbox startup-to-code time is under 0.5 seconds in their benchmark.
Modal's approach involves making everything horizontally scalable by default.
The system uses a fleet of scheduling servers operating in parallel to scale.
Why It Matters
If you're running large-scale AI workloads, Modal's new infrastructure is a game-changer. It supports 1 million concurrent sandboxes and tens of thousands of sandbox creations per second, pushing the limits of traditional container orchestration systems. This architecture is particularly useful for teams using GPUs and high-performance computing, as it allows for efficient scaling without the limitations of centralized coordination systems like Kubernetes.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,488 builders reading daily.