Skip to content
modular.com·

🤖Modular Platform Goes Prod with Billions of Tokens Served

Your AI models just got a new home on Modular Cloud

TL;DR

The Modular Platform is now production-ready, serving billions of tokens per minute and supporting AWS Trainium, Google TPUs, Qualcomm Cloud AI 100, and more. MiniMax runs M3 on a dedicated deployment.

Modular Platform has gone live in production, handling billions of tokens per minute for real enterprise deployments like MiniMax's M3 model. Developers can now use the fully open-source Mojo language (Apache 2.0) to run models across various hardware types without rebuilding software stacks. This portability reduces engineering effort by over 10x and supports AWS Trainium, Google TPUs, Qualcomm Cloud AI 100, and more.

Modular Platform Goes Prod with Billions of Tokens Served — modular.com

Key Points

1

Modular Platform supports AWS Trainium, Google TPUs, Qualcomm Cloud AI 100, and more alongside CPUs and GPUs

2

MiniMax runs its M3 model on a dedicated Modular deployment serving billions of tokens per minute

3

Mojo language is now fully open source under the Apache 2.0 license for cross-platform portability

4

Modular Cloud offers shared endpoints with pay-per-token pricing, similar to OpenAI's API

5

Native Windows support coming thanks to collaboration with Microsoft Windows team

Why It Matters

If you're developing AI models and need a versatile platform that supports multiple hardware architectures without the hassle of rebuilding software stacks, Modular Platform is your new best friend. With reduced engineering effort by over 10x, it allows developers to focus on model development rather than infrastructure.

Modular PlatformMojo LanguageAI ModelsHeterogeneous Compute

Frequently Asked Questions

Why does this matter?

If you're developing AI models and need a versatile platform that supports multiple hardware architectures without the hassle of rebuilding software stacks, Modular Platform is your new best friend. With reduced engineering effort by over 10x, it allows developers to focus on model development rather than infrastructure.

What happened?

The Modular Platform is now production-ready, serving billions of tokens per minute and supporting AWS Trainium, Google TPUs, Qualcomm Cloud AI 100, and more. MiniMax runs M3 on a dedicated deployment.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,187 builders reading daily.

Also get