Skip to content
GitHub·

🤖Kimi K3 Runs Full 2.8T Model on MacBook Pro

Running a 2.8T model on consumer hardware? Spoiler: it's possible

TL;DR

ARGODRIVE's benchmark shows Kimi K3 running a full 2.8T model on a MacBook Pro. Steady decode rates vary, with 1.00 tok/s for 512-token answers and 0.96 tok/s for 17-token prompts. The catch? It's slow for large prompts, taking 6.3 minutes for the first token.

ARGODRIVE's benchmark report reveals the Kimi K3 model running a full 2.8T parameter model on a single M5 Max MacBook Pro. Steady decode rates for 512-token answers are 1.00 tok/s, while 17-token prompts see 0.96 tok/s. The bottleneck? Prefill re-reads each layer's experts 8×, causing slow prefill times. For 512-token prompts, the first token takes 6.3 minutes. This matters for anyone pushing the limits of consumer hardware with large models. The report details the performance on one, two, and three drives, showing diminishing returns beyond three drives. The full model, with no pruning, runs on 128 GB of RAM, taking 6.635 GiB on disk and 4.49 GiB at runtime.

Kimi K3 Runs Full 2.8T Model on MacBook Pro — GitHub

Key Points

1

ARGODRIVE's benchmark shows Kimi K3 running a full 2.8T model on a single M5 Max MacBook Pro, with 128 GB of RAM.

2

Steady decode rate for 512-token answers is 1.00 tok/s, while 17-token prompts see 0.96 tok/s.

3

The slowest of each layer's 16 reads sets the pace, not total bandwidth, causing prefill times to be slow.

4

For 512-token prompts, the first token takes 6.3 minutes, highlighting the bottleneck in consumer hardware.

5

The full model runs on 128 GB of RAM, taking 6.635 GiB on disk and 4.49 GiB at runtime, with no pruning.

Why It Matters

If you're pushing the limits of consumer hardware with large models, Kimi K3's performance on a MacBook Pro shows the trade-offs. The full 2.8T model runs on 128 GB of RAM, taking 6.635 GiB on disk and 4.49 GiB at runtime, with no pruning. The steady decode rates vary, with 1.00 tok/s for 512-token answers and 0.96 tok/s for 17-token prompts. But the 6.3-minute wait for the first token of a 512-token prompt highlights the bottleneck in consumer hardware.

Kimi K3ARGODRIVEMacBook Prolarge modelsconsumer hardware

Frequently Asked Questions

Why does this matter?

If you're pushing the limits of consumer hardware with large models, Kimi K3's performance on a MacBook Pro shows the trade-offs. The full 2.8T model runs on 128 GB of RAM, taking 6.635 GiB on disk and 4.49 GiB at runtime, with no pruning. The steady decode rates vary, with 1.00 tok/s for 512-token answers and 0.96 tok/s for 17-token prompts. But the 6.3-minute wait for the first token of a 512-token prompt highlights the bottleneck in consumer hardware.

What happened?

ARGODRIVE's benchmark shows Kimi K3 running a full 2.8T model on a MacBook Pro. Steady decode rates vary, with 1.00 tok/s for 512-token answers and 0.96 tok/s for 17-token prompts. The catch? It's slow for large prompts, taking 6.3 minutes for the first token.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,470 builders reading daily.

Also get