🔍Pinterest Manas Handles Tens of Billions of Embeddings
Pinterest's Manas powers discovery with massive efficiency gains
TL;DR
Pinterest's Manas platform now handles tens of billions of embeddings using Scalar Quantization and Product Quantization, cutting costs and improving performance. Key: PQ reduces HNSW indices by 93%, SQ by 75%, and SPANN doubles QPS.
Pinterest's Manas platform now handles tens of billions of embeddings using Scalar Quantization (SQ) and Product Quantization (PQ), significantly reducing memory and compute costs. This impacts developers using vector search algorithms, as PQ and SQ slash index sizes by 74% and 93% respectively, while maintaining recall rates of 70-95%. SPANN, an SSD-based serving solution, doubles query performance compared to DiskANN. The baseline HNSW index is 121GB, dropping to 32GB with PQ and 50GB with SQ, achieving 77-93% recall. SPANN with PQ saves 40% CPU time for production queries, making it a game-changer for large-scale vector search.

Key Points
Pinterest's Manas handles tens of billions of embeddings using PQ and SQ, reducing HNSW indices by 74% and 93%.
PQ reduces HNSW index size to 32GB with 77.25% recall, while SQ drops it to 50GB with 92.92% recall.
SPANN with PQ achieves 3x QPS of DiskANN with 1/3 latency and a minor 5% recall drop, optimizing SSD serving.
SPANN saves over 40% CPU time for production queries compared to full in-memory HNSW, cutting costs.
Pinterest is piloting multi-vector Late Interaction models like ColBERT for fine-grained relevance match.
Why It Matters
If you're handling large-scale vector search, Manas's PQ and SQ reduce index sizes by 74% and 93%, respectively, while maintaining high recall. SPANN doubles QPS, saving 40% CPU time. This is a must-watch for teams scaling vector search in production.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,482 builders reading daily.