🚀Netflix Switches to Flink Autoscaler for 30K+ Streaming Jobs
Netflix saves millions with smarter Flink autoscaling
TL;DR
Netflix is moving to Apache Flink Autoscaler for over 30,000 streaming jobs, reducing compute costs by 58% and saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters.
Netflix is switching to the Apache Flink Autoscaler for more than 30,000 streaming jobs across multiple AWS regions, aiming for better cost efficiency. This move follows their previous cluster-level approach, which was less effective for complex, stateful pipelines. One team at Netflix saw a 58% reduction in annualized Flink compute expenditure, saving around $1.1 million annually. The new autoscaler uses job metrics to estimate processing rates and adjust parallelism for individual operators, leading to significant cost savings. Netflix integrated the autoscaler with its internal control plane and modified JobManager metric collection to support jobs with up to 3,000 subtasks.

Key Points
Netflix uses Apache Flink Autoscaler for over 30,000 streaming jobs across multiple AWS regions.
One team at Netflix reduced annualized Flink compute expenditure by 58%, saving $1.1 million annually.
Netflix has run Apache Flink since 2017 and built its first autoscaler around 2019.
The new autoscaler uses job metrics to estimate processing rates and adjust parallelism for individual operators.
Netflix modified JobManager metric collection to support jobs with up to 3,000 subtasks.
Why It Matters
If you're managing complex, stateful pipelines with Apache Flink, Netflix's new autoscaler approach could save you millions. One team at Netflix saw a 58% reduction in Flink compute costs, saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters, making it a game-changer for resource-intensive workloads.
Frequently Asked Questions
Why does this matter?
If you're managing complex, stateful pipelines with Apache Flink, Netflix's new autoscaler approach could save you millions. One team at Netflix saw a 58% reduction in Flink compute costs, saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters, making it a game-changer for resource-intensive workloads.
What happened?
Netflix is moving to Apache Flink Autoscaler for over 30,000 streaming jobs, reducing compute costs by 58% and saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,462 builders reading daily.