Skip to content
InfoQ·

🚀Netflix Switches to Flink Autoscaler for 30K+ Streaming Jobs

Netflix saves millions with smarter Flink autoscaling

TL;DR

Netflix is moving to Apache Flink Autoscaler for over 30,000 streaming jobs, reducing compute costs by 58% and saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters.

Netflix is switching to the Apache Flink Autoscaler for more than 30,000 streaming jobs across multiple AWS regions, aiming for better cost efficiency. This move follows their previous cluster-level approach, which was less effective for complex, stateful pipelines. One team at Netflix saw a 58% reduction in annualized Flink compute expenditure, saving around $1.1 million annually. The new autoscaler uses job metrics to estimate processing rates and adjust parallelism for individual operators, leading to significant cost savings. Netflix integrated the autoscaler with its internal control plane and modified JobManager metric collection to support jobs with up to 3,000 subtasks.

Netflix Switches to Flink Autoscaler for 30K+ Streaming Jobs — InfoQ

Key Points

1

Netflix uses Apache Flink Autoscaler for over 30,000 streaming jobs across multiple AWS regions.

2

One team at Netflix reduced annualized Flink compute expenditure by 58%, saving $1.1 million annually.

3

Netflix has run Apache Flink since 2017 and built its first autoscaler around 2019.

4

The new autoscaler uses job metrics to estimate processing rates and adjust parallelism for individual operators.

5

Netflix modified JobManager metric collection to support jobs with up to 3,000 subtasks.

Why It Matters

If you're managing complex, stateful pipelines with Apache Flink, Netflix's new autoscaler approach could save you millions. One team at Netflix saw a 58% reduction in Flink compute costs, saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters, making it a game-changer for resource-intensive workloads.

Apache Flinkautoscalingcost optimizationNetflixopen-source

Frequently Asked Questions

Why does this matter?

If you're managing complex, stateful pipelines with Apache Flink, Netflix's new autoscaler approach could save you millions. One team at Netflix saw a 58% reduction in Flink compute costs, saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters, making it a game-changer for resource-intensive workloads.

What happened?

Netflix is moving to Apache Flink Autoscaler for over 30,000 streaming jobs, reducing compute costs by 58% and saving $1.1 million annually. The new autoscaler optimizes individual operators, not just clusters.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,462 builders reading daily.

Also get