Optimizing Apache Kafka Deployments - Configuration customization is all you need

Swarnabha Das, Sujay Goswami, Subalalitha Chinnaudayar Navaneethakrishnan · 2023

Apache Kafka is an open source publish-subscribe message broker and distributed logging tool that has a widespread adoption in the modern world by companies like Uber, Airbnb, Linkedin and many more. Delivering extreme high volume event data to diverse consumers via a publish-subscribe messaging system with several out of the box optimizations on the 4 major pillars of latency, throughput, durability and availability makes this tool so much popular in the IT landscape. Kafka has the option to customize most of its configuration related parameters to optimize it the most for our use case. Some of the use cases have millions of writes per second, hence it needs to handle huge billions of volumes of data which cannot be done easily by a traditional database or key-value store pair. Some use cases like a chat application demands very low latency where storing the data can be delayed, but the recipient needs to get the sender’s message delivered as soon as possible. Some use cases like a banking application requires high durability over latency and throughput, where the data records must be stored in spite of any system failures. Latencies over 500ms can be tolerated but not data loss in any possible case. And finally some use cases like emergency systems, autonomous vehicles, these are some applications that have to be available for communication all the time and it really doesn’t demand low latency or high throughput or high durability.So the crux of this is, we should know to configure the application depending on our use case. In this paper we are comprehensively going to talk about how we can optimize Apache Kafka correctly depending on these 4 major pillars. After immense testing and brainstorming we have come up with a set of configuration parameters for each of these pillars which would provide a headstart to optimize the kafka deployment. Kafka’s ability to scale efficiently, but it can become a bottleneck if not properly configured. We have benchmarked the results along with those provided by Microsoft and Confluent. Instead of proposing, using graphs and metrics we imply it in a prudent way.

Read the paper · More papers on PaperTik