Confluent Developer ft. Tim Berglund, Adi Polak & Viktor Gamov

10 Years of Kafka Streams with Matthias J. Sax | Ep. 32

Confluent Season 2 Episode 32

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 29:41

Tim Berglund talks to Matthias J. Sax (Confluent) about 10 whole years of Kafka Streams! Matthias’ first job: electrician-in-training on BMW’s assembly lines. His challenge: reflecting on 10 years of Kafka Streams growth, major milestones, and what comes next.

SEASON 2
Hosted by Tim Berglund, Adi Polak and Viktor Gamov
Produced and Edited by Noelle Gallagher, Peter Furia and Nurie Mohamed
Music by Coastal Kites 
Artwork by Phil Vo 

  •  🎧 Subscribe to Confluent Developer wherever you listen to podcasts. 
  • ▶️ Subscribe on YouTube, and hit the 🔔 to catch new episodes.
  • 👍 If you enjoyed this, please leave us a rating. 
  • 🎧 Confluent also has a podcast for tech leaders: "Life Is But A Stream" hosted by our friend, Joseph Morais.
SPEAKER_02

And so Kafka Streams is running these, you know, very, very mission-critical workloads for certain companies. The main limitation to topology test driver is that it mimics uh input topic with a thingy partition. It's a mountain to climb, but if you really want it, you can get it.

SPEAKER_00

Hey Tim Berlin here. Welcome to another episode of the Confluent Developer Podcast. Kafka Streams turns 10 years old this year, that's 2026 for the kids at home. And who better to celebrate its birthday with us than Matias Sachs? Matias has been a key contributor this whole time, and he's a great guy besides. He and I get to talk about how the project has grown over that time in the last 10 years, who some of the big community influences have been, some key milestones like exactly once, and of course the recent consumer group rebalance changes. And finally, where Matias thinks Streams is going looking forward. Let's join that conversation already in progress. I want to say it's May uh 13th, uh 2026, and something very special is going on right about now. Kafka Streams is turning 10. So who better to talk to than Matthias Sachs?

SPEAKER_02

Well, I would say as a broad community. I mean, I I did a couple of things, but obviously it's always a village, not a single person doing things like that.

SPEAKER_00

And that's uh a key point. Um this is an open source project, and there are a lot of people who've contributed to it. And I I kind of want to ask you about that. But um what what's it like here on the uh the the 10th birthday of this thing that you've you've given a you know a good chunk of your career to?

SPEAKER_02

Yeah, it's very exciting. I mean, I actually wasn't even realizing it for a long time, and a couple of weeks back, it's kind of oh, wait, it's going to be 10. This is great. We need to celebrate. And yeah, it's super exciting. I mean, when when we started this project, I would not have thought that we might you know come this far. Um and also the growth of the community in general and the customers and use cases we see. Um, it's just been a wild ride. And I also learned a ton about distributed systems and stream processing. Um, especially at Confluent, we always had a very strong engineering team. Um, so that was also you know a real, real good learning experience for myself. And and now we also have a good team. Confluent is investing a lot in Kafka streams nowadays. We grew the team in the last year, we basically doubled it. Um we are we are going full theme.

SPEAKER_00

Okay, I love that. Um I think it's good that you say that because uh people have asked questions, right? Um, you know, now that Confluent, this is more a couple years ago, but people were saying, well, Confluent has a managed Flink service. Is Kafka Streams really gonna continue to be a thing? And it really, really is, is the answer. Yeah, absolutely. Um what what uh you know, with with a a project like this, this is uh this is Apache Kafka, right? It's a part of Apache Kafka. It's a a project of the Apache Software Foundation. Um and you know, you and I have confluent coworkers who are engineers who get paid to work on it. Um but it's not confluence, you know, it's not like it's a thing that we do, it's a thing that the ASF does, uh that we're just a part of. Um talk about, you know, who who would you celebrate or what kind of contributions would you celebrate that come from the community, whether whether Confluent or not, like especially people outside of Confluent. Um what are what are some of the cool things you've seen?

SPEAKER_02

Yeah, uh there's so many. Um very, very hard to pick. Um one thing top of mind that is something we started in the last two, three years is uh we we collaborated much more closely with Michelin. Um so Michelin is a very heavy Kafka streams user, and they actually built some scaffolding around Kafka streams to make it easier internally to use. And they open sourced this as an own standalone open source thing on GitHub. And we learned about it and we said, hey, this is great stuff, but this should not be outside of Kafka Streams. We want to have this in Kafka streams. And now they're contributing kips um and we are we are merging these things in, right? Um so one example is improved exception handling um that landed already. Um we also added a dead letter queue that was also part of their project that is now a feature in Kafka streams. Um at the moment they have a new kip um to extend the topology test driver. Um that is also very exciting. Um, so as the main limitation in the topology test driver is that it mimics an input topic with a single partition. So you can't really test multi-partition setups, your keying strategies and stuff like that. And so they propose to make the test driver multi-partition aware.

SPEAKER_00

It might be obvious just from the name, but um for anybody who's not an active Kafka streams user, uh tell us what topology test driver is.

SPEAKER_02

It's basically uh a unit test utility um that mimics a Kafka streams runtime. Where you say, well, I have my Kafka streams topology. That is where the name comes from. Um and well, you have a test driver. So you basically say I create input topics and I create output topics, and then I send data into the input topics, then I can get the data of the output topics, then can I make my assertions on top of that, right? And test my business. Exactly, exactly. Um and otherwise it's very hard to test, right? This with having this external tool. And then before that, people would basically spin up like test containers or whatever, right? Um and you know, run against the real Kafka cluster very heavyweight, and now it's a unit test utility with millisecond latency, right? Basically, you pipe it through and you immediately so you have very, very quick testing.

SPEAKER_00

That is what we want. I love it. Um yeah, I've heard of that limitation before. This has been a thing that has has pushed people into test containers. Exactly. Which is good. I mean, you you want a mix of both, right? It's it's integration tests and unit tests and everything. That's a that's a old distinction that is still a very valid one. Um but you really want stuff that will uh uh run quickly to give you to give you uh good good feedback. Yeah, yeah, yeah, absolutely, absolutely. Of course, um to give I should say to give the the the AI agent good feedback on uh whether it's its code is is correct. Yes, absolutely, yeah.

SPEAKER_02

Yeah, yeah, that's that's one of the things. I mean, what what else came from the community? As I said, especially in the last I would say 12 months, we saw actually a huge spike in community contributions, and at the moment you're a little bit overwhelmed to review all the kips and you know trying to catch up.

SPEAKER_00

Um but it's about it because why has there been a huge spike in contributions?

SPEAKER_02

Uh I don't know. Maybe there is some agent helping. I don't know.

SPEAKER_00

Maybe those are AI PRs. Uh I and I I don't want to sidetrack you on that, but I mean I I think it'll be a question folks are are interested in. Um as a reviewer of those PRs, have you seen I mean you're seeing volume and uh you know all things being equal, that's good. That's engagement from the community. Is it AI slop? Is it are they bad? Are they not thought out? I mean, it are they what what's the quality like uh in the face of the volume?

SPEAKER_02

Yeah, so far I wouldn't say it's negative. I mean, we have seen one or two PRs, or I personally have seen one or two PRs, or maybe a little bit more. Very clearly, like AI slop, it's kind of somebody changed 10,000 files. It's just like okay, this is not gonna happen, right? Um but but but overall, not a problem. I mean, uh from the things I see, I could not even necessarily tell this is AI generated code. Um so quality, quality is good. And I mean we keep the review high, so it's it's fine.

SPEAKER_00

Good, good. That is really good to hear. 10,000 files, yeah. Yeah. And and your comment is could you rebase that into one commit?

SPEAKER_02

Yeah, something like that. Still not merging. Yeah, yeah, yeah. Also, some PRs where it's totally unclear what the PR actually does. There's no PR description, it just pushes code. What is this about? Um, and then there's also no response from the contributor, and we let it sit for a couple of weeks, and then yeah, we just close it.

SPEAKER_00

It's yeah. Yeah, yeah. Well, I mean, that's open source, right? It's exactly you're a high, extremely high profile open source project. Uh there's going to be a distribution, probably a power law distribution of quality there, uh, with a few really good ones and a lot of a lot of not good ones. Yeah, yeah, yeah. Or maybe it's normal, I don't know, but that's that's um that's that's life. And you know, it's important, I guess it's a good note to take that when you're contributing to an open source project. Be thoughtful. You know? Yeah, yeah, yeah, yeah. Prepare uh uh like like it's like you cooked dessert for somebody, you know. Here, I made this for you. You would like it to be presented nicely and be understandable and all those things. That's that's that's pretty important.

SPEAKER_02

Yeah, yeah. And and overall, that's that works in general, right? For example, we we did uh did a larger kip, actually, two kips that are shipping with the upcoming 4.3 release to add header support for state stores in Kafka streams. And this was a huge project, and actually, I think four or five different people from the broader community jumped on this project because this was driven by Confluent, right? And we put a couple of engineers on it, and we could, you know, give them a lot of work and they helped us, and it was a great collaboration.

SPEAKER_00

Nice. You said header support in state stores. Could you expand on that a little bit?

SPEAKER_02

Yeah, so so Kafka Streams uses RocksDB to store state, right, for more complex computations. And very historically, this is a plain key value store. Then very early on, we did a larger project where we said, well, we also want to be able to store a timestamp to you know get better event time semantics and processing guarantees. Um that's already landed, I don't know, six years ago. So it's also not like you anymore, seven years ago.

SPEAKER_00

That's a Kafka message, key value, and timestamp.

SPEAKER_02

Timestamp, exactly. And now we also want to get headers into that. I mean, that was a long, long-standing issue. A lot of people ask for this for a long time. We never had time to do it, and now it got a priority at Confluent. And now we have these new stores where you can also put the headers in and Kafka streams stores for you, and you can modify them, and yeah, it's it's it's great. It's still not completely finished. Uh still need to do a couple of follow-up uh kips um to make it really feature complete. But it ships with uh upcoming 4.3 release, and yeah, um takes us closer to to how Kafka works.

SPEAKER_00

The API abstraction there in the state store is I am storing a value and a timestamp and a collection of headers, and maybe that gets munged into the value in RocksDB, but like I don't I don't need to think about that. I I I see those separate things.

SPEAKER_02

Exactly, exactly, exactly. That is exactly how it works.

SPEAKER_00

Nice. Yeah, better than having to fuss with all that myself.

SPEAKER_02

Yeah, yeah, exactly. And some people did it, right? I mean, especially when you fall back to the processor API, you can do all these things yourself because you can define your custom value type, right? And if you want to add headers, you can add headers to it. But it's obviously not pleasant than a lot of like, you know, boilerplate code and scaffolding you don't want to feel this.

unknown

Right.

SPEAKER_00

Because process the processor API life is is a lot more you're you're probably thinking about the state store access more. Now a quick word from our sponsor. Confluent developer the podcast is brought to you by Confluent Developer the website, which has everything you need as a developer of data streaming systems. And it's completely free. We've got curriculum, hands-on exercises, executive tutorials, the online data streaming engineer certification, also free. A way to find a meetup near you, those are free. Everything is there. I really want you to be successful in your journey as a data streaming engineer, and this is the site that has Witch Data. Check it out at developer.confluent.io. That's developer.confluent.io. Now back to the show. And let me um I want to set that up a little bit, because again, we're just talking like everybody knows Kafka streams really well. And if you're listening to this, you may not, and I want you to be able to learn from Matias if you don't. So uh there are and you I'm gonna say some stuff and you correct me if I'm wrong or just tell me I'm right. Um there are two ways to use Kafka streams. There is one that's called the DSL, which is uh a fluid API fluent API where uh you know you start off by creating a stream out of a topic, and then maybe you filter it and you uh group it by some field and compute an aggregation and all that kind of stuff. And these are these are like method calls uh that you're passing lambdas into and it's super readable and and beautiful and wonderful. And some of those might be stateful, and so they're using the state store, but you don't have to worry your little self about that. That all just gets done. Um and then you spit out the results into an output topic. Then there's the processor API, which is a uh the the abstraction is here's another message that just came in, and here's your state store. Um you you do what you want. Um which sounds bad, right? But I I number one, you know, would you correct anything that I just said? And number two, how has that worked out? Like it would you have a sense of what Kafka Streams developers prefer between those two things?

SPEAKER_02

Yeah, no, everything you said is absolutely correct. Um and both are very, very valuable. Um it depends a little bit what kind of you know applications they're building, but we know of a couple of actually users who explicitly only use the processor API because they really want to have full control of everything. And the the the DSL, while it's like super powerful, um, is obviously opinionated. And once in a while people say, well, you know, I I don't like that. And I I want to do want to do roll my own, right? It's more work. Um but but that is why we have it. We actually also allow you to integrate both into each other. So if you write a DSL program, you can always plug in a custom processor in the middle and say, Okay, now I roll my own for this particular step. Um, it's also what a lot of people use, but but some people basically don't touch the DSL at all and just you know use the lower level one because they want to be in full control, as you know, developers very often like to do.

SPEAKER_00

Yeah. We've been burned enough. We know, like, okay, this high-level abstraction looks nice, but I know what's gonna happen. Let me just get into the pain and do the low-level thing.

SPEAKER_02

Yeah, exactly, exactly. And I mean the DSL is to some extent inspired a little bit by SQL semantics and stuff like that. And I think it's very good for data pipelines, and there's a lot of use cases using it, right? Um, but once in a while, it's just like because it's stream processing, you also need to have a couple of you know, audit in it to bridge between the nature of stream processing and SQL semantics, and then things fall apart once in a while for some use cases, and then people are not so happy. Um was actually interesting. A couple of weeks back, there was a discussion, I think, on LinkedIn, about exactly these topics. And I think there is, I made a strong case that the DSL is useful, and especially for one operator that you do not want to roll by your own, and that's uh foreign key join.

SPEAKER_01

Ah, okay.

SPEAKER_02

Because that is very, very complicated, and I would not recommend anybody to try to re-implement that. So if you need to have that, I would always say go with the DSL. For everything else, if you don't like what the DSL does, I think it's reasonable enough to roll your own. Um, but the foreign keychain is something I I I would die on that hill.

SPEAKER_00

Yeah, yeah. Um yeah, that sounds like a potentially unpleasant thing to do. Um because you get yeah, time you get windows to contend with and everything. I just I wouldn't I wouldn't want to do that. So looking back over 10 years, what are the big milestones in the development of Kafka Streams? I I can think of a couple big things, but what are the things that stick out to you? Because you've been there the whole time.

SPEAKER_02

Yeah, I mean, obviously the first big milestone was uh was the O1000 release. The Kafka Streams saw the light of the boat. Um and yeah, I think it was interesting because I think the first releases were actually pretty much like still prototypes. Um personally, I think the first actual production ready release was O102.1.

SPEAKER_00

Okay.

SPEAKER_02

Um where we fixed a couple of more bug fix bugs in the bug fix release. So I'm very specific about the version.

SPEAKER_00

Yes.

SPEAKER_02

Uh yes, or maybe 2017, early 17, I can't remember exactly. Um, where we were basically now saying, okay, we are we are now in a in a position to yeah, to to really recommend it for production use cases. Um obviously people put into production before, but many of them were actually then stateless applications, which are much easier. And for stateless applications, Kafka streams was production ready early on, but obviously that's that's the boring part.

SPEAKER_00

Right. But it's not it's not where your investment goes.

SPEAKER_02

No, exactly. I mean you start somewhere, right? Um yeah, and then well, I think the the the next big milestone was obviously exactly once processing that came with the 011 release.

SPEAKER_00

Spring of uh 17. I remember that. That was that was kind of right when I was joining the community.

SPEAKER_02

Yeah, yeah, exactly. Um where I mean there was basically two main kips for Kafka itself, uh adding the the item potent producer and then adding these Kafka transactions. Um and then we did a dedicated kit for Kafka Streams to also use these transactions inside the Kafka Streams runtime. Um yeah, uh interestingly, that is why I want to stretch it, because Kafka Streams is kind of a client, right? So it was always kind of, you know, well, how does it fit into the whole ecosystem? But this was interesting because transactions were actually designed not just for Kafka, but actually for Kafka streams in particular. So the way how transactions work is really tailored to how Kafka streams would use them and how they they work with Kafka streams. Um and that also shows the importance of Kafka streams in the ecosystem because it's not that you know Kafka brokers doing whatever they do, but actually Kafka brokers are also changed for Kafka streams to make sure the Kafka streams works the way it is supposed to be. So so um and that was the first like major kip where we where we get got changes in the broker that really are tailored for Kafka streams.

SPEAKER_00

Gotcha, gotcha. And there was another one recently, right?

SPEAKER_02

With uh there was there was exactly once and um Yeah, the latest one is KIP 1071 with the streams rebalance protocol. Yeah, I mean this is uh this is a huge milestone. Um maybe we can talk about this at the end um because that takes a little longer.

unknown

Okay.

SPEAKER_02

Um yes, then the other part was was the was the new incremental rebalancing protocol. Um so it was actually a pain for everybody anyway, for plain consumers as well, but in particular for Kafka streams. But but with this protocol, we we could you know build more advanced features um like warm-up tasks and high availability assigner. Um and then this this served Kafka streams very well to make the assignment more smooth. Um so this is the main problem, is really if you have a consumer application, you assign partitions, right? And reassigning a petition from consumer A to consumer B is pretty much straightforward. You tell the one consumer, hey, stop, they commit the offsets, you get the partition back, and you give it to the new one. In Kafka streams, it's a little bit more complicated and a little bit more costly because we have state. So if you say we move a task from one instance to a different instance, then the new instance says, uh, yeah, I don't have the state store yet. So we assign it. The first thing they do is they go to the changelog topic and and and reread the data to rebuild the state, and then they can continue processing. And that's a time-consuming process, right? And so the idea of the high availability assigner was to change this behavior and say, well, instead of revoking the task immediately, we first put what we call a warm-up on the target. And then this warm-up task would preheat the state store. And when it's ready and the state is ready, then we do another rebalance, and then we say, okay, now you revoke it, and then we give it to the other guy, and the other guy is already ready to take on processing.

SPEAKER_00

Um messages can be processed immediately.

SPEAKER_02

Exactly, exactly. So you have obviously a latency now when you say task A goes from one member to a different member in the group because it needs to uh uh uh uh uh preheat the state. Um but you have zero downtime, and that's a great thing, right? Um but but this was as complicated, and yeah, and all these changes in the consumer protocol were very necessary to build these things. Um interestingly, this is also kind of I would say almost overstretched the capabilities of the protocol. Um I guess now we can also say. Weighing into kit ten seventy one, where you basically say, well, this doesn't work anymore. There's too many problems. Everything gets redesigned and it gets gets gets done more naturally.

SPEAKER_00

There you go. Who uh Yeah, can you talk about like uh who's using it? And I know sometimes in your position and in my position, we're you're focused on building it, I'm focused on having a team talk about it and and popularize it and and teach people how to get started. And and sometimes we lose sight of of use cases uh and people building real things. But um Does anybody come to mind that is somebody doing something awesome with it?

SPEAKER_02

I mean, in general, Kafka streams is used across all industries, what is actually a nice thing. I mean it's a Kafka thing, right? Kafka is not you know industry specifically. Um and we we we work with telecommunication providers, right? And you know, logistics, bank insurance is everywhere. Um maybe one of the cooler examples is maybe the metronome. Um so the you know SaaS companies that is doing all this kind of online billing as a service, right? Um and and they are heavy Kafka streams users. Um customer actually, so we we know a little bit what they're doing, and uh in a and a large portion of the workload is actually fully powered by Kafka streams. And and this is pretty cool because I mean they do billing, right? And they do live billing, and so Kafka Streams is running these, you know, very, very mission-critical workloads for certain companies.

SPEAKER_00

Uh for a lot, a lot of SaaS billing. I mean, their metronome's got a huge footprint. Yeah. Yeah, yeah, absolutely, absolutely. Yeah, yeah. A lot of a lot of our SaaS life is uh accounted for through Kafka streams. That's pretty cool.

SPEAKER_02

Yeah, yeah, yeah. Yeah. And we actually have a couple of customers where I mean we track their confluent cloud, like network traffic, right? And we can attribute it to different clients where they have like 50, 60, 70, 80 percent of traffic coming from Kafka streams. So these companies are entirely powered by Kafka streams, and and that's obviously great to see. So yeah.

SPEAKER_00

Yeah, no, you have to love that. What's the future got? What's what's on the roadmap? What can people look forward to? I mean, uh do you feel like it's mature? Do you feel like it's you you still so much more to do?

SPEAKER_02

We're oh, there's always much more to do, obviously. I mean, first of all, software's never done, right?

SPEAKER_01

No, really.

SPEAKER_02

Um but but no, we have we have a couple of couple of interesting things. First of all, um we need to we need to finish the KIP 1071 story, right? So we we did ship it and we called it GA with the Petchikafka 4.2 release. But it's not feature complete compared to the classic rebalance protocol. So we're working on that one. Um at the moment, in particular, on the on the high availability assignar story, and then we have a couple of other gaps with like static group members that are not supported yet, and and then a couple of smaller things. Um so this will keep us busy for a couple of months for sure. Um we also integrate much more into the broker. This is a general client story for better observability. So a couple of years back, there was this KIP 714 allowing to send metrics to the broker for centralized metric collection. That is something we later extended also for Kafka streams. And now we are we are extending this story, and the next step would be to allow to monitor configurations of clients. So there's a kip already out that was published a couple of weeks back, um, and it will cover the regular consumer and producer clients, but also Kafka streams, so that they can send over their configurations, and then an operator can basically monitor these applications. And if something sticks out, they can find out which team is it and say, hey, you cannot run with this configuration. It's bad, you need to change it.

SPEAKER_00

Right, right, right. This is a part of the there's a general trend in the last, I don't know, five years of changes in core Kafka that make it more cloud native. So if you're operating a cloud service, then you know you've got some number of and then this can be true of an of a platform team with at a large organization with lots of people building Kafka Streams applications. So it's not just if you're a cloud service, but you know, if you sort of don't have as much control as you might like over your clients, um getting that information into the broker where the person who owns the cluster can see it is uh that's just happened in various ways over the last year.

SPEAKER_02

Yeah, exactly, exactly. And say you can see the age of Kafka in general. I mean Kafka is 15 years old now, I don't know.

SPEAKER_01

Yeah, right.

SPEAKER_02

Back in the days, cloud was not really a thing, you know, and it was was developed for a specific use case at LinkedIn, and there was always this thing about like the sick client riding Kafka. Clients have a lot of logic. And now with with modern cloud development, it's kind of yeah, it's still valuable, but we we need to pull in more logic into the broker. That's it's it's just just better nowadays because the architecture changed.

SPEAKER_00

Right, right. Um, that'll be an interesting needle to thread as we go forward because some of these things are obvious, like you know, uh being able to observe what a client is doing. Yeah. Historically, it just wasn't available to the broker. Now we're gonna pull that in because we just might need that in order for the world to work. But the dumbness of the broker has been, I think, a reason for Kafka's success.

SPEAKER_02

I totally agree.

SPEAKER_00

And so it's it's it's a there's a temptation to make it a lot smarter, and uh, that'll be uh fun to see how the community navigates that over the next five years.

SPEAKER_02

Absolutely, absolutely. It's it's a fine line to walk, and I also don't think we should we should pull in everything. But um, I mean, for example, with the rebalancing and KIP 8 for 8 and 1071, we pull in rebalance logic into the broker, and this is a huge benefit, we know that already. Um and but yeah, we we need to be careful not to not to overshoot, absolutely.

SPEAKER_00

Right, right. But yeah, a lot of good things ahead, uh, an amazing decade behind us and um a lot to look forward to.

SPEAKER_02

Yeah, absolutely. And I mean, with the community being on full steam, we we just need to make sure to, you know, to to get all these kips reviewed and merged. And personally, I always hope that somebody sticks along, uh sticks around long enough that we can we can give some committership. Um, because that's something that's personally what I always find a little bit annoying. Um Confluent is the only big player that invests heavily in Kafka streams. And all the active committers work at Confluent. And I think it's not healthy for the community, but we we were not lucky enough so far to find somebody who is not at Confluent, or maybe then we hire them later. That also happened. Um, you know, to give them committership and focus on Kafka streams, I think that that would help significantly the broader community. Because at Confluent, we always have our own priorities, and there's a lot of good stuff coming from the community, and very often we don't have time then to do it. Um so if you're a contributor and you know, if you say, hey, I want to do this, you know, yeah, let us know. Um we would be happy to work with you. It's a it's a mountain to climb, but um, if you really want it, you can you can get it.

SPEAKER_00

My guest today has been Matthias Sax. Matthias, thanks for being a part of the Complex Developer Podcast. Thanks for having me.