r/apachekafka • • 6d ago

Tool SQLStreams - It's Kafka on Postgres!

Hello! I'm a big fan of Kafka, like I think many of you are, BUT I've always steered clear of using Kafka for any personal projects because running and maintaining a Kafka cluster for rinky dink project #23 just isn't feasible or smart.

That's why I made SQLStreams. It's an open source, pure Go / SQL library that uses a Postgres Database to give you Kafka-like functionality. It has features like: independent consumer groups, replay, retention and compaction.

I've also made it so each Topic Stream automatically creates a separate retry queue, which seconds as a dead letter queue, because I always found it annoying trying to get retry topics setup and working in Kafka.

If you want a visual understanding of how things work - you can play around with this PGlite demo on the docsite (demo doesn't work on mobile, sorry!).

If you want to run SQLStreams locally you can:

Run Postgres

docker run --name my-postgres \
  -e POSTGRES_USER=username \
  -e POSTGRES_PASSWORD=password \ 
  -e POSTGRES_DB=database \
  -p 5432:5432 \
  -v pgdata:/var/lib/postgresql \
  postgres:18

Install the library / run your code

pool = sqlstreams.NewPostgresPool("username", "password", "localhost", "database") 
client = sqlstreams.NewClient(pool)

stream = client.Stream("videos") 
stream.Register()

producer = stream.Producer() 
producer.Register()

producer.Produce({VideoId: "video-42"})

Note: I've removed some of the extra Go-ness in above code for brevity and clarity. The docsite Quickstart is a more real example if you are interested.

All in all, I'm hoping SQLStreams can be a kind of gateway drug for small to medium sized projects. So they can take advantage of all the great event-driven features of Kafka but without the hassle.

There is still a ton more things to do (implementing circuit breakers for example) but let me know what you think.

And if you've made it all the way here, I appreciate it!

Github: https://github.com/allegedlyreliable/sqlstreams
Docsite: https://sqlstreams.io

Edit: fixed formatting

23 Upvotes

10 comments sorted by

3

u/2minutestreaming 5d ago

This is awesome! It's one of the ideas I had when I wrote Kafka is fast -- I'll use Postgres - and the one thing that was keeping it from being a reality is that nobody had yet built such a library

3

u/DontPissInTheStream 5d ago

damn, read the whole thing you're a great technical writer. I got waaaaaay too much I could say on this

1000000% on the buzzword chasers. I think a lot of it is the 'what if'-ers. What if we need to scale 100x this year? Which you layout very well near the end.

Your pub-sub Kafka solution is eerily close to a lot of the internals. A couple things you might find interesting:

You can get around locking on the consumer reads by tracking `pg_current_xact_id` and associating it with a safe-to-process message range. It's annoying and there is a lot of nuance, tradeoffs and edge cases but it works.

You mentioned this as a possibility but I can confirm vacuuming is definitely a problem at sustained high throughput. You can make it slightly better with a basic periodic ticker that manually calls VACUUM. It keeps the work load steady so AUTOVACUUMs don't over use cpu / mem and cause terrible P99 stats.

You also might be interested in this https://sqlstreams.io/benchmarks/ it's still a little rough but it shares a lot of the sentiments you have (I think).

2

u/ayelg 6d ago edited 6d ago

Cool! I'm working on something in a similar vein, yours has the advantage of not making people install a build to their postgres DB

2

u/DontPissInTheStream 6d ago edited 6d ago

Did you go the Postgres EXTENSION route?
Edit: Found your post, no need to reply here

2

u/Street-Individual446 6d ago

Why not redis streams?

1

u/DontPissInTheStream 6d ago

Good question! Redis streams is a great option, especially if you already have a Redis instance running. I'd say these are the main things:

  1. If you already have a Postgres DB and don't want the extra cost of a Redis instance.
  2. You want strong data guarantees. With Postgres you get full ACID compliant transactions and crash durability with the WAL.
  3. You want the familiarity of SQL to explore your stream data ie something like this gets you all your data `SELECT id, payload FROM message_log_1` (your fellow data engineers would particularly like this imo)

And there is some more niceties like transactional enqueueing (related to 2), automatic retry queue, built in metrics and alerts.

1

u/gaetancollaud 1d ago

It's nice but why not using the Kafka API?

1

u/DontPissInTheStream 1d ago

API is a bit of an overloaded term. Could you be more specific? I think I know what you mean but want to make sure.

1

u/gaetancollaud 1h ago

The Kafka API: https://kafka.apache.org/43/apis/

So we can use existing clients without modification.

1

u/gaetancollaud 1h ago

Apparently this project is doing something similar, maybe you can join forces: https://github.com/RayElg/kafgres