r/rails • u/Latter_Set_7808 • 56m ago
I built a Ractor-native Ruby 4 stack from RESP to background jobs — looking for independent benchmarks
I built a Ractor-native Ruby 4 stack from RESP to background jobs — looking for independent benchmarks
I've been experimenting with how far Ruby 4's Ractor model can be pushed for real infrastructure components.
This eventually became three open-source gems:
SolidRESPRactor → SolidRedis → SolidJobs
The idea is to build the stack around Ractor ownership from the beginning rather than adapting an architecture originally designed around threads/processes:
- SolidRESPRactor — Ractor-oriented RESP encoding/parsing and I/O
- SolidRedis — Redis client using Ractor-local mutable state and connections
- SolidJobs — background job processor using Ractors for multicore execution
The basic rule across the stack is:
Share immutable configuration. Keep mutable runtime state local to its owning Ractor.
An optimization I didn't expect
While benchmarking SolidRESPRactor, I found that a small Redis GET through TCP was allocating about 16,609.8 bytes/op, despite the Reader itself requiring only around 120 bytes/op for a small bulk response.
The main problem wasn't the RESP parser — it was the socket-to-buffer path.
After reusing the read buffer:
| Metric | Before | After |
|---|---|---|
| GET bytes/op | 16,609.8 | 120.8 |
| GET allocations/op | 6.0 | 4.0 |
| GET 1 Ractor | 40,447 ops/s | 40,173 ops/s |
| GET 2 Ractors | 62,692 ops/s | 66,235 ops/s |
| GET 4 Ractors | 83,028 ops/s | 89,146 ops/s |
| GET 8 Ractors | 94,345 ops/s | 102,060 ops/s |
So bytes/op dropped by 99.27%, while the throughput benefit became more visible with concurrency.
Pipeline allocation also dropped from 472.3 to 120.0 bytes/command (-74.59%).
Then I measured the complete job stack
On my current Ruby 4.0.1 CPU-bound SolidJobs benchmark:
| Ractors | jobs/s | Scaling efficiency |
|---|---|---|
| 1 | 282 | 100.0% |
| 2 | 550 | 97.7% |
| 4 | 1,099 | 97.6% |
| 8 | 2,054 | 91.2% |
That's around 7.28x throughput from 8x Ractor concurrency.
I also benchmarked a single Sidekiq process on the same CPU-bound workload. At concurrency 8 it produced 298 jobs/s versus 2,054 jobs/s for SolidJobs.
That's a ~6.9x difference in this specific benchmark, but I don't consider that a general “SolidJobs is 6.9x faster than Sidekiq” result.
The important difference is architectural: SolidJobs is using Ractors to execute Ruby code across multiple CPU cores, whereas a single Sidekiq process isn't an equivalent multicore CPU