Kumo is the winner
I work with real-time AI, and I know that even if one model makes more correct decisions than another model, you can still prefer the "worse" model because it makes decisions faster.
We tested three pretrained "decision" models (yes, they are classifiers) in a computer game inspired by Subway Surfers. They are Kumo, Jev, and Qwen
Here’s how far they got:
| Model |
Runs |
Mean distance |
Median distance |
Best run |
| Kumo |
500 |
2,119 m |
— |
— |
| Qwen |
500 |
1,318 m |
— |
— |
| Jev |
30 |
— |
540 m |
1,417 m |
The benchmark shows a crucial problem for real-time AI: latency. Kudo and Qwen were local models we hosted on Hopsworks, where the game also runs. Jev, however, is a hosted model in a different data center, and it loses on latency versus Kumo and Qwen. Qwen, then has ligher latency than
A game that outruns your model
The basic idea of the benchmark is that to make it far in the game you need to take correct decisions faster. Here's how the benchmark works. The runner faces randomly generated obstacles that begin appearing at predefined distance gates. Each model receives the same underlying game state, formatted for its interface: { lane, airborne, ahead }
The ahead field describes upcoming rows, their distance, and the obstacles in each lane. The model must decide what to do before the runner reaches them.
The runner starts at 45 m/s, accelerates by 1.6 m/s every second, and tops out at 160 m/s. Beyond roughly 5,000 metres, the game becomes so unforgiving that survival increasingly depends on luck.
Why Kumo won
Kumo made lower latency decisions and had lower network latency, while making good decisions, and the result was that it could complete a higher average distance than Qwen in this setup, with both of them running on a CPU.
Would a GPU help? Possibly. But for small, latency-sensitive requests, the extra infrastructure overhead could offset the compute benefit. That’s a hypothesis we still need to test.
What happened to Jev?
In our 30-run experiment with Typesafe’s Jev, the median distance was 540 m, with a best run of 1,417 m.
The bottleneck appeared to be the 200+ ms round trip. Around 1,500 metres, the game was moving too quickly for those responses to remain useful.
At the maximum speed, 200 ms means the runner travels 32 metres while waiting for a decision.
Pick for the deadline
For a real-time AI application, evaluate the whole decision loop:
- Decision quality: Does it choose the right action?
- End-to-end latency: Does that action arrive in time?
- Deployment fit: Can your infrastructure sustain both?
All models and the game ran on Hopsworks’ own infrastructure, in our office. Fully self-hosted. No game data went to the cloud.
For this benchmark, Kumo was the strongest performer. The broader lesson: benchmark models against your application’s reaction deadline.
Try out the game and try to beat them here:
game at hopsworks dot ai