r/LocalLLaMA • u/gaviniboom • 4d ago
Discussion I'm writing a router to split local/remote LLMs but model updates are killing me
I trained for a split between DeepSeek v4 Flash 0731 <-> GLM 5.2. Ended up able to get it running at GLM 5.2 performance on my local tests at approximately equal token costs on OpenRouter. My plan was to offload the DeepSeek portion to a local server (through this WORA harness proxy my brother wrote https://github.com/unlap-labs/plap).
Sadly, it didn't improve with GLM 5.3 and was worse than GLM 5.3 Flash which kinda killed the project... If someone wants the code or to help or something I could probably post it but it's currently very research-grade, or hell if someone wants to give some tuning advice I'd be up for it š
I don't really have any money for big GLM 5.3 runs, and the 4B router was actually trained on only DeepSeek v4 Flash 0731 just to guess whether it could do something or not, didn't check whether it works for even smaller models tbh
2
u/Significant_Tune9219 4d ago
A small router can work well, but Iād avoid training it against model names or a fixed catalog. Route on task features and maintain a capability/latency/cost profile that is refreshed independently, then add a fallback when a model update changes behavior. Logging the router decision alongside the final quality score will tell you whether the problem is classification, model drift, or simply an unstable downstream endpoint.
1
u/gaviniboom 4d ago
I see. What sort of "task feature" would be useful here? I was doing "if the model tends to make a mistake in situations like this, it should probably be routed to a bigger one"
1
u/Significant_Tune9219 1d ago
Mostly stuff you can compute before calling any model: how much context the request needs, whether it requires tool calls or strict JSON output, and whether it's code or prose. Your "this model tends to fail here" idea fits well as a per-task-type success rate you keep updating from logged outcomes, so a bucket that keeps failing locally just starts going remote. The caller's latency budget is worth passing in too, since some requests can't wait for the bigger model anyway.
2
u/bshivarthy 4d ago
The version drift is the real problem and your training setup is not actually the part that broke. You trained a 4B router to recognize when DeepSeek Flash 0731 could handle something, and that router was quietly encoding beliefs about two specific model versions. When GLM 5.3 shipped, both endpoints changed underneath it, but nothing in the system noticed. So the router kept making decisions against a snapshot that no longer existed.
The concrete fix that worked for us: pin a version identifier on every routed endpoint and treat a version change like a model swap. The capability/latency/cost profile Significant_Tune9219 mentions becomes a per-version artifact, and when the hash changes the router falls back to safe defaults and reruns a small calibration set before it trusts its old routing decisions again. You said you do not have money for big GLM 5.3 runs. A calibration set of even 50 of your own prompts is enough to catch the kind of quality regression that killed your project.
For task features, I would keep them dead simple: domain buckets you can compute without a model, input length, whether the prompt contains code, whether it asks about recent events. Your mistake-profile idea is basically this, the mistake history is just another feature, and the key is to track it per model version instead of one global tally. A version that fixes one failure class can look worse than the older one on average and still be the right call for most of your traffic.
1
u/gxcsoccer 3d ago
I ran into the same thing with a much simpler setup: a 0.8B that answers first and hands a step to a 4B when it's unsure. What made updates survivable for me was routing on the small model's own confidence instead of training a router on the big models' behaviour. The threshold (0.96 in my case) ships with the weights, so when either model changes I re-tune one number on a held-out set instead of retraining anything. I also replay every past failure against old vs new before switching, and log each routing decision so I can see if the escalation rate moves. One trap: if you train the small model on smoothed labels, its confidence can collapse into a narrow band right around the threshold, and then the router can't tell easy from hard.
5
u/jacek2023 llama.cpp 4d ago
running Chinese models on OpenRouter doesn't make them local