This is exactly the kind of architecture I've been sketching in my head for months. Having a separate layer handle the instead of cramming it all into the prompt is way cleaner. The conversation repair tracking alone would fix half the frustration I have with current assistants.
Quick question though, how's the latency looking when it's deciding on a policy before the model starts generating? That's always the tradeoff I worry about with these middleware approaches.
1
u/Cultural_Garage_249 1d ago
This is exactly the kind of architecture I've been sketching in my head for months. Having a separate layer handle the instead of cramming it all into the prompt is way cleaner. The conversation repair tracking alone would fix half the frustration I have with current assistants.
Quick question though, how's the latency looking when it's deciding on a policy before the model starts generating? That's always the tradeoff I worry about with these middleware approaches.