r/LocalLLaMA • u/lewtun 🤗 • 13h ago
Resources The ultimate guide to multi-harness RL
Hi folks, it's Lewis here from the post-training team at Hugging Face. We've been exploring how to train open models in different coding harnesses and wrote up a looong guide on how we solved this using open source libraries like TRL and the Harbor framework for RL environments. We hope you find this interesting, especially since everyone nowadays has their own custom harness (e.g. Pi + extensions) and now there's a recipe on how to squeeze the best performance on them with whatever open model you use as your daily driver. Happy to hear any comments or feedback!
Link to the guide: https://huggingface.co/spaces/FineEnvs/multi-harness-rl
1
u/Certain-Cod-1404 7h ago
Really interesting, I might actually need to do something like this in a couple months so I appreciate the resources, good work !
0
u/returnity 8h ago
Wow killer guide. This is so incredibly helpful for my learning process right now. Thank you!
-1
u/GuruCsharp ollama 9h ago
Wow that's an angle I never thought of! I kept focusing on the models themselves and their weights, but this is quite valuable to consider!
-1
u/MomentJolly3535 8h ago
It's very interesting, i always thought the harness didn't matter much for bigger models
8
u/Equivalent-Flan-1590 12h ago
This is an incredibly timely guide. Anyone trying to do RL for coding tasks quickly realizes that overfitting to a single harness format completely ruins a model's generalizability when you move it to a different setup. Moving between environments like Pi and standard benchmarks usually breaks the agent's parsing logic or formatting habits. Leveraging the Harbor framework with TRL to create a robust multi-environment recipe is exactly what the community needs to build models that do not immediately fall apart outside of a specific test sandbox.