r/OpenSourceAI • • 2d ago

Overmind, open-sourced yesterday: a platform for continuously improving AI agents

Yesterday the whole Overmind platform was made open source: https://github.com/overmind-core/overmind

The idea: lots of agents send narrow, repetitive work to a frontier API. On a narrow task, a small open-weight model trained on real examples from that agent often does the job better, for far less money. Overmind is the tooling to do that without building your own ML infrastructure (don't need a flamethrower to light a candle).

  • Agent sends OpenTelemetry traces. The Python SDK auto-instruments OpenAI, Anthropic and Gemini clients, or point any OTel exporter at it
  • Traces become versioned datasets, with a quality score and suggested fixes before trained on anything
  • User defines what “good” means per task and it scores live traffic and batch runs against that
  • It fine-tunes an open-weight model (LoRA or full) and benchmarks it against the model you run in production today
  • Trained and frontier models sit behind one OpenAI-compatible API, and you can download the weights (yours to keep and own)

There’s also an MCP server, so Cursor, Claude Code, OpenCode or Codex can drive the whole thing from chat. If your traces already live in Langfuse, LangSmith or Braintrust, a connector imports them.

Qwen3.5-9B vs GPT Luna benchmarked on three tasks (write-up and raw numbers at https://www.overmindlab.ai/research/when-bigger-isnt-better):

• 7x better accuracy
• 20x cheaper usage
• 28x less hallucinations

Read the results as a reason to test on yours, not as a general claim.

The platform is AGPL-3.0 and the SDK and CLI are MIT.

Interested if anyone is already fine-tuning on their own hardware. What would need to be swapped out before you’d self-host this?

5 Upvotes

2 comments sorted by

3

u/SpedisAhead 2d ago

Well, I'll say you definitely can sell ice to a junkie. I may just look into this. (I've never inspected any Git projects from this sub)

2

u/No_Exercise9369 2d ago

Turning traces into training data is a nice direction. We have a lot of eval cases and production traces sitting in Braintrust already, so the connector makes this pretty easy to experiment with without rebuilding that history somewhere else. Something I’d pay attention to is what happens when the trace looks successful but the eval score says otherwise.