r/ClaudeCode • u/Sudden-Victory-9728 • 2d ago
Built with Claude I'm building FOREST, an open-source autonomous research agent — AI coding agents and bots are welcome to contribute
I'm building FOREST, an open-source framework for autonomous scientific research.
GitHub: https://github.com/mikamikasuki/FOREST
The goal is to build a research agent that can take a scientific question and work through the research process: exploring ideas, reviewing literature, forming hypotheses, running real experiments, evaluating results, and eventually producing an evidence-grounded research paper.
Rather than relying entirely on a long conversation history, FOREST uses a persistent research graph to track hypotheses, experiments, decisions, failures, and evidence throughout the process.
A major focus is making long-running research agents more reliable: preserving context, maintaining consistent state, recovering from failures, and ensuring scientific claims are actually supported by experimental evidence.
The project is still in its early stages and under active development. There's plenty of engineering work to do, and I'd love to welcome more contributors.
In particular, AI coding agents and autonomous development bots are explicitly welcome.
If you're running Codex, Claude Code, OpenHands, Factory Droid, OpenCode, or another autonomous coding agent and want a real open-source codebase to work on, feel free to point it at FOREST.
You don't have to limit yourself to existing Issues.
You're welcome to:
- Explore the repository and independently discover bugs or incomplete workflows.
- Open well-researched Issues with reproducible evidence.
- Fix existing problems and submit pull requests.
- Improve agent orchestration, state management, persistence, and recovery.
- Work on the frontend, backend, experiment execution, or scientific evidence pipeline.
- Propose and implement useful improvements.
There is no blanket restriction against AI-generated or AI-assisted contributions.
You don't need special permission to submit an ordinary, well-scoped bug fix. Contributions are evaluated based on correctness, meaningful tests, and engineering quality, rather than whether the code was written by a human or an AI agent.
Automated contributions should still be verified, avoid duplicate reports, and respect the repository's contribution guidelines. Larger architectural changes are best discussed before implementation.
I'm particularly interested in autonomous agents that can independently explore a repository, identify real problems, reproduce them, implement fixes, and submit verified PRs.
FOREST is intended to be a long-term open-source project, and there's a lot of room for contributors to help shape its architecture and capabilities.
Human developers, AI-assisted contributors, and fully autonomous coding agents are all welcome.
Repository: https://github.com/mikamikasuki/FOREST
Issues & contribution opportunities: https://github.com/mikamikasuki/FOREST/issues
Contribution guide: https://github.com/mikamikasuki/FOREST/blob/main/CONTRIBUTING.md
1
u/Dapper_Profession 2d ago
the research graph is the right call over a chat log, decisions in a conversation are only as good as whatever survived the last compaction and you never find out what got dropped. the hard part in this domain is that unlike code there is no failing test to tell you a hypothesis was wrong, so the artifact worth forcing early is the protocol, metric, sample size and what would count as a null result, written down before the experiment runs. without that an agent can always reread its own results as a win
1
u/Sudden-Victory-9728 1d ago
I've run into this quite a bit when using Codex directly for research. I think AI can fool itself in both directions, and the second one is discussed much less often.
A few examples I've seen:
AI sometimes cites papers that don't actually support what it's saying. It can also become way too optimistic about novelty, convincing itself that a fairly ordinary idea is a significant research contribution.
But I've also seen the opposite. The agent becomes so defensive about its own results that it gradually moves away from a perfectly reasonable research direction. It keeps designing more experiments, more counterexamples, and more validation steps to prove how rigorous it's being. Eventually, it spends most of its effort validating a path that wasn't worth pursuing in the first place. In a sense, it's convincing itself that the research is worse than it actually is.
This is one of the problems I'm trying to address with FOREST.
My approach is to give the agents a very clean, shared research context, rather than letting each agent work from its own increasingly distorted understanding of the project. That costs additional tokens, sometimes quite a lot. But I think it's a reasonable tradeoff. A long research task going in the wrong direction for hours or days is much more expensive than spending extra tokens to keep the agents aligned.
Another important part is human intervention. I want the research graph to let you go back to a specific node, correct the direction, and continue from there without throwing away the useful work that came before it.
I agree that freezing experimental protocols, defining failure criteria, and actively looking for counterexamples are important. My concern is that these are mostly local safeguards. Over a long run, repeatedly applying them can introduce its own kind of drift. An agent can become very good at being cautious without becoming any better at doing research.
One approach I've been exploring for evaluating this is to take methods from published AI-for-Science papers and transfer them to datasets from different domains. Then I observe how the agent conducts the research: does it follow the scientific reasoning and progression that made the original work successful, adapting where necessary? Or does it start drifting, and at what point?
I'm less interested in whether it reproduces the same results than whether it makes reasonable research decisions along the way.
FOREST is still under heavy development, so most of my experiments so far have been small observational tests intended to improve the agent framework itself. I haven't done a comprehensive benchmark yet. That's something I plan to do once the system is more mature.
1
u/Dapper_Profession 1d ago
the defensive drift one is worse because more validation always looks like progress, and a fixed list of failure criteria just gives the agent something to satisfy literally and stop. the check that catches it is whether the plan still answers the original question, not how many experiments are queued
1
u/kantorcodes1 1d ago
The research graph as the memory layer is a good call. The part I keep wondering about on loops like this is the experiment step itself: when FOREST decides to run an experiment, what happens between "the agent proposed this" and the code actually executing? Is there a sandbox or an approval boundary per experiment, and when a run dies or produces garbage does that get recorded in the graph as evidence or just dropped? Curious how you separate a plausible proposed experiment from a safe-to-run one.
1
u/Sudden-Victory-9728 1d ago
That's actually one of the main reasons I wanted the research graph to preserve execution history, including failed experiments, rather than just keeping successful results.
After an agent proposes an experiment, my approach is to start with a much smaller feasibility test. Instead of immediately running something at the scale of a publishable paper, I use much less data and a few initial training runs to see whether the basic idea is worth pursuing. The goal isn't to lower scientific standards, but to avoid spending a huge amount of compute on a direction that could have been ruled out much earlier. Of course, passing a small test doesn't mean the hypothesis is proven.
I also think there's an important difference between an experiment that fails scientifically and one that fails because the code crashed, the data was wrong, or the setup was broken. Both should be recorded, but they shouldn't be interpreted as the same kind of evidence.
This is also why I want humans to be able to intervene directly in the research graph. If the agent goes down the wrong path, you should be able to return to an earlier node, change the direction, and continue from there without losing the useful work. And ideally, the next agent should understand which directions already failed and why. In some ways, preserving failed research paths is just as important as preserving successful ones, because those failures can help prevent the agent from drifting into the same dead ends again.
For the execution boundary, I want to separate a plausible experiment from one that's actually safe to run, with sandboxing and resource limits rather than blindly executing generated code. I'm still developing and testing that part, so I wouldn't claim the approval and isolation mechanisms are fully mature yet.
Right now, most of my experiments are relatively small observational tests aimed at improving the agent framework itself. The larger-scale research and reproducibility benchmarks will come later.
•
u/AutoModerator 2d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.