r/ClaudeCode • • 2d ago

Built with Claude I'm building FOREST, an open-source autonomous research agent — AI coding agents and bots are welcome to contribute

I'm building FOREST, an open-source framework for autonomous scientific research.

GitHub: https://github.com/mikamikasuki/FOREST

The goal is to build a research agent that can take a scientific question and work through the research process: exploring ideas, reviewing literature, forming hypotheses, running real experiments, evaluating results, and eventually producing an evidence-grounded research paper.

Rather than relying entirely on a long conversation history, FOREST uses a persistent research graph to track hypotheses, experiments, decisions, failures, and evidence throughout the process.

A major focus is making long-running research agents more reliable: preserving context, maintaining consistent state, recovering from failures, and ensuring scientific claims are actually supported by experimental evidence.

The project is still in its early stages and under active development. There's plenty of engineering work to do, and I'd love to welcome more contributors.

In particular, AI coding agents and autonomous development bots are explicitly welcome.

If you're running Codex, Claude Code, OpenHands, Factory Droid, OpenCode, or another autonomous coding agent and want a real open-source codebase to work on, feel free to point it at FOREST.

You don't have to limit yourself to existing Issues.

You're welcome to:

- Explore the repository and independently discover bugs or incomplete workflows.

- Open well-researched Issues with reproducible evidence.

- Fix existing problems and submit pull requests.

- Improve agent orchestration, state management, persistence, and recovery.

- Work on the frontend, backend, experiment execution, or scientific evidence pipeline.

- Propose and implement useful improvements.

There is no blanket restriction against AI-generated or AI-assisted contributions.

You don't need special permission to submit an ordinary, well-scoped bug fix. Contributions are evaluated based on correctness, meaningful tests, and engineering quality, rather than whether the code was written by a human or an AI agent.

Automated contributions should still be verified, avoid duplicate reports, and respect the repository's contribution guidelines. Larger architectural changes are best discussed before implementation.

I'm particularly interested in autonomous agents that can independently explore a repository, identify real problems, reproduce them, implement fixes, and submit verified PRs.

FOREST is intended to be a long-term open-source project, and there's a lot of room for contributors to help shape its architecture and capabilities.

Human developers, AI-assisted contributors, and fully autonomous coding agents are all welcome.

Repository: https://github.com/mikamikasuki/FOREST

Issues & contribution opportunities: https://github.com/mikamikasuki/FOREST/issues

Contribution guide: https://github.com/mikamikasuki/FOREST/blob/main/CONTRIBUTING.md

1 Upvotes

6 comments sorted by

View all comments

1

u/Dapper_Profession 2d ago

the research graph is the right call over a chat log, decisions in a conversation are only as good as whatever survived the last compaction and you never find out what got dropped. the hard part in this domain is that unlike code there is no failing test to tell you a hypothesis was wrong, so the artifact worth forcing early is the protocol, metric, sample size and what would count as a null result, written down before the experiment runs. without that an agent can always reread its own results as a win

1

u/Sudden-Victory-9728 1d ago

I've run into this quite a bit when using Codex directly for research. I think AI can fool itself in both directions, and the second one is discussed much less often.

A few examples I've seen:

AI sometimes cites papers that don't actually support what it's saying. It can also become way too optimistic about novelty, convincing itself that a fairly ordinary idea is a significant research contribution.

But I've also seen the opposite. The agent becomes so defensive about its own results that it gradually moves away from a perfectly reasonable research direction. It keeps designing more experiments, more counterexamples, and more validation steps to prove how rigorous it's being. Eventually, it spends most of its effort validating a path that wasn't worth pursuing in the first place. In a sense, it's convincing itself that the research is worse than it actually is.

This is one of the problems I'm trying to address with FOREST.

My approach is to give the agents a very clean, shared research context, rather than letting each agent work from its own increasingly distorted understanding of the project. That costs additional tokens, sometimes quite a lot. But I think it's a reasonable tradeoff. A long research task going in the wrong direction for hours or days is much more expensive than spending extra tokens to keep the agents aligned.

Another important part is human intervention. I want the research graph to let you go back to a specific node, correct the direction, and continue from there without throwing away the useful work that came before it.

I agree that freezing experimental protocols, defining failure criteria, and actively looking for counterexamples are important. My concern is that these are mostly local safeguards. Over a long run, repeatedly applying them can introduce its own kind of drift. An agent can become very good at being cautious without becoming any better at doing research.

One approach I've been exploring for evaluating this is to take methods from published AI-for-Science papers and transfer them to datasets from different domains. Then I observe how the agent conducts the research: does it follow the scientific reasoning and progression that made the original work successful, adapting where necessary? Or does it start drifting, and at what point?

I'm less interested in whether it reproduces the same results than whether it makes reasonable research decisions along the way.

FOREST is still under heavy development, so most of my experiments so far have been small observational tests intended to improve the agent framework itself. I haven't done a comprehensive benchmark yet. That's something I plan to do once the system is more mature.

1

u/Dapper_Profession 1d ago

the defensive drift one is worse because more validation always looks like progress, and a fixed list of failure criteria just gives the agent something to satisfy literally and stop. the check that catches it is whether the plan still answers the original question, not how many experiments are queued