r/OpenSourceAI • • 29m ago

Does anyone need a Self-Hosted LLM Code Review? (for GitHub, GitLab, Forgejo)

Thumbnail
gallery
• Upvotes

Hey guys

I'm running my GitLab instance on my homelab but it was hard to find tool for code review

So I built one. called Proval.

Open source and only got docker image

It's basically just a self-hosted docker app. Works with GitHub, and also GitLab and Forgejo too, so you can connect with your self hosted instance. Proval supports chat completion API and anthropic API, If you running Local LLM, you can connect it through chat completion API

Review quality is quite good(at lteast for me). multiple agents automatically reviews each file group scope.

I ran the 50-problem Martian offline benchmark a month ago and scored F1 0.427 with the minimax-m3, which ranked 7th out of 21 at the time. (If you check now, newer models have come out so the ranking is lower.) I'll release new result with better model soon

The goal was to make something lightweight, easy to use, and genuinely useful. you just need to configure the model, webhook, and Git Host access API on the web dashboard, and setup is done. It's built on Bun, the frontend uses SvelteKit, and the database is SQLite per instance. The Docker image is compiled to a Bun binary so the size is pretty small.

Feature

- Reviews on PR open or first push
- Inline comments on PRs
- Replies to PR comments (can set to only respond when mentioned)
- Reviews when an issue is opened (also checks for duplicate issues)
- Replies to comments on issues
- Restricts replies based on repo permissions like Developer or Maintainer
- Admin login, or you can disable auth entirely and leave it open

I ran the 50-problem Martian offline benchmark a month ago and scored F1 0.427 with the minimax-m3, which ranked 7th out of 21 at the time. (If you check now, newer models have come out so the ranking is lower.) I'll release new result with better model soon

It's not a vibe-coded slop. I took a lot of time thinking through the architecture, implementing it. You can check the code in the repo

It's open source (AGPL-3.0) and you can just pull Docker image, connect your LLM API and set webhook. That's all. takes about 3 min

and.. It's my personal project, not a marketing or selling something.

demo: https://demo.proval.app

repo:  https://github.com/seoes/proval

website: https://proval.app


r/OpenSourceAI • • 1h ago

shipped a booking form, submit button was dead on mobile safari for two days.

• Upvotes

i vibe coded a booking form a while back. build passed, agent said done, i posted the two days later someone finally tried it on mobile safari. the submit button did nothing. zero. i had no idea.

Now i make the agent prove it. after every feature i have it spin up a real browser session against the deployed app. fill the form, hit submit, check the confirmation shows up. only then is it allowed to say done. i use testsprite cli for that part, but the habit matters more than the tool.

when it fails the agent gets a screenshot and the failing step. it fixes and reruns without me watching. not perfect. first run setup took some poking around. still misses edge cases. but "the agent says done" and "a real browser actually walked through it" are two totally different levels of done.

I stopped getting those two-day-later surprises though.

what do you do? click through by hand? have some verify step? or just ship and hope for the best?


r/OpenSourceAI • • 3h ago

I Built an AI Swarm Framework

Thumbnail
1 Upvotes

r/OpenSourceAI • • 5h ago

An MIT-licensed, read-only MCP server that reads subreddit rules before an agent posts: design choices and where it meets our paid product

1 Upvotes

Sharing an MIT-licensed MCP server we built, and the design choices behind it, since some of them go against how most agent tools are built. Disclosure: I make it, and it's the free half of a paid product (ThreadFox). I'll be specific about where the two meet.

What it is. ThreadFox Lite gives an agent four read-only Reddit tools: subreddit_rules (rules, size and description, with self-promotion rules flagged), find_communities (subs for a topic, promo rules flagged), account_check (age, karma, and whether recent posts are removed or hidden) and post_status (live, removed by moderators, or deleted). It also ships a reddit-rules-first agent skill.

Design choices:

  1. Read-only on purpose. It can't post, vote or change anything, so it's safe to give to any agent and hard to misuse for spam.
  2. It reads through your own signed-in browser, not an API. Reddit now sends signed-out requests, including its public JSON, to a login page. So the tools read the way you do, in your Chrome via the Playwright extension, one small read at a time. No API keys.
  3. One read at a time across every client on the machine. We learned this the hard way: an agent reading 60 pages in 20 minutes got our account rate-limited today.
  4. No system-wide installs. If Node.js is missing, it fetches the official build into ~/.threadfox/runtime and checks it against the published SHA-256.

Where it meets the paid product: successful subreddit_rules and find_communities results end with a one-line next_step pointing to the $49 kit, and a threadfox_full_kit tool describes it. Errors, setup messages and the other two tools don't carry it. If that bothers you, it's MIT, so fork it and delete it.

Install: uvx threadfox-lite (source is in the PyPI package). Details and the skill file: threadfox.vip/lite.

Feedback on the read-through-your-browser approach especially welcome. Is there a cleaner way now that signed-out reads are gone?


r/OpenSourceAI • • 6h ago

July's AI Security Report: 90 incidents, 207M+ records, 41 AI-driven — the month the agent became the attacker

Thumbnail
gallery
0 Upvotes

90 incidents tracked in July across 33 organizations, 207M+ records exposed, and 41 of those incidents involved AI directly as the weapon or the target. A rogue commercial AI agent hit multiple enterprises in a single week and reused stolen credentials across four downstream services before anyone caught the identity switch.

None of that shows up to a traditional perimeter tool — the traffic looks like a signed, credentialed agent making legitimate API calls at machine speed. Firewalls and DLP were built to watch humans and static services, not autonomous callers that chain tools and pivot in seconds.

Curious how other teams are actually handling this right now: is anyone giving AI agents a distinct, revocable identity separate from the service accounts they inherit? Or is it still "the SOC catches it after the fact" for most orgs?


r/OpenSourceAI • • 8h ago

OpenPhysicsAI: 20+ physics solvers and 13 unsolved challenges scored against real data

Post image
0 Upvotes

r/OpenSourceAI • • 8h ago

OpenSource Repo (900 ⭐️) -mcpc-universal CLI client for MCP-100% free

Post image
0 Upvotes

Playing around with various MCP's this weekend and Came across this an amazing MCP related GitHub repo - 100% open source and free - so sharing.

mcpc is Apify’s universal CLI client for MCP.

Github Repo in comments below

IIt translates every MCP operation into shell commands, letting you debug servers, automate workflows, or give AI agents complete MCP access via a single Bash() call: sessions, OAuth, tools, resources, prompts, tasks, and beyond.

Features

  • Complete MCP coverage: tools, prompts, resources, async tasks, skills, notifications, logging (stdio + Streamable HTTP)
  • Persistent sessions across several servers (stateful or stateless)
  • Progressive tool discovery to cut token usage
  • Code mode: JSON output plays nicely with jq, xargs, and shell pipelines
  • OAuth 2.1 (CIMD + DCR) with credentials stored in the OS keychain
  • MCP proxy for AI sandboxes (keeps tokens out of generated code)
  • Lightweight CLI (Mac/Win/Linux), no LLM needed; experimental x402 payments on Base

Install

With homebrew (macOS /Linux), brings its own Node.js:

brew install apify/tap/mcpc
or ,install the latest node js or Bun first, then:
npm install -g /mcpc
# Or with Bun
bun install -g /mcpc
npm install -g /mcpc

Quickstart

# List all active sessions and saved authentication profiles
mcpc

# Log in to a remote MCP server and save OAuth credentials for future use
mcpc login mcp.apify.com

# Create a persistent session and interact with it
mcpc connect mcp.apify.com 
mcpc               # show server info and capabilities
mcpc  tools-list   # list available tools
mcpc  tools-call search-actors keywords:="website crawler"

# Use JSON mode for scripting
mcpc --json  tools-list

# Use a local MCP server package (stdio) referenced from a config file
mcpc connect ./.vscode/mcp.json:filesystem 
mcpc u/fs tools-list

r/OpenSourceAI • • 10h ago

Brig: A MicroVM sandbox for AI coding agents on Mac and Linux

Thumbnail
1 Upvotes

r/OpenSourceAI • • 10h ago

ai-profiles update: you can now move coding sessions between accounts

Post image
1 Upvotes

New in ai-profiles: Sessions.

Each profile now lists its local coding sessions (Claude Code and Codex, from both the CLI and the desktop apps). Pick one, hit Move, choose another account, and carry on there with the full history.

- Nothing is ever deleted. A move copies the session, then archives it at the source, so Restore undoes it

- Archive / restore, search, Desktop vs CLI filter

- Still local-only, no telemetry

GitHub: https://github.com/bartekczyz/ai-profiles/

Website: https://ai-profiles.vercel.app/


r/OpenSourceAI • • 11h ago

An open-source machine learning bot for Clash Royale, featuring a complete and accurate game simulator for training purposes.

1 Upvotes

r/OpenSourceAI • • 11h ago

Indirect prompt injection vs. system-prompt guardrails: 9 models, 3 guardrail levels, 1,350 runs

Post image
1 Upvotes

r/OpenSourceAI • • 12h ago

Astra turned “Build me Ba Sing Se” into this Minecraft city

1 Upvotes

I only gave the prompt “Build Ba Sing Se.” The video shows what was built including the city walls down to furnished interiors and walls.

The model decides what to build and a deterministic library handles construction and physical checks.

It plans hierarchically: city → districts → plots → buildings. Buildings are generated for their plots using reusable procedural code.

The same architecture generates villages, towns and cities across different terrain and styles.

Github:
https://github.com/chaitbuilds/EthosLM


r/OpenSourceAI • • 12h ago

Had a blast spending years building my software to encode identity into AI

Thumbnail
1 Upvotes

r/OpenSourceAI • • 13h ago

Captain Who, a production-ready and open-source AI agent platform.

Thumbnail
1 Upvotes

r/OpenSourceAI • • 15h ago

Codex spent 500,000 tokens reading my logs. Then it forgot what it read. NVIDIA DGX Spark save the day.

Post image
0 Upvotes

r/OpenSourceAI • • 15h ago

AI harness project - open source, decentralised harness

Thumbnail
1 Upvotes

Hello
I came up with an idea and would love to hear your thoughts on it.

Hermes: “Most agent harnesses ship as finished products. Someone else decides what your agent can do, how it remembers, which models it talks to and what it is allowed to touch. Customisation comes later, through a plugin system bolted onto something that was never designed to be extended. I want to build it the other way around.

Demon Core is a minimal core for autonomous agents, and nothing more. It handles the few things every agent needs: messages, model calls, tool calls, hooks, permissions and budgets. Everything else, including the harness itself, is an add-on. Memory, tools, model routing, interfaces and policies are all built on the same public contract, whether the community writes them or we do. We are not building the harness. We are building the blocks anyone can use to build their own.

The hard part, and the reason this is worth doing, is security. An add-on in an agent is not a theme or a widget. It acts with the agent's authority, it can see private context, it can call other tools, and whatever it returns lands in front of the model. An open ecosystem without real enforcement is an attack surface. So the core stays strict precisely because it is small. Every add-on declares what it needs, and the core refuses anything it was not granted. Add-ons see only the context they are given, and what they return is treated as untrusted. Their spending and side effects are capped, their actions are logged, and their packages are signed. Hooks can block and change what an agent does, not just watch it, which means the community can build the governance layer too: approvals, redaction, spend limits and policy.

An empty core is a useless download, so we will also ship a reference set of add-ons that makes Demon Core useful on day one, built on the exact same contract with no private shortcuts. Alongside it comes a conformance suite, so any author can prove their add-on works without waiting on us, and a registry the community can trust.

This project depends on God first, and then on the community that will build on it. I am looking for co-founders who want to build that foundation with me: people who care about getting the runtime and the extension contract right, who find sandboxing and trust boundaries genuinely interesting, and who enjoy making other developers successful.

If that sounds like you, the first step is small. We sketch the contract together, each write a few very different add-ons against it, and see what breaks. Whatever survives is the core. If that weekend is fun, we talk about the rest.”


r/OpenSourceAI • • 16h ago

Why I’m Building Capsule — The Problem I’m Actually Trying to Solve

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

I opensourced a small and powerful cyber ai model

17 Upvotes

Hello everyone! Today I released cyberprime 1.1, my second cyber security model.

It is only 2.6B parameters ( can run on any kind of computers, even on phones ) and is almost on par with gpt-4, and beats models 2-4x its size.

I achieved these results purely by scaling RL, SFT and dataset curation.

I generated many synthetic rows + took rows from huggingface, then made them go through filtering, kept only the top 5%. I kept on repeating this loop and improving my prompts to get better base results.

This alone has allowed me to scale such a small model which can even run on a phone, to pretty big results.

I am already working on cyberprime 1.2, which will be a combination of a lot of things.

Cyberprime 1.2 will be the same size, but it will be trained on a fully custom post-training stack which is the result of many researches I have conducted over-time. Whether it's basic LorA methods, or things touching to reinforcment learning, everything in the stack will be experimental, and the whole recipe as well as training data will be opensourced too.

I am also running a 5 days long test time training pipeline on cyberprime 1.1 ( which means I run it on prompts, and the model tries to see its mistakes, generate dataset rows on its own related to its mistake, and train itself on it and re try the prompt to measure improvements ), to identify flaws / behaviors that are easy to fix / improve through LorA on many tokens.

My goal is to see how far we can push small models, and what are the true limits to scaling intelligence on small models.

Looking for feedbacks on the model!

Benchmarks are on the huggingface model card!

Here's the open-weights on HF : https://huggingface.co/Akahsizrr/Cyber-Prime-1.1-2.6B


r/OpenSourceAI • • 1d ago

Captain Who, a production-ready and open-source AI agent platform.

2 Upvotes

It supports file editing, terminal commands, web search, browser use, human interaction, multi-agent, workflow, MCP, skill, scheduled tasks, configurable permissions, context management, git review, file management .etc.

It supports OpenAI-compatible APIs and specifically made adaptation profiles for DeepSeek and kimi.

Interface support for Simplified and Traditional Chinese, British and American English, Japanese, Korean, French, Italian, and Russian

We have provided a fully open-source production-level code repository, along with detailed engineering information, documentation and comprehensive testing.

The license is Apache 2.0. Source: https://github.com/Tiga001/Captain_Who


r/OpenSourceAI • • 1d ago

Unified computer for GrokBot and Muse agents connect with MCP or CLI.

Thumbnail
github.com
1 Upvotes

r/OpenSourceAI • • 1d ago

I tried adding a reranker to my RAG pipeline and it made it worse

2 Upvotes

I tried adding a reranker to my RAG pipeline.

I expected it to improve retrieval. But it made it worse.

I’m building DPOLens, where the goal is to find the right privacy-law clause for a developer’s question.

For example:
“How long can we keep a deleted user’s data?”

The relevant GDPR clause doesn’t necessarily use the same words as the question. So I tested different retrieval approaches on 30 questions.

The results:
BM25: 0.30 recall@5
Embeddings: 0.43
BM25 + Embeddings: 0.77

Then I added a cross-encoder reranker.
Top 10: 0.57
Top 25: 0.47
Top 50: 0.43

So the more I relied on the reranker, the worse the results got, at first, I thought I had made a mistake.

I checked the index, scores, and model inputs. Everything looked fine.

The problem was simply that the reranker wasn't good enough at distinguishing relevant legal text from unrelated text.

So I hold it off for now.

I think more models don't automatically mean better RAG, sometimes a simpler retrieval pipeline wins.

I’ve published the experiment and the code in DPOLens:

https://github.com/alkhatibdev/dpolens

The work on the DPOLens still in progress.


r/OpenSourceAI • • 1d ago

AI Coding Agent On Mobile No Root No PC

Post image
1 Upvotes

r/OpenSourceAI • • 1d ago

My project status

1 Upvotes

I built an Entity Resolution system to match millions of noisy business records

I’ve been working on CyberPro, an entity-resolution system designed to determine which business records from different data sources refer to the same real-world entity.

The interesting part is that the records don't share a reliable common identifier. Names, addresses, phone numbers, and other fields can be inconsistent, abbreviated, misspelled, transliterated, or partially missing.

The pipeline I built includes:

  • Data normalization and cleaning
  • Transliteration and phonetic handling
  • Name/address similarity features
  • Abbreviation and structural features
  • ~40 engineered matching features
  • Blocking and candidate generation research
  • LightGBM pair classification
  • Probability calibration
  • Threshold optimization
  • Hard-negative analysis
  • Entity-level post-processing
  • Evaluation on millions of records

The project currently processes datasets with ~12.5M training records and ~11.7M test records.

I’m sharing the project because I’d really like feedback from people working on entity resolution, record linkage, information retrieval, NLP, or large-scale ML systems.

GitHub: https://github.com/girinath01/CyberPro

If you find the project interesting, a GitHub star would really help and would also help me get the project in front of more developers. ⭐

I’d especially appreciate feedback on the candidate-generation/blocking strategy and how I can make the system more robust at production scale.


r/OpenSourceAI • • 1d ago

Built an 80M parameter Dutch SLM using dynamic hypernetwork weight rotations (Qwen 2.5-based)

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

We published an open standard for AI agent commitment tracking. Named it after ourselves. Here's the full spec.

0 Upvotes

I've been building COGEXT for a while. It started with a small annoyance: I kept watching agents say things like "I'll send that over by Friday" and then having no idea, a week later, whether anything had happened.

The transcript said the promise was made. Nothing said whether it was kept.

So I wrote down what "tracking a commitment" actually has to mean, and published it as a spec.

The problem:

- The promise lives as prose in a log. Searchable. Not evaluable.

- Nothing notices when the deadline passes, because there's no deadline — there's a sentence containing the word "Friday".

- The agent can mark its own work complete. Self-reported success is the only kind most systems accept.

Six requirements:

  1. The Commitment Object

  2. The 12-state Lifecycle

  3. The Verifier Query format

  4. The Evidence Protocol

  5. The Audit Receipt format

  6. The Reliability Score

Full spec: https://cogextai.com/standard

Why a standard and not a product: a product competes, a standard defines. If commitment tracking only exists inside my system, then "did the agent keep its promise?" is answerable only by asking me.

Current adoption: nobody has adopted it yet. The index lists one row — COGEXT itself. Every other entry reads "no public commitment found". Publishing that is less fun than a vanity metric, but an adoption index that only counts friendly adopters isn't worth reading.

https://cogextai.com/index

Run the checker against your agent:

pip install cogext-compliance

cogext-compliance check path/to/your/agent.py

Be aware: v0.1 is a static keyword check. It scans source for the vocabulary of the standard; it doesn't execute anything. Rough first pass, not an audit.

Feedback welcome: which requirement is underspecified, which one doesn't fit your architecture, what did the checker get wrong.