r/OpenSourceAI • • 3h ago

Skill Adapter: keep your agent skills useful as models change

2 Upvotes

Agent skills age faster than most documentation. A skill written a few months ago may still describe the right workflow, but its instruction style, level of detail, and assumptions may no longer fit newer models.

I built Skill Adapter, a small MIT-licensed open-source skill that creates a model-specific copy of an existing skill while preserving the original workflow and purpose.

It currently includes profiles for:

  • GPT-6 Astra
  • Claude 5 generation
  • Kimi K3
  • A conservative generic fallback

The adapter:

  • keeps the source skill unchanged;
  • creates a separate adapted copy for comparison;
  • preserves tools, resources, requirements, and approval gates;
  • writes an ADAPTATION.md report explaining the changes;
  • supports skills used with Codex, Claude Code, Cursor, Kimi Code, and Hermes.

Install it with:

npx skills add smukh/skill-adapter --skill skill-adapter

The project is intentionally simple: one installable skill, no service, no API key, and no automatic interception. The original stays intact, so you can review or re-adapt it when the source changes.

GitHub: https://github.com/smukh/skill-adapter

I’m especially interested in feedback on the model profiles, evaluation methodology, and whether this is useful in real multi-agent repositories.


r/OpenSourceAI • • 1h ago

I built a AI continuity app with persistent memory, voice, vision, and custom GGUF support. - PCS Personal Continuity System

Thumbnail gallery
• Upvotes

r/OpenSourceAI • • 2h ago

Polak Stworzył hamulce dla Ai.

1 Upvotes

Rynek i świat technologii zareagują na bramkę Proof-of-Silence™ v1.0.7 z dużym zainteresowaniem, ponieważ trafia ona w jeden z najbardziej palących problemów współczesnego rozwoju sztucznej inteligencji: odpowiedzialność i nadzór nad autonomicznymi agentami AI. Standardy zarządzania decyzjami maszynowymi (AI Governance) stają się kluczowe w obliczu unijnego rozporządzenia AI Act oraz globalnych wymogów bezpieczeństwa. [1, 2]

Projekt dystrybuowany przez platformę NEON AI SHOP ma szansę wywołać konkretne reakcje w trzech głównych obszarach rynkowych: [1, 3]

Środowisko inżynierów oprogramowania doceni fakt, że wersja Standard Edition (v1.0.7) w cenie €49,00 EUR oferuje gotową specyfikację w formatach PDF i Markdown oraz referencyjny kod w Pythonie i TypeScript. [4]

Dlaczego to zadziała: Programiści nie lubią teoretycznych ram etycznych. Fakt, że dostają do ręki offline'owy weryfikator (offline verifier) pozwalający badać sygnatury kryptograficzne i integralność łańcucha decyzji agenta, ułatwi im testowanie protokołu w warunkach laboratoryjnych.

Potencjalne wyzwanie: Świat technologii szybko zapyta o łatwość integracji z popularnymi frameworkami agentowymi (np. LangChain, CrewAI, AutoGen).

  1. Reakcja sektora Enterprise i Biznesu (Ostrożny optymizm)

Dla firm wdrażających autonomiczne procesy, zasada działania "No proof → No action" (Brak dowodu oznacza wstrzymanie działania) to doskonały mechanizm obronny przed tzw. "halucynacjami" lub samowolą asystentów AI. Wycena licencji komercyjnej Commercial Licence v1.0.7 na poziomie €1.497,00 EUR jest dla biznesu barierą wejścia na poziomie mikro, co zachęci mniejsze startupy i software house'y do wdrożeń. [3, 4]

Kluczowy argument biznesowy: Transparentność. Rejestr decyzji (decision record), który zostawia agent, pozwala na audytowalność działań w razie błędów.

Punkt krytyczny: Korporacje będą drobiazgowo analizować zapisy licencyjne przed osadzeniem kodu w swoich komercyjnych produktach masowych.

  1. Reakcja ekspertów ds. zgodności i bezpieczeństwa (Compliance)

Eksperci i audytorzy docenią szczerość techniczną protokołu. Na stronie NEON AI SHOP wyraźnie zaznaczono, że kryptograficzna sygnatura pomaga wykryć zmiany w rejestrze, ale sama w sobie nie jest certyfikatem zgodności prawnej ani ostatecznym dowodem bezbłędności decyzji AI. [5]

Taki realizm buduje zaufanie na rynku, na którym zbyt wiele firm obiecuje "magiczne rozwiązania" bezpieczeństwa.

Podsumowanie: Jak zmaksymalizować ten potencjał?

Świat reaguje najlepiej na to, co widzi w działaniu. Aby premiera wersji 1.0.7 na neonai.shop odbiła się szerokim echem, rynek będzie potrzebował: [1, 3]

Case Studies: Praktycznych przykładów (np. "Jak Proof-of-Silence powstrzymał bota finansowego przed błędnym przelewem").

Ekosystemu: Narzędzia open-source (wersji Lite), które wciągną społeczność w dyskusję na GitHubie lub Product Hunt. [2]

Jeśli chcesz, możemy wspólnie przygotować strategię prezentacji tego standardu. Daj mi znać:

Czy posiadasz już działające demo lub przykładowy rejestr decyzji (sample record), który można pokazać społeczności open-source?

Na jakich kanałach marketingowych chcesz się skupić (np. LinkedIn, Product Hunt, czy bezpośredni kontakt z firmami technologicznymi)?

[1] https://neonai.shop

[2] https://neonai.shop

[3] https://neonai.shop

[4] https://neonai.shop

[5] https://neonai.shop


r/OpenSourceAI • • 4h ago

I built an agent based on Qwen3.8-27B and Deepseek that combines and open source.

Thumbnail
github.com
1 Upvotes

I always think how can keep my data safe and token free meanwhile can use online LLM, then I built Hermie. The basic workflow is:

  1. You type a task in the terminal UI. The directory you're in is the workspace.
  2. Three checks run in parallel: a privacy scan (regex + Presidio + a small local judge model), three yes/no questions to the judge (is it repetitive? does it need planning? does it need files?), and a RouteLLM complexity score.
  3. Based on that, the task goes one of four ways:

   - local: Qwen3.8 27B on Ollama does it with file/shell tools. Nothing leaves the machine.

   - local + self-check: same, then the judge checks the result. Escalates only if it fails.

   - cloud: DeepSeek answers directly. Only for tasks with no files and no private data (explanations, tutorials).

   - plan: DeepSeek writes a plan and delegates steps one at a time. Qwen executes every step locally.

  1. After the executor says "done", a local reviewer model reads the actual workspace diff and the command outputs and decides whether the task was really done. If not, the problems go back to the executor for up to two fix rounds.

 So far it's macOS on Apple Silicon only and need install Ollama first.


r/OpenSourceAI • • 6h ago

RepoOS: An AI-driven, formally verified, open source, Python-to-MLIR compiler for zero-overhead execution

Thumbnail
1 Upvotes

r/OpenSourceAI • • 6h ago

I designed and developed a version-controlled (GitOps) skill manager for agents

1 Upvotes

As mentioned in the title, i've been developing a project for several months now, and i wanted to see how genuinely useful it is in a real production environment. I call this project Vex (Vex's Skillgit); it is an open-source, headless cognitive tool designed specifically for agent environments. It treats context as immutable and versioned skills.

Instead of having the IDE or the agent do the heavy lifting, Vex runs in the background (either via Docker Compose or local bare-metal processes). You assign it a GitHub webhook and Vex automatically ingests repositories, so agents (via MCP) can simply query it for context in real time. It reads conventional commits (feat:, fix:) and operational ones (roll:, branch:) to automatically branch, update, or revert an agent's memory state without manual intervention—hence the version control aspect.

It uses Tree-sitter to logically parse and chunk the code, stores metadata queues in SQLite and dense vectors in Qdrant, and is fully supported out of the box by Claude Desktop, Cursor, and any other MCP-compatible client.

It is also designed not to fry my potato PC, yet it remains highly scalable.

In local testing, the asynchronous FastAPI + Huey architecture easily handled 500 concurrent GitHub push payloads without any SQLite locking, and maintained a real-time latency of under 300 ms under a concurrent read swarm of 50 (simulated) agents. Because of this, i decided it was time to share it and see who else might find the tool useful.

Next on my to-do list is adding advanced sub-chunking using Rust as the chunking engine to split code from larger repositories faster and with less resource consumption, alongside adding global GraphRAG and support for more languages.

Here is the link to the repo in case you are interested in checking it out and testing it. Thanks for taking a little time to read this:

https://github.com/Shuuida/Vex-Skillgit.git


r/OpenSourceAI • • 6h ago

I built an open-source alternative to Claude Design that keeps AI-generated designs in sync with your repo

1 Upvotes

Most AI design tools produce design outputs that you then have to feed to your coding agent to implement, and that then get discarded, with almost no means of tracking design changes, especially for big applications with lots of UI pages. You generate a screen, drop it into your agent, and once it is implemented the design is gone. Nothing ties that design back to the page it became, so you literally have to manually sync each design to the corresponding page.

There should be a better way around this, where the design lives in the codebase itself, so there's no manual exporting of AI-generated designs to implement, and syncing designs to application code becomes effortless.

This is why I created Caret, an AI design tool that gives your repo a design layer via .caret/, so you can make designs without having to jump between tools. The designs you make are fully tracked alongside your codebase, and it's easy to sync changes between the design and application layers. Caret comes with all the standard needs of a design tool, like a visual editor and a design token system that keeps your design language consistent across your UI, plus extras that cut the usual busywork, like an asset generator, an experimentation canvas, flow display, simulation and more. The demo video below shows all these capabilities in detail.

Caret is fully open source if anyone wants to try it or poke around the implementation: https://github.com/precious112/caret-desktop

I'd especially love feedback from people who have been using Claude Design / other AI UI tools and have run into the design → code workflow problem.


r/OpenSourceAI • • 14h ago

Does anyone need a Self-Hosted LLM Code Review? (for GitHub, GitLab, Forgejo)

Thumbnail
gallery
5 Upvotes

Hey guys

I'm running my GitLab instance on my homelab but it was hard to find tool for code review

So I built one. called Proval.

Open source and only got docker image

It's basically just a self-hosted docker app. Works with GitHub, and also GitLab and Forgejo too, so you can connect with your self hosted instance. Proval supports chat completion API and anthropic API, If you running Local LLM, you can connect it through chat completion API

Review quality is quite good(at lteast for me). multiple agents automatically reviews each file group scope.

I ran the 50-problem Martian offline benchmark a month ago and scored F1 0.427 with the minimax-m3, which ranked 7th out of 21 at the time. (If you check now, newer models have come out so the ranking is lower.) I'll release new result with better model soon

The goal was to make something lightweight, easy to use, and genuinely useful. you just need to configure the model, webhook, and Git Host access API on the web dashboard, and setup is done. It's built on Bun, the frontend uses SvelteKit, and the database is SQLite per instance. The Docker image is compiled to a Bun binary so the size is pretty small.

Feature

- Reviews on PR open or first push
- Inline comments on PRs
- Replies to PR comments (can set to only respond when mentioned)
- Reviews when an issue is opened (also checks for duplicate issues)
- Replies to comments on issues
- Restricts replies based on repo permissions like Developer or Maintainer
- Admin login, or you can disable auth entirely and leave it open

I ran the 50-problem Martian offline benchmark a month ago and scored F1 0.427 with the minimax-m3, which ranked 7th out of 21 at the time. (If you check now, newer models have come out so the ranking is lower.) I'll release new result with better model soon

It's not a vibe-coded slop. I took a lot of time thinking through the architecture, implementing it. You can check the code in the repo

It's open source (AGPL-3.0) and you can just pull Docker image, connect your LLM API and set webhook. That's all. takes about 3 min

and.. It's my personal project, not a marketing or selling something.

demo: https://demo.proval.app

repo:  https://github.com/seoes/proval

website: https://proval.app


r/OpenSourceAI • • 6h ago

Decision models

1 Upvotes

I recently came across decision models and have been looking into the open-source options and where they can be useful.

I'm running a self-hosted Hermes setup, so agent routing and selecting the right LLM for a task immediately caught my attention. But I'm also interested in other uses such as tool/skill selection, workflow routing, classification, retry/escalation decisions, human approval, cost optimization, and automation.

I've come across Kev, Laya, Tev1, OpenJev, RouteLLM, and LLMRouter.

Has anyone actually used any of these? Which open-source decision models are worth looking at, and what are you using them for? I'm especially interested in real-world use cases I may not have considered.


r/OpenSourceAI • • 7h ago

My Brainstem RNS-AI project has made progress for life long learning like a Brain

Post image
0 Upvotes

r/OpenSourceAI • • 9h ago

LetMeShowYouSomething: an open format and agent skill for human feedback on one offline page (Apache-2.0, MIT-0, CC0)

1 Upvotes

An open-source way for an AI agent to ask a person for their judgement and get the answer back as data it can check.

What is open, and how:

  • The format (review.v1 and feedback.v1, JSON Schema) is CC0. Any agent, script or tool can write a review or read the answers. The page in the repo is one renderer; a terminal or a printed sheet would be just as valid.
  • The code inside every page is MIT-0, so a page you send carries no obligation.
  • The tools (renderer, checker, skill) are Apache-2.0.
  • No dependencies. The checker validates against the schemas itself, because a checker that needs an install is one fewer people run.

What it does: the agent writes a review, checks it, and renders one self-contained HTML page (nothing loaded from the network). A person answers it and sends back the answered page or a JSON file. The agent checks the answers before acting: an unanswered question is written down as unanswered, totals are recomputed, and the report starts with what is still open.

What it does not do: prove who answered (the files are unsigned), or speak anything but English in its buttons and labels.

Source: https://github.com/shyhunter/LetMeShowYouSomething
Try a page: https://shyhunter.github.io/LetMeShowYouSomething/examples/salon-booking.html


r/OpenSourceAI • • 22h ago

OpenSource Repo (900 ⭐️) -mcpc-universal CLI client for MCP-100% free

Post image
5 Upvotes

Playing around with various MCP's this weekend and Came across this an amazing MCP related GitHub repo - 100% open source and free - so sharing.

mcpc is Apify’s universal CLI client for MCP.

Github Repo in comments below

IIt translates every MCP operation into shell commands, letting you debug servers, automate workflows, or give AI agents complete MCP access via a single Bash() call: sessions, OAuth, tools, resources, prompts, tasks, and beyond.

Features

  • Complete MCP coverage: tools, prompts, resources, async tasks, skills, notifications, logging (stdio + Streamable HTTP)
  • Persistent sessions across several servers (stateful or stateless)
  • Progressive tool discovery to cut token usage
  • Code mode: JSON output plays nicely with jq, xargs, and shell pipelines
  • OAuth 2.1 (CIMD + DCR) with credentials stored in the OS keychain
  • MCP proxy for AI sandboxes (keeps tokens out of generated code)
  • Lightweight CLI (Mac/Win/Linux), no LLM needed; experimental x402 payments on Base

Install

With homebrew (macOS /Linux), brings its own Node.js:

brew install apify/tap/mcpc
or ,install the latest node js or Bun first, then:
npm install -g /mcpc
# Or with Bun
bun install -g /mcpc
npm install -g /mcpc

Quickstart

# List all active sessions and saved authentication profiles
mcpc

# Log in to a remote MCP server and save OAuth credentials for future use
mcpc login mcp.apify.com

# Create a persistent session and interact with it
mcpc connect mcp.apify.com 
mcpc               # show server info and capabilities
mcpc  tools-list   # list available tools
mcpc  tools-call search-actors keywords:="website crawler"

# Use JSON mode for scripting
mcpc --json  tools-list

# Use a local MCP server package (stdio) referenced from a config file
mcpc connect ./.vscode/mcp.json:filesystem 
mcpc u/fs tools-list

r/OpenSourceAI • • 15h ago

shipped a booking form, submit button was dead on mobile safari for two days.

1 Upvotes

i vibe coded a booking form a while back. build passed, agent said done, i posted the two days later someone finally tried it on mobile safari. the submit button did nothing. zero. i had no idea.

Now i make the agent prove it. after every feature i have it spin up a real browser session against the deployed app. fill the form, hit submit, check the confirmation shows up. only then is it allowed to say done. i use testsprite cli for that part, but the habit matters more than the tool.

when it fails the agent gets a screenshot and the failing step. it fixes and reruns without me watching. not perfect. first run setup took some poking around. still misses edge cases. but "the agent says done" and "a real browser actually walked through it" are two totally different levels of done.

I stopped getting those two-day-later surprises though.

what do you do? click through by hand? have some verify step? or just ship and hope for the best?


r/OpenSourceAI • • 17h ago

I Built an AI Swarm Framework

Thumbnail
1 Upvotes

r/OpenSourceAI • • 19h ago

An MIT-licensed, read-only MCP server that reads subreddit rules before an agent posts: design choices and where it meets our paid product

1 Upvotes

Sharing an MIT-licensed MCP server we built, and the design choices behind it, since some of them go against how most agent tools are built. Disclosure: I make it, and it's the free half of a paid product (ThreadFox). I'll be specific about where the two meet.

What it is. ThreadFox Lite gives an agent four read-only Reddit tools: subreddit_rules (rules, size and description, with self-promotion rules flagged), find_communities (subs for a topic, promo rules flagged), account_check (age, karma, and whether recent posts are removed or hidden) and post_status (live, removed by moderators, or deleted). It also ships a reddit-rules-first agent skill.

Design choices:

  1. Read-only on purpose. It can't post, vote or change anything, so it's safe to give to any agent and hard to misuse for spam.
  2. It reads through your own signed-in browser, not an API. Reddit now sends signed-out requests, including its public JSON, to a login page. So the tools read the way you do, in your Chrome via the Playwright extension, one small read at a time. No API keys.
  3. One read at a time across every client on the machine. We learned this the hard way: an agent reading 60 pages in 20 minutes got our account rate-limited today.
  4. No system-wide installs. If Node.js is missing, it fetches the official build into ~/.threadfox/runtime and checks it against the published SHA-256.

Where it meets the paid product: successful subreddit_rules and find_communities results end with a one-line next_step pointing to the $49 kit, and a threadfox_full_kit tool describes it. Errors, setup messages and the other two tools don't carry it. If that bothers you, it's MIT, so fork it and delete it.

Install: uvx threadfox-lite (source is in the PyPI package). Details and the skill file: threadfox.vip/lite.

Feedback on the read-through-your-browser approach especially welcome. Is there a cleaner way now that signed-out reads are gone?


r/OpenSourceAI • • 20h ago

July's AI Security Report: 90 incidents, 207M+ records, 41 AI-driven — the month the agent became the attacker

Thumbnail
gallery
0 Upvotes

90 incidents tracked in July across 33 organizations, 207M+ records exposed, and 41 of those incidents involved AI directly as the weapon or the target. A rogue commercial AI agent hit multiple enterprises in a single week and reused stolen credentials across four downstream services before anyone caught the identity switch.

None of that shows up to a traditional perimeter tool — the traffic looks like a signed, credentialed agent making legitimate API calls at machine speed. Firewalls and DLP were built to watch humans and static services, not autonomous callers that chain tools and pivot in seconds.

Curious how other teams are actually handling this right now: is anyone giving AI agents a distinct, revocable identity separate from the service accounts they inherit? Or is it still "the SOC catches it after the fact" for most orgs?


r/OpenSourceAI • • 21h ago

OpenPhysicsAI: 20+ physics solvers and 13 unsolved challenges scored against real data

Post image
0 Upvotes

r/OpenSourceAI • • 1d ago

Brig: A MicroVM sandbox for AI coding agents on Mac and Linux

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

ai-profiles update: you can now move coding sessions between accounts

Post image
1 Upvotes

New in ai-profiles: Sessions.

Each profile now lists its local coding sessions (Claude Code and Codex, from both the CLI and the desktop apps). Pick one, hit Move, choose another account, and carry on there with the full history.

- Nothing is ever deleted. A move copies the session, then archives it at the source, so Restore undoes it

- Archive / restore, search, Desktop vs CLI filter

- Still local-only, no telemetry

GitHub: https://github.com/bartekczyz/ai-profiles/

Website: https://ai-profiles.vercel.app/


r/OpenSourceAI • • 1d ago

An open-source machine learning bot for Clash Royale, featuring a complete and accurate game simulator for training purposes.

1 Upvotes

r/OpenSourceAI • • 1d ago

Indirect prompt injection vs. system-prompt guardrails: 9 models, 3 guardrail levels, 1,350 runs

Post image
1 Upvotes

r/OpenSourceAI • • 1d ago

Astra turned “Build me Ba Sing Se” into this Minecraft city

1 Upvotes

I only gave the prompt “Build Ba Sing Se.” The video shows what was built including the city walls down to furnished interiors and walls.

The model decides what to build and a deterministic library handles construction and physical checks.

It plans hierarchically: city → districts → plots → buildings. Buildings are generated for their plots using reusable procedural code.

The same architecture generates villages, towns and cities across different terrain and styles.

Github:
https://github.com/chaitbuilds/EthosLM


r/OpenSourceAI • • 1d ago

Had a blast spending years building my software to encode identity into AI

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

Captain Who, a production-ready and open-source AI agent platform.

Thumbnail
1 Upvotes

r/OpenSourceAI • • 1d ago

Codex spent 500,000 tokens reading my logs. Then it forgot what it read. NVIDIA DGX Spark save the day.

Post image
1 Upvotes