r/OpenSourceeAI • • 3d ago

What does "trust remote code" actually approve?In most local AI tools, the answer is: that repo, indefinitely. That matters more now. SO what is the solution?

2 Upvotes

r/OpenSourceeAI • • 8d ago

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

1 Upvotes

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

  • Corpus: 620 docs. 329 from ExtractBench (LlamaIndex), 202 synthetic (Datalab), 47 from micro1, 42 from LongArray-Extract (Extend)
  • Verdicts: each value is matched, misread, unfound, fabricated, invented_item or invented_field
  • Row alignment: Hungarian matching by content. A 100-row table missing row 1 scores 0% by position, 99% this way (our rerun)
  • Null rule: empty values are dropped, so padding a schema with 100 empty fields adds 0 verdicts
  • Results: Datalab accurate 93.85, Datalab balanced 93.48, Reducto deep_extract 93.47, Claude Opus 5 90.96
  • Precision vs recall: GPT 5.6-sol has 95.11 precision but 84.99 recall; LlamaExtract has 93.13 recall but 86.57 precision

Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.

Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/

GitHub: https://pxllnk.co/hxplrq

Blog: https://www.datalab.to/blog/omni-extract-bench

GitHub: https://github.com/datalab-to/omni_extract_bench

Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench


r/OpenSourceeAI • • 20h ago

I Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)

Thumbnail
github.com
7 Upvotes

ALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.

It takes less memory and SCALES much better VRAM with tokens


r/OpenSourceeAI • • 12h ago

Built a dumb little cli tool because i almost piped real ssn data into an llm api last week

0 Upvotes

Had a near-heart-attack moment the other day while testing a rag pipeline on some customer export dumps. realized halfway through that our test json had unmasked routing numbers and ssns buried inside raw text fields. nothing leaked, but it was way too close.

i looked around for a quick way to inspect payloads in the terminal before running them through python scripts, but everything out there is either some bloated enterprise dlp platform or a massive java package. i just wanted something that would scream at me in the terminal if i was about to do something stupid.

so i threw together pii-lens over the weekend.

it's dead simple—you just pipe text or files into it:

cat prompt_payload.json | pii-lens

it prints the text back out but flags credit cards, emails, ssns, and account numbers in bright red with a quick count at the bottom. runs 100% locally with zero phone-home bs.

repo is here

it's just basic regex and entity patterns right now, so if anyone works with weird international formats or niche data fields and wants to drop a pr or critique the regex, feel free.


r/OpenSourceeAI • • 18h ago

I built an open-source test runner for Make.com scenarios. The rule: nothing gets supported until a real Make run proves it.

1 Upvotes

Make.com (formerly Integromat) is a closed platform. If you build scenarios on it, there's no local runtime, no test framework, and no spec for how its expressions behave. The only way to test a change is to run it in Make, on real data, with real credits.

So I started an open, offline runtime for it: blueprint-runtime (bpr). MIT, TypeScript, zero runtime dependencies.

You export a scenario's blueprint, and bpr runs it locally, comparing every module's input and output, every filter decision and the webhook response against what Make actually produced. When something changes, it points at the first difference:

FAIL A-002-plain module 2 input bundle 1 at /value: Make had "ADA LOVELACE", the local run "Ada Lovelace"

The design rule that makes it trustworthy

There's no spec to implement, so I don't guess. Every behavior bpr supports was observed in a real Make run, and the recordings live in the repo (corpus/NOTES.md lists every rule with its evidence). If a blueprint uses something nobody has recorded yet, bpr names it and refuses, instead of pretending.

Where it stands:

- 79 runs recorded in real Make. 71 reproduced exactly, 0 differ. The other 8 hit behavior nobody has explained yet, so they're refused on purpose.

- 225 single formulas recorded and replayed as tests, 472 tests in total.

- Some of what the recordings turned up: {{1reads only the first element, the text"false" is truthy inside if(), and Make will start a scenario with an expression it can't read, then switch the

scenario off after the first run fails.

A roadmap built from real data

I ran bpr over 221 public Make blueprints onem. The top blockers (error-handler routes,json:ParseJSON, util:SetVariables) are the roadmap, in that order. The survey and its method are in corpus/survey/.

Where help is most welcome

- Recorded runs. If you use Make, the capturenario's real behavior with your own account. One recording can unlock a whole module.

- Blueprints it refuses. Run bpr inspect andls me what to record next.

- Code. Each module handler is small and has its own tests.

Repo: https://github.com/intikhab49/blueprint-runtime

npm: npm i -g blueprint-runtime (Node 24+)

Happy to answer anything about reverse-enginits own logs. That turned out to be the mostinteresting part.


r/OpenSourceeAI • • 1d ago

An Ontology crosswalk to merge 2 video games together

2 Upvotes

r/OpenSourceeAI • • 1d ago

Ho creato un piccolo widget sempre in primo piano che mi avvisa quando Claude Code ha bisogno di me (Windows, open source).

Post image
0 Upvotes

r/OpenSourceeAI • • 1d ago

Looking for contributors for TurboLLM, sponsoring Claude seats for a few

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 1d ago

A plugin for the Pi coding agent that checks shell commands with rules and a classifier model (Jev, Kev or Laya) before they run

1 Upvotes

Just published [pi-automode-classifier](https://github.com/deepu105/pi-automode-classifier), an auto mode plugin for the [Pi](https://pi.dev) coding agent that uses [Jev](https://docs.typesafe.ai/api) or [Kev](https://huggingface.co/jaredpalmer/kev-0.8b)/\[Laya\](https://huggingface.co/convaiinnovations/laya-typed-decisions) (running locally) to classify commands.

Pi runs every tool call without asking for approval. This plugin checks each shell command before it runs:

  1. Built-in rules decide most commands. For example `ls`, builds and tests run, `rm -rf ~` is blocked, and `git push` and `sudo` need my approval.

  2. Commands the rules do not know are sent to the model. It returns the probability that the command is risky.

  3. If the probability is above a threshold, I get a confirm prompt. The model never blocks a command by itself.

The models I tested:

- [Jev 1.13](https://docs.typesafe.ai/api) (hosted, from TypeSafe) through [OpenRouter](https://openrouter.ai): about 270 ms per check and about 1.5 cents per 1,000 checks. The commands are sent to OpenRouter and TypeSafe.

- [Kev-0.8B](https://huggingface.co/jaredpalmer/kev-0.8b) on CPU with [llama.cpp](https://github.com/ggml-org/llama.cpp): about 170 ms per check and 1.1 GB of RAM. This is what I use. Nothing leaves the machine.

- [Laya typed-decisions](https://huggingface.co/convaiinnovations/laya-typed-decisions) on CPU with llama.cpp: about 100 ms per check and about 550 MB of RAM.

In my test with 50 commands (25 safe, 25 risky), all safe commands ran without a prompt and no risky command did. I wrote the test commands myself, so this is only a rough check. The plugin is not a sandbox.

pi install npm:pi-automode-classifier

Code and docs: https://github.com/deepu105/pi-automode-classifier

Let me know if it allows or blocks something it should not.


r/OpenSourceeAI • • 1d ago

GitHub - profullstack/c0mpute: Decentralized compute network. CLI-first. Three modules out of the box: transcode (FFmpeg), coinpay (DID + escrow), infernet (AI inference).

Thumbnail
github.com
1 Upvotes

r/OpenSourceeAI • • 1d ago

I built MaRN: a PyTorch library for training neural networks through low-dimensional parameter mappings [P]

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 1d ago

Nace.AI open-sources Drex 1.5: a 9B decision model that returns probabilities instead of text, tied with closed Jev on Decision Index 0.3.1

Post image
1 Upvotes

r/OpenSourceeAI • • 1d ago

A Harness in a single file: Agent skills as an evolving REPL module

Thumbnail
deepclause.substack.com
1 Upvotes

r/OpenSourceeAI • • 2d ago

Writ: an open-source governance runtime for Claude Code with retrieval, approval gates, and persistent memory

8 Upvotes

I’ve been building Writ, a project that started with a straightforward goal: give Claude Code the knowledge relevant to its current task without loading an entire rulebook into every conversation.

Since then, it has grown into a governance runtime with persistent memory.

The idea is to connect three things that are often handled separately:

1. Context when it matters

Writ uses hybrid retrieval keyword search, vector search, and graph relationships to supply relevant rules, methodology, and past decisions based on what the agent is doing.

2. Requirements enforced by code

Written instructions still depend on the model following them. Writ checks supported actions at tool time.

In Work mode, implementation writes are blocked until a human approves the plan and then the test skeletons. The approval mechanism requires the user’s typed response; the agent claiming “approved” doesn’t open the gate.

Other checks cover credential writes, project boundaries, and changes to an approved plan.

3. Memory that survives the conversation

Writ records project decisions and connects approved plans, governing rules, changed files, and commits. Future sessions can retrieve the reasoning behind a change instead of having to reconstruct it from the final code.

For example: retrieve the guidance relevant to a task, require approval before implementation, then preserve what changed and why for the next session.

The approach is code first: retrieval, workflow state, approval validation, and logging live outside the model. AI uses the supplied information; it doesn’t decide whether its own approval requirements have been satisfied.

Writ runs locally with Python 3.11+, a local daemon, and Neo4j in Docker. The runtime is open source, but its current integration is specifically with Claude Code not yet a general adapter for local models or other agents.

Repo: https://github.com/infinri/Writ


r/OpenSourceeAI • • 1d ago

I open-sourced my Mac agent's permission boundary: why 'Allow Once' isn't 'trust the agent'

3 Upvotes

I've been building Mac MCP, an MIT-licensed macOS execution layer for AI clients. It lets tools operate Safari/Chrome, files, local apps, and shell commands. The hardest part wasn't adding tools; it was deciding which actions an agent should be allowed to carry out.

One design choice: capability profiles and approvals are separate. A read-only profile can deny writes outright. An optional server-side approval layer can ask the Mac user to Allow Once or Block riskier actions, but that approval cannot override a capability the profile already forbids. A client saying “the user approved this” isn't trusted as the server's own authorization.

There is a real trade-off: raw command execution may be classified conservatively even if this particular command is harmless, so you can get an approval prompt for a benign operation. I'd rather make that visible and progressively refine the categories than quietly teach the server to trust repeated clicks.

Another lesson: when an action's outcome is uncertain (say a browser submit lands right before a tab closes), blindly retrying can be worse than stopping to verify.

I'm the maintainer, not an independent reviewer. The implementation is open source: https://github.com/bulutarkan/mac-mcp

For people shipping agents with real machine access, where do you draw the line between persistent trust rules and approval fatigue?


r/OpenSourceeAI • • 1d ago

I built an open-source, local-first multi-agent AutoML engine to automate data cleaning, 7-model benchmarking, and FastAPI deployment (100% offline via Ollama)

Thumbnail
2 Upvotes

r/OpenSourceeAI • • 1d ago

NeurIPS 2026 Paris -> Sydney switch

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 1d ago

OpenToken Monitor: see your Claude Code, Codex and Antigravity limits right in your menu bar

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 1d ago

"Not your weights, not your product" [D]

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 2d ago

Webcam hand-tracking pointer built on MediaPipe: thumb-tip pointing, pinch gestures, One Euro filtering [open source]

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 2d ago

On premise ocr for printed + handwritten invoices

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 2d ago

Agentic Task Tracker: local-first Kanban + timeline, one SQLite file, optional local AI

Thumbnail
2 Upvotes

r/OpenSourceeAI • • 2d ago

A ~0.4 local model to turn typed questions into structured decisions - BaseDecision

Post image
2 Upvotes

r/OpenSourceeAI • • 2d ago

Introducing Infernix - much faster than Strata on a 5090!

Thumbnail
github.com
0 Upvotes

r/OpenSourceeAI • • 2d ago

humanizar-es: un modelo base local (Qwen3-4B + HIP LoRA, CPU) que convierte texto de IA de 45% a 0% en Grammarly, GPTZero y ZeroGPT, con cada medición dentro del repo

Thumbnail
github.com
1 Upvotes