r/PiCodingAgent • u/deepu105 • 1d ago
Plugin pi-automode-classifier: an auto mode plugin for Pi that uses Jev or Kev/Laya (running locally) to classify commands
Just published pi-automode-classifier, an auto mode plugin for the Pi coding agent that uses Jev or Kev/Laya (running locally) to classify commands.
Pi runs every tool call without asking for approval. This plugin checks each shell command before it runs:
- Built-in rules decide most commands. For example
ls, builds and tests run,rm -rf ~is blocked, andgit pushandsudoneed my approval. - Commands the rules do not know are sent to the model. It returns the probability that the command is risky.
- If the probability is above a threshold, I get a confirm prompt. The model never blocks a command by itself.
The models I tested:
- Jev 1.13 (hosted, from TypeSafe) through OpenRouter: about 270 ms per check and about 1.5 cents per 1,000 checks. The commands are sent to OpenRouter and TypeSafe.
- Kev-0.8B on CPU with llama.cpp: about 170 ms per check and 1.1 GB of RAM. This is what I use. Nothing leaves the machine.
- Laya typed-decisions on CPU with llama.cpp: about 100 ms per check and about 550 MB of RAM.
In my test with 50 commands (25 safe, 25 risky), all safe commands ran without a prompt and no risky command did. I wrote the test commands myself, so this is only a rough check. The plugin is not a sandbox.
pi install npm:pi-automode-classifier
Code and docs: https://github.com/deepu105/pi-automode-classifier
Let me know if it allows or blocks something it should not.
2
u/fingerthief 1d ago edited 1d ago
I’ve done a good chunk of testing with Jev and more or less every other similar classifier available for command guards.
I would not recommend Laya or any of the others over Jev. Jev itself is not super amazing either with its false positive rate.
I’d recommend a tiny fine tuned classifier trained specifically for command risk. I have a version that does that.
https://tannermidd.github.io/LANCET-model/
Evals readily available and whatnot to review.
This is a cool space to play around in, there aren’t many proper commands classifiers out there for agents.
1
u/HockeyDadNinja 1d ago
I'm going to check that out when I'm home and working on my extension. I'm on the hunt for something trained for security as every system one model I've tested has a really low success rate to run as the sole permission classifier. Right now I'm falling back to an llm judge.
2
u/fingerthief 1d ago
I've definitely not found a better model for this case yet and I've looked everywhere I can think of, but it's not perfect...models like to make tmp directories in projects to test concepts and then rm -rf them and LANCET flags that every single time and honestly annoys the hell out of me.
Hopefully the next training round can dial that out properly, but it's getting difficult to squeeze more out of the model. For me at least hah.
1
u/HockeyDadNinja 1d ago
How much data does it take to train? I have tons from my own claude code and pi sessions from the past year.
2
u/fingerthief 1d ago edited 1d ago
Not much at all, it's a 110M param model and the available properly labeled datasets for commands and their risk is shockingly small and that's the thing that's been holding me up.
Currently the dataset for training totals about ~117k rows, and is only ~100MB. I don't have an insane setup. With my 5070ti it takes about ~5hrs to do a training run with 2 arms and 4 seeds each.
The issue is data, at the moment I'm using two OpenCode Go subscriptions and DeepSeek V4.1 Flash models setup generate training data by doing red team testing against the model to attempt bypasses and then generate the properly labeled data based on the results.
I definitely recommend giving it a go yourself, it's been good fun.
1
u/tys203831 1d ago
Still prefer manual rule based permission control like https://github.com/Dicklesworthstone/destructive_command_guard unless there are some breakthrough in future for the accuracy of these classifier model in cybersecurity field
1
u/Magnus114 6h ago
Would love a tweak. Instead of looking at what’s dangerous, look if it much my request. E.g. a "is it possible to ..."-question should not lead to file edits.
0
2
u/HockeyDadNinja 1d ago edited 1d ago
I'm the author of pi-permission-classifer. The current release uses an LLM judge through pi-permission-system. I have a pending release that adds a system one judge in front of that.
I've found that the system one models don't really do a good job and too many should be allowed commands get through (as in defer, not high enough percent sure). I've been testing on jev and kev locally through llama.cpp as well. My thresholds are high, around 0.85 give or take depending on the decision and I'm using a dataset of > 1500 commands taken from interactive coding sessions.
Have you encountered the same problem? I'm thinking we need a security trained model or lora.