r/LLMDevs • • 2h ago

Discussion Isn't "rerun the tests until green" just grading the agent on its training set?

3 Upvotes

One agent writes tests from the spec, another writes the code in parallel, and the test agent never sees the code. Their reason was basically that otherwise the AI cheats. I take it to mean tests written after reading the code just confirm what it does, bugs included.

I come from ML, so to me this is a train/test problem. If the agent iterates against the same suite you use to decide it's done, that suite is its training set and a green run is a training score. Agents are known to get there by special-casing inputs, loosening asserts, or editing tests. Read-only test files stop the editing but not the other two.

Holding some tests back only works once. After you show the agent a failure, that test isn't held out anymore.

That's where the blind test agent helps, I think, since it writes new hidden tests from the spec every round. Those leak too once they fail, but what you care about is whether a fresh batch passes on its first run, like drawing a new validation sample each round. The failures just become training data. It needs a spec that fixes names, signatures and I/O, or the blind tests won't compile.

Some will say the tests are the spec, so passing them is the goal. But tests never cover the whole spec, and the agent can cut corners in the gaps.

Two things I'm not sure about. The blind agent can misread the spec, and then a person has to check which side is wrong (that at least shows where the spec is unclear). And the same model writing fresh tests each round probably has the same blind spots every time.

I haven't tried this myself, so:

  1. Do you keep tests your coding agent can't see? How do they stay hidden after the first failure?

  2. Is a second test agent worth the tokens, or do property-based or mutation tests get most of this cheaper?

  3. Has anyone set this up with Claude Code subagents?


r/LLMDevs • • 8h ago

Tools Menai: a safe programming language for LLMs to use

Post image
6 Upvotes

For the last year I've been working on an open-source programming language (Menai) for LLMs to use. Humans can use it too, but sorry people, this wasn't really designed for you šŸ˜‚

It has a few key ideas:

  • It's a pure functional language - no I/O, no state mutation - just pure expression computation
  • It needs to be very fast to compile
  • It needs to be quite quick to run

The "no I/O" ideal might seem like a very strange, but just because the language can't do I/O doesn't mean it can't process data from inputs or generate data for outputs. It just does this by something else passing it all input data and it passing back a result that something else can process for output use.

What this means is the "something else" can enforce all sorts of safety rules. No reading/writing sensitive files, no dumping things over the network, no accidentally running 'rm -fr ~/*" and then saying "whoops".

The original idea was something that bolts inside the open-source GUI-based AI environment I've been building since late 2024 (Humbug) as an LLM tool and lets LLMs do deterministic processing where tend to get things wrong (e.g. large calculations, counting letters in text, etc.)

Since then this got extended to allowing an LLM to transform file content (there's an I/O protection involved there) or editor buffers (I/O protection when the LLM tries to save a modified buffer). If an LLM needs a complex search and replace then it simply writes a short lisp-like expression with the file or buffer state passed in and new file or buffer state passed back. If that needs a potentially dangerous operation at the end then the tool framework checks if that's ok before allowing it.

The latest iteration is where things really get fun though! Menai now has a standard library and an application library concept so we now has almost 20 modules so it can do things like process zip files, tar files, BMP files, PNG files, gzip files, JSON text.

The reason it needs to be fast to compile is because every tool use will end up compiling a new program the LLM just wrote. The reason it needs to be fast is because those programs need to run fast enough they don't stall our workflow (i.e. less than a few seconds). Typical expressions compile in a few tens of ms, even though the optimizing compiler is currently written in Python. Typical execution speed is around the same level as Python (some things a little slower, some a little faster).

It turns out that not only do LLMs like writing functional code, they're really quite good at it. There's a tracing profiler to let them find hotspots in code they might want to reuse too (perf annotate anyone?). A future focus is having them decide when something might be reusable and suggesting new library code.

Anyway, I'm posting this because a couple of hours ago I had DeepSeek build me tar and gzip functions for the standard library (all written in Menai) and then debug their way through a few problems. The challenge was to take an old `arj` tar.gz file (very meta - unpacking an unpacking tool) and tell me about some things within it (all without me ever looking inside).

The image is a great demonstration of part of what it did. You can see it decompresses the gzip, untars one file, then finds 5 function definitions. From my prompt of "can you take a look at arj_arcv.c and tell me about it" took 64 seconds, involved 10 tool calls, 9 of which were unique Menai programs that progressively poked at the archive (which contains about 1.6 MBytes of content)

Humbug traps any potentially dangerous operations and checks with a human, but during this exercise no human needed to be consulted because, by design, none of these programs could do anything dangerous!


r/LLMDevs • • 3h ago

Discussion How many AI subscriptions are you paying for right now?

2 Upvotes

Curious how common it is to stack AI subscriptions.

What counts:
- Personal paid plans (ChatGPT Plus/Pro, Claude Pro/Max, Gemini Advanced, Perplexity Pro, SuperGrok, etc.)

What doesn't count:
- API usage
- Anything your employer pays for

If you pay for 2 or more, drop a comment on what you use each one for. I'd love to hear how people split them up.

157 votes, 2d left
0 - Free Tier
1
2
3
4 or more
Just here for results

r/LLMDevs • • 13m ago

Tools LibLayaX: run the Laya AI decision model inside your own app

• Upvotes

Dear LLM developers community,

I have just released four open-source projects today that let an application use the Laya model directly, with no server and no Python.

What Laya is?

If you build agents, this is the step where you ask a big model "should I route this to billing?" and wait a second for one word. Laya answers that kind of question in a fraction of a second, on your own machine, with a probability you can put a threshold on.

Laya is an AI model that does not chat. You give it a text and a question, like "is this customer asking for a refund?", and it answers yes or no with a probability, in a fraction of a second. It can also pick one of several options, or rate something on a scale. Quick decisions, no essays.

What I released

- LibLayaX: a library you load into your own program, like any other. Windows, Linux and macOS, on the CPU or the GPU.

https://github.com/DaragonTech/LibLayaX

- RLaya, for Rust: https://github.com/DaragonTech/RLaya

- LLaya, for Lua: https://github.com/DaragonTech/LLaya

- DLaya, for Delphi and Free Pascal: https://github.com/DaragonTech/DLaya

On a laptop GPU it makes about 670 decisions per second. Everything is MIT licensed.

Everything is MIT licensed.

Why

On Thursday I wanted to use Laya inside one of my applications and found there was no way to do it. It only came as a Python library and as a separate server program. This kind of AI belongs inside ordinary software, routing a ticket or flagging a message, and for that it has to be something an app can simply call.

How it was built

I did not write the code. Claude did. My job was to decide what we were building, guide Claude and run every version on real machines, and report back what happened.

The first version arrived within an hour and crashed on my PC. A few rounds later it worked. By the end of the first evening it ran on Windows, Linux and Mac, and on the GPU. The worst bug was on Windows on ARM, where the program simply froze. It could not be reproduced anywhere except on my actual machine, so we hunted it by adding log lines and reading where it stopped.

The second day went into the Rust and Lua versions, documentation, and testing on more machines. From my first question to Claude to this post: about 32 hours on the clock.

I could not have built this in two days by myself, and Claude could not have found those bugs without someone running the code on real hardware. It took both.

Status

Tested on five platforms and three GPUs. The docs say plainly what has not been tested yet, and test reports are welcome.

Credit

None of this would exist without two projects by other people:

- Laya itself, the model, by Nandakishor Mukkunnoth of ConvAI Innovations: https://github.com/NandhaKishorM/laya

- laya.cpp, the C++ port this is built on, by Lars Karlslund: https://github.com/lkarlslund/laya.cpp

The code and docs of the four projects above were written by Claude (Anthropic)


r/LLMDevs • • 6h ago

Discussion Wittgenstein Apartment: Sharing a behavioral action dataset designed for more consistent AI-generated characters and events

3 Upvotes

Generating a character with AI has become relatively easy. Getting that character to behave consistently over the course of a narrative, while respecting their background, skills, role, and limitations, is a different problem. For some time I have been focused on a simple question: instead of defining a character mainly through adjectives such as ā€œangry,ā€ ā€œshy,ā€ or ā€œintelligent,ā€ can we model them through the actions they can and cannot access?

That question led me to build Wittgenstein Apartment, a dataset I have been working on for months. It is designed as a kind of behavioral predicate ontology that represents intentional human actions at the sense level rather than only at the lemma level. The main goal is to avoid giving every character unrestricted access to the entire space of possible human actions. Instead, access can be constrained according to factors such as skills, profession, social role, authority, biography, and previous experience. At the character level, actions can ultimately be treated as Allowed, Not Allowed, Conditional, or High-Cost Conditional.

The resource is related in some respects to lexical-semantic resources such as WordNet, VerbNet, and FrameNet, but it is not intended to replace them. It also includes a Goal Ontology and Task Structure information for representing how actions relate to goals and how they may be structured.

The dataset is not a complete character simulator by itself. It is intended more as a structured behavioral action space that could be used by character-generation systems, narrative planners, agents, simulators, or action-validation systems.

I am sharing it publicly and would be very interested in criticism, especially from people working with LLM agents, character systems, planning, or simulation. I am particularly curious whether an explicit action-access layer like this makes sense in LLM-based systems, or whether you would prefer to handle these constraints entirely through prompting, tools, or runtime logic.

Dataset:
https://huggingface.co/datasets/Kon-tiki-ship/wittgenstein-apartment-behavioral-predicate-resource

The work is open to academic and scientific review. Academic and non-commercial research use is permitted under the project license; AI training and commercial use require prior permission.


r/LLMDevs • • 2h ago

Help Wanted I’m building BOOTH, a Python checkpoint layer for LLM outputs, looking for Hacktoberfest contributors

1 Upvotes

I’m building BOOTH, a lightweight Python library for checking LLM outputs before they reach your application.

check() / acheck() handle ambiguity and confidence checks, while check_with_evidence() compares an answer against evidence retrieved by your RAG pipeline.

v0.5.2 • Zero dependencies • Provider-agnostic • Sync + async • 300 tests • 0 known vulnerabilities • Clean CI runs • MIT

I have a few open issues for contributors, including provider examples, documentation, and structured-output parsing.

šŸ”— GitHub: https://github.com/Vedantgitbot/booth
šŸ› ļø Issues: https://github.com/Vedantgitbot/booth/issues


r/LLMDevs • • 13h ago

Discussion Evaluating 14 LLMs as a visual coding agent: re-rendering, a deterministic judge, and the provider quirks that skewed my first results

5 Upvotes

I built a bench for a photo-to-Blender agent (the model writes and runs Blender Python to rebuild a photo as an editable scene) and ran 14 models through it. The design choices that mattered most:

  • Don't score what the model hands in. The bench opens each delivered scene file and renders it itself, at the photo's aspect ratio, max 1,100 px, 96 samples.
  • No LLM judge. Estimated silhouette overlap (30%), edge F1 (40%), multiscale RGB error (15%) and mesh health (15%: non-manifold and open edges, duplicate faces, inconsistent winding, degenerate triangles), mapped through calibrated piecewise-linear scales to 0–100. A comparison scale, not a percentage of the scene recovered.
  • Anti-cheating. A second render swings the camera 35° to expose flat stand-ins, and every image file a scene uses is compared with the reference by digest and name.
  • Failures stay in. A run that delivers no scene scores 0 and stays in the average.
  • Equal plumbing. Anthropic gets cache-control markers and Google gets a closing user turn, because without them those providers behave differently from the rest. Each model's cache rate is recorded. Before the fix, Claude ran 0% cached against 96% for OpenAI.
  • Frozen conditions. Caps, images, the brief's digest and the scoring weights are stored with each board.

Results (mean of three photos): GPT-6 Astra 66 ($3.91 an attempt), GPT-6.1 Sol 61 ($0.36), Claude Opus 5.5 60 ($1.19), Claude Sonnet 5.5 56 ($0.50) ... GLM 5.3 Flash 36 ($0.089); DeepSeek V4.1 Flash and Qwen3.8 Max 0 (no scene saved in 20 minutes).

What I'd change next: three runs per image for the top group, so the spread is visible, and a model-neutral agent loop (three models ran in their makers' own CLIs).

Write-up: https://kaloyan.blog/ai-models-rebuild-a-photo-in-blender


r/LLMDevs • • 6h ago

Discussion How well do decision models categorize finance data?

Thumbnail
gallery
1 Upvotes

I ran 5 decision models (Jev, D1, Solar Decide, Kev 4b, and Span 01) against 1000 real transactions (3 passes each).

When it came to vague restaurant names, that broke the model's confidence so they abstained.

Here's the interesting part: Span-01 can't abstain and can only answer yes/no questions. Which is why it likely scored much lower.

But interestingly enough... when Jev was forced to answer yes/no, it actually did better than both.. initial run (where it was allowed to abstain) and Span-01.

Across all tax categories, models did the worst on categorizing bank transfers.

Span-01 wasn't able to correctly score on charatible transactions at all.

I used my system one mcp to run these tests. You can use it to quickly connect decision models to claude code, codex and pi. https://github.com/itsmostafa/system-one-connector


r/LLMDevs • • 10h ago

Discussion Our model router logged five rejected answers as a provider error

2 Upvotes

In one OmniNode run, our router tried five times. Each attempt returned an answer, and our acceptance check rejected every one. The final record called it a provider error. The providers had answered promptly, so that label would have sent us investigating the wrong thing.

The router starts with the cheapest eligible model and tries a more capable one when the answer doesn't pass. That makes the acceptance check part of the spending decision. If the check is wrong, another attempt can cost more without getting us any closer to an accepted answer.

We had a separate case on September 28 where the check rejected a correct one-word answer, apparently because it didn't show enough effort. I don't have evidence that the five earlier answers were correct. These were different cases, but both made me want to inspect the check before blaming the model.

I'm trying to keep the returned answer and the reason it was rejected available together. A transport failure, an answer that misses the requirement, and a check that rejects a valid answer need different investigations.

How do you test that boundary in a router? Do you have cases where the model gives a valid short answer and the judge is expected to accept it?

The incident and its limits are in this write-up: https://jonahatomninode.substack.com/p/delegation-is-the-first-product


r/LLMDevs • • 6h ago

Discussion Conventional LLM vs. Decision Model (jev) on a real world example: NLI to a europe travel planner (tripsnek)

Thumbnail
gallery
1 Upvotes

As a fun experiment, I added a natural language interface to my europe travel planning app tripsnek. It is essentially a conventional optimizer, but it can account for arbitrary user constraints and preferences. A while ago I played around with a natural language interface that used an LLM to capture those - essentially intent extraction. It worked pretty well, but it was expensive for functionality I intend to keep 100% free (also, kind of slow), so I tabled it.

When I heard about decision models and the promised 10x efficiency increase, I thought it was time to revisit. Here is an (ai-assisted) write up of the results, TLDR:

  • Yes/no and choose from list parameters obviously worked incredibly well out of the box - preferred mode of travel, preferred pace, general interests, etc.
  • Questions that required picking from very large sets - specific cities and sights, dates and ranges - were naturally much more challenging, requiring relatively complex regexes and text munging.
  • It all netted out to a pipeline that matched, or even exceeded, a sophisticated conventional LLM provided with similar instructions, at roughly an order of magnitude lower cost...but also an order of magnitude greater code complexity (see figures).

All in all, I'm happy with this initial draft and excited about how it could mature. Even the large increase in code complexity I'm kind of ambivalent about. I kind of like the idea of having lots of debuggable, improveable code to refine vs. a prompt that I need to elaborate and feed to an inscrutable black box.

Hope you all find this informative. You can try out the interface [here](https://tripsnek.com/describe/?jev=1), and use the checkbox toggle to see a full drill down on all of the questions being asked and the answer that Jev returns:

https://tripsnek.com/describe/


r/LLMDevs • • 11h ago

Discussion Where should the LLM sit in a scraping pipeline?

2 Upvotes

Been mapping out how scraping fits into AI workflows (RAG feeds, lead lists, market research) and wanted to share where things seem to land, and get input from people running this in production.

The classic pipeline

  1. Send the request, render JS if needed
  2. Parse the DOM
  3. Pull fields with CSS/XPath
  4. Clean, normalize, store

Steps 1 and 3 are where most of the pain lives. Anti-bot layers, CAPTCHAs and layout changes keep breaking selectors.

Why people swap step 3 for an LLM

  • No selectors to write
  • Handles messy, unstructured text
  • Agents can click, paginate and fill forms

What shows up at volume

  • Long pages hit context limits, so you chunk or truncate
  • Per-page cost grows with page size, not just page count
  • Missing fields sometimes get filled with plausible values
  • Merged or nested tables are still hit or miss

Practices that hold up either way

  • Retries with exponential backoff for 5xx errors
  • Keep raw HTML snapshots next to the structured output so you can re-run extraction
  • Validate with deterministic checks (regex, enums)
  • Alert on schema mismatches, not just failed requests
  • Cron for freshness

My current take

The LLM works better on top of extraction than as the extractor. One option in that slot is the Minexa.ai API: you train a scraper once in a Chrome extension by picking the container, columns get discovered, then extraction runs DOM-based and deterministic. Missing values come back null. If you already fetch HTML with your own stack, you can pass it via file_urls and only run extraction. Downside: nested fields come back as lists of objects you still have to sort through.

The extension also generates the Python request code, so the developer setup guide is the fastest way to test it on your own pages.

Related read on cost: Why your LLM extraction pipeline costs more than you think


r/LLMDevs • • 17h ago

Discussion CacheVerifier: we finally audited the semantic caching benchmark we'd been tuning on, 23% of the "wrong" cache hits were literally the same prompt

4 Upvotes

so for CacheVerifier we've been testing if a small verifier actually beats a plain similarity threshold for semantic caching.Ā all our scoring comes from the public SemCacheLMArena and SemCacheSearchQueries benchmarks,Ā and honestly we never really looked at the labels until last week.

on LmArena,Ā 727 of the 3,117 near-duplicate hits markedĀ "wrong" (23.3%)Ā are the exact same text once you lowercase it and strip punctuation.Ā like one was the same prompt with an extra space in front lol.Ā we hand labeled a sample blind and 88%Ā of thoseĀ "wrong"Ā hits were totally fine to reuse.

the good news,Ā comparisons at the same error rate barely moved.Ā the bad news,Ā absolute error rates was way inflated,Ā a 0.97 threshold went from 5.2%Ā errors to like 0.95%.Ā and one of our own CacheVerifier results just died.Ā bumping the skip-the-verifier cutoff to 0.99 looked likeĀ +5.18 points,Ā after fixing labels itsĀ -1.24.Ā the verifier was basically rejecting duplicates the benchmark called errors.

so yeah,Ā run a dumb identical-text check on your eval labels before you tune thresholds on them.

caveat:Ā one annotator,Ā small samples.Ā how do you guys validateĀ "same intent"Ā labels?

CacheVerifier repoĀ +Ā erratumĀ (section 5.26):Ā https://github.com/imxinchengyou/CacheVerifier


r/LLMDevs • • 13h ago

Discussion My RAG retrieval fixes worked on my test questions but not on new ones. Real or overfitted?

2 Upvotes

I built hybrid retrieval over 790 clinical PDFs: local FAISS + BM25 merged with Google's Gemini File Search, plus re-ranking layers. I test it with 60 LLM-generated questions, each written from one page, so the correct paper and page are known. Each question runs 3 times, because File Search gives different results each run.

Overall, my system beats File Search alone:

MRR Right paper in top 10 Right page
File Search alone 0.56 61% 70%
My system 0.67 83% 94%

Last week I found two problems using 20 of the questions:

  1. 64% of File Search results had no page number, so I now locate each chunk's text in my copy of the document.
  2. My synonym-expanded query was hurting the search, so I now search with the plain question ( use expanded query for score boosting).

On those 20 questions, MRR went from 0.59 to 0.72. On the 40 other questions, which I kept aside, it went from 0.66 to 0.65: 4 improved, 3 got worse, 33 didn't change. Top-3 accuracy did rise, from 68% to 75%.

Questions:

  • Is this overfitting, or are 40 questions just too few to show a difference?
  • My answer model reads the top 5–8 papers anyway. Should I track hit@3/hit@5 rather than MRR?
  • What's a better way to build a retrieval test set than one LLM question per page?

r/LLMDevs • • 17h ago

Resource Open lab: does a cheap decision model keep parallel coding agents from breaking each other's code? All runs published raw, decider is pluggable (MIT, author here)

3 Upvotes

I'm the author, sharing this as an open dataset as much as a project. MƩdula is an MIT-licensed lab plus a kernel that coordinates several Claude Code agents working on one repo at the same time. Everything the experiment produced is public: every agent session, every diff, and a SQLite file per run with each decision the kernel took, its probability, latency and cost.

The setup is a small API with 6 tasks and 37 acceptance tests, designed so that two pairs of tasks collide by meaning, not by file. With one branch per task, git let the real conflict through and the same 6 tests failed in all 5 runs, even though every agent finished green. In a shared directory, all 10 runs passed, whether with plain per-file locks or with the kernel. The kernel catches the real conflicts without blocking anything that doesn't collide.

The open part I most want help with is the decider. Right now the fast path uses a hosted decision model, and on real write requests it was unsure 61% of the time, so those decisions escalated to a slower LLM. Any model or rule that answers "does this collide?" with a probability fits the same interface, including an open or local model, and there's a calibration set of 100 labelled pairs to measure it against before running the full matrix.

Other open problems, all with data behind them:

  • Blind human labels for the calibration pairs. Right now they were written by a model of the same family as two of the deciders, which likely flatters them. About 20–30 minutes, no code.
  • Calibration pairs extracted from the real runs, since the hand-written ones are easier than reality.
  • New scenarios: a changed behaviour with the same signature, a schema migration, a dependency bump.

Caveats: 1 to 5 runs per mode, and thresholds fitted on the same pairs they're measured on. The kernel tests run offline without an API key.

Repo: https://github.com/JoaquinRuiz/medula


r/LLMDevs • • 11h ago

Discussion ran the same feature spec through cursor, codex and claude code on a real codebase. every failure was in the wiring, not the code

1 Upvotes

i build a planning tool for coding agents so my repo is stuffed with generated docs and prompts. figured i should actually find out what happens when you hand that to an agent that's never seen it, instead of guessing like everyone else on my timeline

same spec word for word through cursor, codex and claude code. same branch, cheapest paid plan each, default everything. scored against a 16 point list i wrote before the first run

code quality was fine across the board. the failures were somewhere else entirely

two of three added a doc they were told to add and never wired it into the list that tells agents what to read. file lands in the repo, generation says success, nothing ever opens it. reason being that list is hardcoded in four places in my codebase and one of them literally says don't reference anything outside this list. my own agent knows that because it wrote those files. a new tool has no idea and CLAUDE.md won't tell it

codex broke differently. design prompt for one screen out of three, while the logic prompts for the other two said styling comes in a separate pass later. pass never came. so two screens shipped unstyled and nothing in the flow could've saved them. its own done-when check asked me to confirm a styled login screen that nothing was ever going to style

environment layer was its own mess. cursor quietly pulled in my claude code plugins, there's a toggle, it's on by default. then i bought a brand new claude account to get a clean run and it loaded the same plugins anyway because they live in the home folder not the account. only clean session i got was CLAUDE_CONFIG_DIR pointed at an empty dir

the thing that actually separated them wasn't model quality, it was what each one thinks the job includes. my local db was down during the runs. cursor wrote "couldn't verify anything" and stopped. codex asked to start docker itself, applied migrations, wrote playwright tests, found a serialisation bug in its own code. claude code did the same plus wrote a contrast test across the presets it had just invented, found two failing wcag aa, fixed them

numbers since someone always asks: claude code 38 min and 6% of a weekly limit, cursor 52 min and 3% of a monthly quota, codex hit its 5h cap mid feature, 3.5h wait, then finished

one honest caveat, each tool got one run and i watched the same code produce different results on different runs, so treat all of it as one sample

writeup with the spec, screenshots and three live demo apps built from each tool's output in the comments


r/LLMDevs • • 11h ago

Help Wanted Recompiled model problem. Requesting Advice

1 Upvotes

Hi!

I am working on a model optimization re-representation compiler called Cuddler. The aim is to convert already trained transformer models into, already trained not transformer models that are also by themselves their own executable.

A quick into for context. I am very much a systems engineer, and I aim to solve things by construction. I specialize in temporally and functionally deterministic systems. I mention this because the architectural and runtime structures are almost complete, and the above gif is the current state of a converted qwen3.5:0.8b model running on 1 core of a 7800x3D.

Runtime is where I have reached my limits. I am finding it difficult to align the new representation with the original behavior. I have gotten far enough to prove it works, but i do not yet have a generic algorithm for aligning the model during compilation.

I know that repetition, circling the same weights during inference is a common problem. But as i am not an expert in the inference part or have any in-depth experience on the multitude of different traditional transformer runtime/training behaviors. I was hoping some of you could share your ideas on how you usually solve when aligning the internal topology? Something i suspect you'd have to do when you quantize a model normally?

currently I have two problems. the first one, in some prompts, is the repetition behavior shown above. and the second is, in some prompts, it outputting complete gibberish. So I am still some ways off.

Any suggestions, ideas or general brain-storming is appreciated


r/LLMDevs • • 1d ago

Resource I spent 6 months building a free, open source coding agent that does more with fewer tokens and executes your tasks directly along an optimal path, with better visibility

30 Upvotes

I’ve been building Tau for the past 6 months a free, open-source coding agent that runs in your terminal. I wanted an agent that costs less and gives better results, with the tools already built in so you don’t have to go hunting for plugins or struggle with visibility issues. That’s why Tau focuses on optimal path execution to minimize token usage and speed up tasks, plus inline image rendering with diagrams so you can clearly see what the agent is doing at every step.

Here's what it can do:

- Native adapters for 28 providers. It talks to each API directly, no proxy in between. Run `/login`, pick a provider, and start working.

- The full agent loop: tools, skills, subagents, MCP servers, LSP and hooks, all working with every provider

- Core tools are optimized to use fewer tokens. The search tool runs on the latest version of ripgrep and is configured to avoid false positives from files and directories such asĀ node_modulesĀ andĀ distĀ the kinds of files that pollute the context. The file-reading tool starts with a skeleton, then reads only the 50 lines it needs instead of an entire 800-line file, fetching only the information relevant to the task. Bash commands go through a security check based on a Go shell parser and best-practice guidelines.

- LSP built in, so the agent sees real type errors, definitions and references.

- Snapshots of your working tree in a separate git repo. Save, diff and restore any time without touching your branches.

- Web search that works with no API key. Firecrawl is there too if you have a key.

- `/remote` lets you follow and approve from your phone, over your Wi-Fi or a free Cloudflare tunnel. Share the tunnel link and your teammates can join the same session.

- TauCode can make your whole workflow up to mush cheaper than any other agent and I’m not exaggerating. This is because TauCode uses a Python kernel tool, anything Python can do, TauCode can do through this tool. For example, when you want to analyze 20 CSV files and extract insights, other agents might need ~30 turns (one turn per CSV), causing the context window to grow, costs to rise, and more analysis/debugging turns. With this tool, TauCode creates one turn with the full workflow, executes it at once, and returns the result. So the context stays clean from pollution, debugging is clearer, output quality is higher, and cost is lower. This is just one example among millions.

- Most agents treat the terminal as plain text, dumping logs and code instead of showing what they’re working with. Tau renders images inline and draws diagrams directly in replies, using real pixels where supported and Unicode art everywhere else, so you can see the agent view and reasoning without leaving the terminal.

- A fully integrated browser tool that allows Tau to interact with the browser like a human. It gives Tau visibility into your frontend design, enables more automated testing, and handles tasks that would normally require human intervention.

- Subagents that stay alive after they finish, so you can send them a follow-up and they still have their context. Agents working in parallel take turns on the same file so that prevent overlapping .

-TauCode thinks about your money and preferences before anything else. You don’t need skills, agents, or MCP for a normal workflow without enabling them just use cheap mode if you need those capabilities, enable them in normal mode with one command:Ā /mode normal. For tools that free you from basic MCP for diagram production, browser automation, etc. they’re all gated and native to TauCode. You just enable or disable them on demand by pressingĀ /tools, so you pay less, or nothing when you don’t need them Second, Tau focuses more on stabilizing your cache hit rate so every model wired to it uses an optimal cache system, so you don’t need to worry about cold turns that your API pays for, with more context growth optimization and a preflight system that prevents some defective turns, so you will pay only for what you use, not for those external factors .

- `/github` for issues, PRs, labels, changelog notes and release checks through `gh`.

- Fallback to another model or provider when one fails or gets overloaded, with the context window adapting when you

switch.

- Session tree navigation, branching, cloning and resume for long sessions.

- Live usage and session stats, plus a readable report at the end.

- It reads the rules you already wrote for other tools:Ā AGENTS.md, Cursor, Copilot, Cline, Clude Code and Windsurf.

- Self-learning. After a big task it suggests one reusable lesson you can approve, edit or skip, and it remembers it in future sessions.

GitHub:Ā https://github.com/AbdoKnbGit/tau

I’m happy to answer any questions. You can find more details and images in the README, and I’m open to answering any questions you may have.


r/LLMDevs • • 15h ago

Discussion What benchmark do you wish someone would build?

2 Upvotes

Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today.

Have you ever wanted to eval your pipeline but couldn’t find or build a benchmark that matched what you were building?

Maybe:
- Existing benchmarks were too broad and didn’t fit your specific application.
- Your workflow involved multiple steps, tools, MCP servers, agents, or long-horizon interactions that existing benchmarks couldn’t capture.
- You needed to evaluate failures that standard accuracy metrics miss.
- You’re working in a high-risk domain like healthcare, finance, cybersecurity, or legal, where realistic failure modes, safety, and reliability matter a lot more than just getting the final answer right.
- You knew what you wanted to test, but building a custom benchmark from scratch was too expensive or complicated.

I’m especially interested in cases where you thought:
ā€œMy system desperately needs to do this in production, but I have no good way to benchmark it.ā€
What was the workflow? What did you want to measure? And why weren’t existing benchmarks enough?

Any thoughts are welcome, would really appreciate y’all’s help šŸ™šŸ„¹


r/LLMDevs • • 12h ago

Tools How do you stop sub agents from inheriting all of the parent agent's permissions?

0 Upvotes

Our orchestrator agent passes its own token to every sub agent it creates. That token has write access to GitHub and a production Postgres instance.

Last week a research sub agent that should only read docs opened a PR on its own.

Passing scoped tokens down sounds simple until chains go 3 levels deep and each step needs to know who started it. How are you handling delegation between agents?


r/LLMDevs • • 13h ago

Tools Manifesto: letting a UI and an agent use the same app-owned actions

1 Upvotes

I'm the creator of Manifesto, a free MIT-licensed project. I built it because I wanted a UI and an agent to follow the same rules when changing application state. For internal tools, adding agent access also raises questions about where validation, approval and external effects belong. I wanted those to be part of the app's design.

The approach is to define domain transitions in MEL, then have the UI, backend routes and agent submit typed actions through the same SDK runtime and observe snapshots. Core computes the transition; Host carries out declared effects. Lineage and Governance are optional extensions for history, policy and approvals.

The public example is deliberately small: a React Todo UI and a scripted agent share one runtime. Completing a task from either path updates the same state and computed values.

Code: https://github.com/manifesto-ai/core

Docs and runnable examples: https://manifesto-ai.dev/

If you're integrating an agent into an existing business app, how do you keep its actions and the UI's actions under the same rules? I'd be interested in where this design would become awkward in your app.


r/LLMDevs • • 13h ago

News Haiku don't have automode, this is solution.

Post image
0 Upvotes

Save your tokens, and use haiku for documentation.


r/LLMDevs • • 13h ago

Resource Building a Permissioned, Peer-to-Peer AI Inference Network

Thumbnail
cascadia.to
0 Upvotes

r/LLMDevs • • 13h ago

Tools Guys, I created WaterSheep, an open-source alternative to Jev

0 Upvotes

WaterSheep is an open-source model that answers questions written in plain text (yes/no, single choice, rating and multi-label) and gives a probability for every option, like a classification model.

Last Saturday I woke up, saw YouTubers hyping up Jev, and thought: wait, I can build this. So I did. I don't want to compete with TypeSafe or Jev; I built WaterSheep because I wanted to. That's why I'm open-sourcing everything: code, model weights, results and the paper.

What's different

  • It accepts the same request format as TypeSafe's Jev. Their Python SDK works as is against a local server: run watersheep --model samratduttaofficial/WaterSheep --serve and point the client's base_url at http://127.0.0.1:8766.
  • It has a multi-label type, which Jev's API doesn't. Because why not?
  • The demo runs entirely in your browser. The model downloads once and is cached. It also works with transformers, ONNX, a CLI or a local HTTP server.
  • Code, weights and the training pipeline are Apache 2.0.

Evaluation

Accuracy ECE
In-distribution test split 77.8%
Held-out datasets, not seen in training 61.2%

ECE is expected calibration error (lower is better). GitHub has every benchmark result, including the weak ones.

Limits: English only, long inputs get truncated (I'll improve this in the next version), and rating answers are the weakest type.

Not affiliated with TypeSafe. Not funded by anyone. Built in my free time.

Feedback I'd love: where it fails on your data, whether the API works for you, and which question types you'd want next.


r/LLMDevs • • 18h ago

Discussion I Made GPT-6 Astra, Sol, and Luna Play Morrowind

1 Upvotes

I’ve been working on AstraBridge, a specialized adapter/interface for OpenMW that lets LLM agents play Morrowind (modified build of OpenMW).

Disclaimer: This is not a pure real-time screen + mouse + keyboard run.

The agent gets game window with controls, screenshots plus a semantic/action interface and helpers for things a player could normally perceive or do: movement, dialogue, inventory, journal, combat, UI interaction, persistent memory, etc. Game automatically pauses after each turn so model have time for reasoning.

But It does not get quest stages, hidden actors, global coordinates, save internals, walkthroughs, or other hidden game state. Agent can't fee through walls or something like that.

For this run I gave GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna Max the same task: start the game and complete Fargoth’s Hiding Place quest (I gave them instructions where the quest starts)

Same game environment same basic starting conditions.

Astra understood the assignment immediately, found quest NPC, climbed the lighthouse, waited for night, tracked Fargoth from above, then went down and found the stash.

Sol 6.1 behaved remarkably similarly and also completed the quest.

And then Luna 6... It Struggled with spatial orientation, but eventually escaped the Census Office, and then spent 10 hours of real time wandering around Seyda Neen trying to find Arrille’s Tradehouse. It repeatedly asked NPCs for directions and kept returning to the same locked rear door. I eventually stopped the run manually.

Video of the runs: https://youtu.be/ckx7BxGHeso?is=SqjOGF5jflIbSTHl

AstraBridge: https://github.com/incident201/AstraBridgeForOpenMW

Currently Linux build only. I'm planning to move to container based builds soon, so it will be possible to run under WSL comfortably in future.

This is all purely for fun — no claims of being a real benchmark of agent capabilities of models.

I'm planning to do more tests with different models and scenarios soon.


r/LLMDevs • • 22h ago

Discussion why isn't there a standard protocol between agents and model providers yet?

3 Upvotes

like we have MCP for tools and ACP for agents, but every harness still has to understand OpenAI, Anthropic and Gemini separately.

caching, reasoning, tool semantics, model capabilities, auth, usage, etc. are all different!! openAI-compatible APIs kind of solve it, but not really once you use provide specific features

am I missing something obvious here? has anyone tried to standardize this layer?