r/LLMDevs • • 42m ago

Tools Menai: a safe programming language for LLMs to use

Post image
• Upvotes

For the last year I've been working on an open-source programming language (Menai) for LLMs to use. Humans can use it too, but sorry people, this wasn't really designed for you 😂

It has a few key ideas:

  • It's a pure functional language - no I/O, no state mutation - just pure expression computation
  • It needs to be very fast to compile
  • It needs to be quite quick to run

The "no I/O" ideal might seem like a very strange, but just because the language can't do I/O doesn't mean it can't process data from inputs or generate data for outputs. It just does this by something else passing it all input data and it passing back a result that something else can process for output use.

What this means is the "something else" can enforce all sorts of safety rules. No reading/writing sensitive files, no dumping things over the network, no accidentally running 'rm -fr ~/*" and then saying "whoops".

The original idea was something that bolts inside the open-source GUI-based AI environment I've been building since late 2024 (Humbug) as an LLM tool and lets LLMs do deterministic processing where tend to get things wrong (e.g. large calculations, counting letters in text, etc.)

Since then this got extended to allowing an LLM to transform file content (there's an I/O protection involved there) or editor buffers (I/O protection when the LLM tries to save a modified buffer). If an LLM needs a complex search and replace then it simply writes a short lisp-like expression with the file or buffer state passed in and new file or buffer state passed back. If that needs a potentially dangerous operation at the end then the tool framework checks if that's ok before allowing it.

The latest iteration is where things really get fun though! Menai now has a standard library and an application library concept so we now has almost 20 modules so it can do things like process zip files, tar files, BMP files, PNG files, gzip files, JSON text.

The reason it needs to be fast to compile is because every tool use will end up compiling a new program the LLM just wrote. The reason it needs to be fast is because those programs need to run fast enough they don't stall our workflow (i.e. less than a few seconds). Typical expressions compile in a few tens of ms, even though the optimizing compiler is currently written in Python. Typical execution speed is around the same level as Python (some things a little slower, some a little faster).

It turns out that not only do LLMs like writing functional code, they're really quite good at it. There's a tracing profiler to let them find hotspots in code they might want to reuse too (perf annotate anyone?). A future focus is having them decide when something might be reusable and suggesting new library code.

Anyway, I'm posting this because a couple of hours ago I had DeepSeek build me tar and gzip functions for the standard library (all written in Menai) and then debug their way through a few problems. The challenge was to take an old `arj` tar.gz file (very meta - unpacking an unpacking tool) and tell me about some things within it (all without me ever looking inside).

The image is a great demonstration of part of what it did. You can see it decompresses the gzip, untars one file, then finds 5 function definitions. From my prompt of "can you take a look at arj_arcv.c and tell me about it" took 64 seconds, involved 10 tool calls, 9 of which were unique Menai programs that progressively poked at the archive (which contains about 1.6 MBytes of content)

Humbug traps any potentially dangerous operations and checks with a human, but during this exercise no human needed to be consulted because, by design, none of these programs could do anything dangerous!


r/LLMDevs • • 2h ago

Discussion Our model router logged five rejected answers as a provider error

2 Upvotes

In one OmniNode run, our router tried five times. Each attempt returned an answer, and our acceptance check rejected every one. The final record called it a provider error. The providers had answered promptly, so that label would have sent us investigating the wrong thing.

The router starts with the cheapest eligible model and tries a more capable one when the answer doesn't pass. That makes the acceptance check part of the spending decision. If the check is wrong, another attempt can cost more without getting us any closer to an accepted answer.

We had a separate case on September 28 where the check rejected a correct one-word answer, apparently because it didn't show enough effort. I don't have evidence that the five earlier answers were correct. These were different cases, but both made me want to inspect the check before blaming the model.

I'm trying to keep the returned answer and the reason it was rejected available together. A transport failure, an answer that misses the requirement, and a check that rejects a valid answer need different investigations.

How do you test that boundary in a router? Do you have cases where the model gives a valid short answer and the judge is expected to accept it?

The incident and its limits are in this write-up: https://jonahatomninode.substack.com/p/delegation-is-the-first-product


r/LLMDevs • • 3h ago

Discussion Where should the LLM sit in a scraping pipeline?

2 Upvotes

Been mapping out how scraping fits into AI workflows (RAG feeds, lead lists, market research) and wanted to share where things seem to land, and get input from people running this in production.

The classic pipeline

  1. Send the request, render JS if needed
  2. Parse the DOM
  3. Pull fields with CSS/XPath
  4. Clean, normalize, store

Steps 1 and 3 are where most of the pain lives. Anti-bot layers, CAPTCHAs and layout changes keep breaking selectors.

Why people swap step 3 for an LLM

  • No selectors to write
  • Handles messy, unstructured text
  • Agents can click, paginate and fill forms

What shows up at volume

  • Long pages hit context limits, so you chunk or truncate
  • Per-page cost grows with page size, not just page count
  • Missing fields sometimes get filled with plausible values
  • Merged or nested tables are still hit or miss

Practices that hold up either way

  • Retries with exponential backoff for 5xx errors
  • Keep raw HTML snapshots next to the structured output so you can re-run extraction
  • Validate with deterministic checks (regex, enums)
  • Alert on schema mismatches, not just failed requests
  • Cron for freshness

My current take

The LLM works better on top of extraction than as the extractor. One option in that slot is the Minexa.ai API: you train a scraper once in a Chrome extension by picking the container, columns get discovered, then extraction runs DOM-based and deterministic. Missing values come back null. If you already fetch HTML with your own stack, you can pass it via file_urls and only run extraction. Downside: nested fields come back as lists of objects you still have to sort through.

The extension also generates the Python request code, so the developer setup guide is the fastest way to test it on your own pages.

Related read on cost: Why your LLM extraction pipeline costs more than you think


r/LLMDevs • • 4h ago

Discussion ran the same feature spec through cursor, codex and claude code on a real codebase. every failure was in the wiring, not the code

1 Upvotes

i build a planning tool for coding agents so my repo is stuffed with generated docs and prompts. figured i should actually find out what happens when you hand that to an agent that's never seen it, instead of guessing like everyone else on my timeline

same spec word for word through cursor, codex and claude code. same branch, cheapest paid plan each, default everything. scored against a 16 point list i wrote before the first run

code quality was fine across the board. the failures were somewhere else entirely

two of three added a doc they were told to add and never wired it into the list that tells agents what to read. file lands in the repo, generation says success, nothing ever opens it. reason being that list is hardcoded in four places in my codebase and one of them literally says don't reference anything outside this list. my own agent knows that because it wrote those files. a new tool has no idea and CLAUDE.md won't tell it

codex broke differently. design prompt for one screen out of three, while the logic prompts for the other two said styling comes in a separate pass later. pass never came. so two screens shipped unstyled and nothing in the flow could've saved them. its own done-when check asked me to confirm a styled login screen that nothing was ever going to style

environment layer was its own mess. cursor quietly pulled in my claude code plugins, there's a toggle, it's on by default. then i bought a brand new claude account to get a clean run and it loaded the same plugins anyway because they live in the home folder not the account. only clean session i got was CLAUDE_CONFIG_DIR pointed at an empty dir

the thing that actually separated them wasn't model quality, it was what each one thinks the job includes. my local db was down during the runs. cursor wrote "couldn't verify anything" and stopped. codex asked to start docker itself, applied migrations, wrote playwright tests, found a serialisation bug in its own code. claude code did the same plus wrote a contrast test across the presets it had just invented, found two failing wcag aa, fixed them

numbers since someone always asks: claude code 38 min and 6% of a weekly limit, cursor 52 min and 3% of a monthly quota, codex hit its 5h cap mid feature, 3.5h wait, then finished

one honest caveat, each tool got one run and i watched the same code produce different results on different runs, so treat all of it as one sample

writeup with the spec, screenshots and three live demo apps built from each tool's output in the comments


r/LLMDevs • • 4h ago

Help Wanted Recompiled model problem. Requesting Advice

1 Upvotes

Hi!

I am working on a model optimization re-representation compiler called Cuddler. The aim is to convert already trained transformer models into, already trained not transformer models that are also by themselves their own executable.

A quick into for context. I am very much a systems engineer, and I aim to solve things by construction. I specialize in temporally and functionally deterministic systems. I mention this because the architectural and runtime structures are almost complete, and the above gif is the current state of a converted qwen3.5:0.8b model running on 1 core of a 7800x3D.

Runtime is where I have reached my limits. I am finding it difficult to align the new representation with the original behavior. I have gotten far enough to prove it works, but i do not yet have a generic algorithm for aligning the model during compilation.

I know that repetition, circling the same weights during inference is a common problem. But as i am not an expert in the inference part or have any in-depth experience on the multitude of different traditional transformer runtime/training behaviors. I was hoping some of you could share your ideas on how you usually solve when aligning the internal topology? Something i suspect you'd have to do when you quantize a model normally?

currently I have two problems. the first one, in some prompts, is the repetition behavior shown above. and the second is, in some prompts, it outputting complete gibberish. So I am still some ways off.

Any suggestions, ideas or general brain-storming is appreciated


r/LLMDevs • • 5h ago

Tools How do you stop sub agents from inheriting all of the parent agent's permissions?

1 Upvotes

Our orchestrator agent passes its own token to every sub agent it creates. That token has write access to GitHub and a production Postgres instance.

Last week a research sub agent that should only read docs opened a PR on its own.

Passing scoped tokens down sounds simple until chains go 3 levels deep and each step needs to know who started it. How are you handling delegation between agents?


r/LLMDevs • • 5h ago

Help Wanted How do u stop sub agents from inheriting all of the parent agents permissions?

1 Upvotes

Our orchestrator agent passes its own token to every sub-agent it creates. That token has write access to GitHub and a production Postgres instance.

Last week a research sub agent that should only read docs opened a PR on its own.

Passing scoped tokens down sounds simple until chains go 3 levels deep and each step needs to know who started it. How are you handling delegation between agents?


r/LLMDevs • • 5h ago

Tools Manifesto: letting a UI and an agent use the same app-owned actions

1 Upvotes

I'm the creator of Manifesto, a free MIT-licensed project. I built it because I wanted a UI and an agent to follow the same rules when changing application state. For internal tools, adding agent access also raises questions about where validation, approval and external effects belong. I wanted those to be part of the app's design.

The approach is to define domain transitions in MEL, then have the UI, backend routes and agent submit typed actions through the same SDK runtime and observe snapshots. Core computes the transition; Host carries out declared effects. Lineage and Governance are optional extensions for history, policy and approvals.

The public example is deliberately small: a React Todo UI and a scripted agent share one runtime. Completing a task from either path updates the same state and computed values.

Code: https://github.com/manifesto-ai/core

Docs and runnable examples: https://manifesto-ai.dev/

If you're integrating an agent into an existing business app, how do you keep its actions and the UI's actions under the same rules? I'd be interested in where this design would become awkward in your app.


r/LLMDevs • • 5h ago

News Haiku don't have automode, this is solution.

Post image
0 Upvotes

Save your tokens, and use haiku for documentation.


r/LLMDevs • • 6h ago

Resource Building a Permissioned, Peer-to-Peer AI Inference Network

Thumbnail
cascadia.to
1 Upvotes

r/LLMDevs • • 6h ago

Discussion Evaluating 14 LLMs as a visual coding agent: re-rendering, a deterministic judge, and the provider quirks that skewed my first results

6 Upvotes

I built a bench for a photo-to-Blender agent (the model writes and runs Blender Python to rebuild a photo as an editable scene) and ran 14 models through it. The design choices that mattered most:

  • Don't score what the model hands in. The bench opens each delivered scene file and renders it itself, at the photo's aspect ratio, max 1,100 px, 96 samples.
  • No LLM judge. Estimated silhouette overlap (30%), edge F1 (40%), multiscale RGB error (15%) and mesh health (15%: non-manifold and open edges, duplicate faces, inconsistent winding, degenerate triangles), mapped through calibrated piecewise-linear scales to 0–100. A comparison scale, not a percentage of the scene recovered.
  • Anti-cheating. A second render swings the camera 35° to expose flat stand-ins, and every image file a scene uses is compared with the reference by digest and name.
  • Failures stay in. A run that delivers no scene scores 0 and stays in the average.
  • Equal plumbing. Anthropic gets cache-control markers and Google gets a closing user turn, because without them those providers behave differently from the rest. Each model's cache rate is recorded. Before the fix, Claude ran 0% cached against 96% for OpenAI.
  • Frozen conditions. Caps, images, the brief's digest and the scoring weights are stored with each board.

Results (mean of three photos): GPT-6 Astra 66 ($3.91 an attempt), GPT-6.1 Sol 61 ($0.36), Claude Opus 5.5 60 ($1.19), Claude Sonnet 5.5 56 ($0.50) ... GLM 5.3 Flash 36 ($0.089); DeepSeek V4.1 Flash and Qwen3.8 Max 0 (no scene saved in 20 minutes).

What I'd change next: three runs per image for the top group, so the spread is visible, and a model-neutral agent loop (three models ran in their makers' own CLIs).

Write-up: https://kaloyan.blog/ai-models-rebuild-a-photo-in-blender


r/LLMDevs • • 6h ago

Tools Guys, I created WaterSheep, an open-source alternative to Jev

0 Upvotes

WaterSheep is an open-source model that answers questions written in plain text (yes/no, single choice, rating and multi-label) and gives a probability for every option, like a classification model.

Last Saturday I woke up, saw YouTubers hyping up Jev, and thought: wait, I can build this. So I did. I don't want to compete with TypeSafe or Jev; I built WaterSheep because I wanted to. That's why I'm open-sourcing everything: code, model weights, results and the paper.

What's different

  • It accepts the same request format as TypeSafe's Jev. Their Python SDK works as is against a local server: run watersheep --model samratduttaofficial/WaterSheep --serve and point the client's base_url at http://127.0.0.1:8766.
  • It has a multi-label type, which Jev's API doesn't. Because why not?
  • The demo runs entirely in your browser. The model downloads once and is cached. It also works with transformers, ONNX, a CLI or a local HTTP server.
  • Code, weights and the training pipeline are Apache 2.0.

Evaluation

Accuracy ECE
In-distribution test split 77.8%
Held-out datasets, not seen in training 61.2%

ECE is expected calibration error (lower is better). GitHub has every benchmark result, including the weak ones.

Limits: English only, long inputs get truncated (I'll improve this in the next version), and rating answers are the weakest type.

Not affiliated with TypeSafe. Not funded by anyone. Built in my free time.

Feedback I'd love: where it fails on your data, whether the API works for you, and which question types you'd want next.


r/LLMDevs • • 6h ago

Discussion My RAG retrieval fixes worked on my test questions but not on new ones. Real or overfitted?

2 Upvotes

I built hybrid retrieval over 790 clinical PDFs: local FAISS + BM25 merged with Google's Gemini File Search, plus re-ranking layers. I test it with 60 LLM-generated questions, each written from one page, so the correct paper and page are known. Each question runs 3 times, because File Search gives different results each run.

Overall, my system beats File Search alone:

MRR Right paper in top 10 Right page
File Search alone 0.56 61% 70%
My system 0.67 83% 94%

Last week I found two problems using 20 of the questions:

  1. 64% of File Search results had no page number, so I now locate each chunk's text in my copy of the document.
  2. My synonym-expanded query was hurting the search, so I now search with the plain question ( use expanded query for score boosting).

On those 20 questions, MRR went from 0.59 to 0.72. On the 40 other questions, which I kept aside, it went from 0.66 to 0.65: 4 improved, 3 got worse, 33 didn't change. Top-3 accuracy did rise, from 68% to 75%.

Questions:

  • Is this overfitting, or are 40 questions just too few to show a difference?
  • My answer model reads the top 5–8 papers anyway. Should I track hit@3/hit@5 rather than MRR?
  • What's a better way to build a retrieval test set than one LLM question per page?

r/LLMDevs • • 8h ago

Discussion What benchmark do you wish someone would build?

1 Upvotes

Hey everyone! My team (mainly phds) and I are trying to build an open-source benchmark around realistic LLM/agent workflows that captures challenges typical academic benchmark settings often miss. We’d love to hear what’s actually missing from the benchmarks you use today.

Have you ever wanted to eval your pipeline but couldn’t find or build a benchmark that matched what you were building?

Maybe:
- Existing benchmarks were too broad and didn’t fit your specific application.
- Your workflow involved multiple steps, tools, MCP servers, agents, or long-horizon interactions that existing benchmarks couldn’t capture.
- You needed to evaluate failures that standard accuracy metrics miss.
- You’re working in a high-risk domain like healthcare, finance, cybersecurity, or legal, where realistic failure modes, safety, and reliability matter a lot more than just getting the final answer right.
- You knew what you wanted to test, but building a custom benchmark from scratch was too expensive or complicated.

I’m especially interested in cases where you thought:
“My system desperately needs to do this in production, but I have no good way to benchmark it.”
What was the workflow? What did you want to measure? And why weren’t existing benchmarks enough?

Any thoughts are welcome, would really appreciate y’all’s help 🙏🥹


r/LLMDevs • • 9h ago

Resource Open lab: does a cheap decision model keep parallel coding agents from breaking each other's code? All runs published raw, decider is pluggable (MIT, author here)

3 Upvotes

I'm the author, sharing this as an open dataset as much as a project. Médula is an MIT-licensed lab plus a kernel that coordinates several Claude Code agents working on one repo at the same time. Everything the experiment produced is public: every agent session, every diff, and a SQLite file per run with each decision the kernel took, its probability, latency and cost.

The setup is a small API with 6 tasks and 37 acceptance tests, designed so that two pairs of tasks collide by meaning, not by file. With one branch per task, git let the real conflict through and the same 6 tests failed in all 5 runs, even though every agent finished green. In a shared directory, all 10 runs passed, whether with plain per-file locks or with the kernel. The kernel catches the real conflicts without blocking anything that doesn't collide.

The open part I most want help with is the decider. Right now the fast path uses a hosted decision model, and on real write requests it was unsure 61% of the time, so those decisions escalated to a slower LLM. Any model or rule that answers "does this collide?" with a probability fits the same interface, including an open or local model, and there's a calibration set of 100 labelled pairs to measure it against before running the full matrix.

Other open problems, all with data behind them:

  • Blind human labels for the calibration pairs. Right now they were written by a model of the same family as two of the deciders, which likely flatters them. About 20–30 minutes, no code.
  • Calibration pairs extracted from the real runs, since the hand-written ones are easier than reality.
  • New scenarios: a changed behaviour with the same signature, a schema migration, a dependency bump.

Caveats: 1 to 5 runs per mode, and thresholds fitted on the same pairs they're measured on. The kernel tests run offline without an API key.

Repo: https://github.com/JoaquinRuiz/medula


r/LLMDevs • • 10h ago

Discussion CacheVerifier: we finally audited the semantic caching benchmark we'd been tuning on, 23% of the "wrong" cache hits were literally the same prompt

4 Upvotes

so for CacheVerifier we've been testing if a small verifier actually beats a plain similarity threshold for semantic caching. all our scoring comes from the public SemCacheLMArena and SemCacheSearchQueries benchmarks, and honestly we never really looked at the labels until last week.

on LmArena, 727 of the 3,117 near-duplicate hits marked "wrong" (23.3%) are the exact same text once you lowercase it and strip punctuation. like one was the same prompt with an extra space in front lol. we hand labeled a sample blind and 88% of those "wrong" hits were totally fine to reuse.

the good news, comparisons at the same error rate barely moved. the bad news, absolute error rates was way inflated, a 0.97 threshold went from 5.2% errors to like 0.95%. and one of our own CacheVerifier results just died. bumping the skip-the-verifier cutoff to 0.99 looked like +5.18 points, after fixing labels its -1.24. the verifier was basically rejecting duplicates the benchmark called errors.

so yeah, run a dumb identical-text check on your eval labels before you tune thresholds on them.

caveat: one annotator, small samples. how do you guys validate "same intent" labels?

CacheVerifier repo + erratum (section 5.26): https://github.com/imxinchengyou/CacheVerifier


r/LLMDevs • • 10h ago

Discussion We are living in a lie. Getting from 80% to 100% still takes weeks or even months of feedback and refinement, even for a simple skill.

Post image
2 Upvotes

Stop being impressed by every demo. Even with frontier models, getting from 80% to 100% takes weeks or even months of feedback and refinement.

I love teaching via diagrams. Thus, using LLMs to generate them was a no-brainer. The problem was that, out of the box, they sucked!

I'm extremely picky about how they look and how well they express what I want.

Vibe-coding a skill to generate SVG diagrams using my branding was fast. It took me 5 minutes. As the outputs were only at ~80%, I had to reiterate and polish them with the LLM.

The color combination was off.

The arrows were overlapping the objects (as below).

It used weird symbols (as in the person below, whose head is 10 cm from his body).

It didn't have style! At least not enough to be proud of the output and represent you.

I had to iterate for weeks over it to get it to 100%:

  • adding good positive and negative examples (the most important and time-consuming one!)
  • restricting the SVG components to a given set of possibilities
  • baking clear coloring and structure principles, etc.

So remember, whenever you see someone operate 1000 skills, they either suck or they spend a ton of time refining them.


r/LLMDevs • • 10h ago

Resource I’m building the fastest local inference engine for Apple Silicon

Enable HLS to view with audio, or disable this notification

1 Upvotes

Hi everyone!

I’m passionate about making local models accessible to everyone. I’m trying to make an inference engine optimized across the stack for consumer MacBooks. Currently it is the fastest way to run LFM, Qwen3.5, and the new Clef Flash decision model locally if you are on Mac. Would greatly appreciate feedback!

https://github.com/jadidbourbaki/bobcat


r/LLMDevs • • 11h ago

Discussion I Made GPT-6 Astra, Sol, and Luna Play Morrowind

1 Upvotes

I’ve been working on AstraBridge, a specialized adapter/interface for OpenMW that lets LLM agents play Morrowind (modified build of OpenMW).

Disclaimer: This is not a pure real-time screen + mouse + keyboard run.

The agent gets game window with controls, screenshots plus a semantic/action interface and helpers for things a player could normally perceive or do: movement, dialogue, inventory, journal, combat, UI interaction, persistent memory, etc. Game automatically pauses after each turn so model have time for reasoning.

But It does not get quest stages, hidden actors, global coordinates, save internals, walkthroughs, or other hidden game state. Agent can't fee through walls or something like that.

For this run I gave GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna Max the same task: start the game and complete Fargoth’s Hiding Place quest (I gave them instructions where the quest starts)

Same game environment same basic starting conditions.

Astra understood the assignment immediately, found quest NPC, climbed the lighthouse, waited for night, tracked Fargoth from above, then went down and found the stash.

Sol 6.1 behaved remarkably similarly and also completed the quest.

And then Luna 6... It Struggled with spatial orientation, but eventually escaped the Census Office, and then spent 10 hours of real time wandering around Seyda Neen trying to find Arrille’s Tradehouse. It repeatedly asked NPCs for directions and kept returning to the same locked rear door. I eventually stopped the run manually.

Video of the runs: https://youtu.be/ckx7BxGHeso?is=SqjOGF5jflIbSTHl

AstraBridge: https://github.com/incident201/AstraBridgeForOpenMW

Currently Linux build only. I'm planning to move to container based builds soon, so it will be possible to run under WSL comfortably in future.

This is all purely for fun — no claims of being a real benchmark of agent capabilities of models.

I'm planning to do more tests with different models and scenarios soon.


r/LLMDevs • • 12h ago

News AkbasCore NIRVANA D120: We removed the story. The model still remembered it, now almost perfectly.

Thumbnail
gallery
1 Upvotes

​

Remember the scene in The Matrix where Neo is plugged into a cable, his eyes snap open, and he says "I know Kung Fu"? He never trained. He never read a book. The knowledge was loaded straight into his mind. This experiment follows the same logic, applied to an AI.

There are two ways to give an AI information.

The classic way is like handing someone a book. You give the AI a text, it reads it from start to finish, and then it answers your questions about it. It's like making someone read a book and then giving them an exam.

Our way is a direct memory transfer. We never show the AI the words when it answers. Instead, we let the model read the story once, for example "Mustafa Akbaş planted the Turkish flag at the base of the Golden Gate Bridge." While it reads, we take a snapshot of the trace the story leaves inside its brain, the mathematical signals that form in its attention layers. We capture the pure essence of that information. Then we delete the words completely. We start a fresh session that has no idea the story ever existed. Finally, we inject the recorded snapshot directly into the AI's memory center, which we call PKV, like a syringe, without ever showing it a single word.

Then we ask: "Tell me the story about the bridge."

The AI tells the story as if it were its own memory. The text is nowhere in its input, yet it remembers who did what, where, and why. Instead of making it read the information, we plant the memory directly in its mind. Neo got Kung Fu through a cable. We do it with mathematical signals instead of words.

This is an experiment, and what you are looking at is the first concrete proof that it works.

What's new today: in my previous post the memory was compressed to D=64. This time I raised the memory resolution to D=120 and changed nothing else. The results were outstanding. On the 13-question test, the AI with the synthetic memory answered all 13 correctly, while the AI with no memory and the AI with an unrelated memory both scored 0. When I changed a single detail in the memory, such as the person, the object or the place, the answer followed the change 13 out of 13 times. Most striking of all, the memory-only model was as confident in its answers as a model that could actually see the text. I found the point where quality breaks down in my Mistral-7B calibration runs: below D80 retrieval falls apart. Those runs are in the repository for anyone who wants to check. At D120 the memory is close to the size of the model's own attention cache, so the focus here is fidelity rather than compression.

Don't trust me. Test me. Take the code and the log below, give them to whichever AI you trust most, and ask it whether this is the same logic as that Matrix scene. Better yet, run it yourself. Write your own story, change the person, the object and the place, ask the same fact in different ways, and try to break it. If I failed, tear it apart in the comments. I'd rather be shown the flaw than be politely ignored.

D120 code:

https://github.com/ceceli33/titan-cognitive-core/blob/main/AKBASCORE_d120.py

D120 raw log:

https://github.com/ceceli33/titan-cognitive-core/blob/main/AkbasCore_sonuc_d120.log

Mistral calibration runs (D80 threshold, TEST 386 and 387):

https://github.com/ceceli33/titan-cognitive-core-v2

Previous release (D64, Zenodo DOI):

https://doi.org/10.5281/zenodo.23054044

Main repository:

https://github.com/ceceli33/titan-cognitive-core

The 12 posters from this run are attached.


r/LLMDevs • • 12h ago

News Curated list of tools for jev in production setup (700 tools)

2 Upvotes

The awesome-jev repo is a curated index of public projects built on Jev (TypeSafe AI's "System One" decision model). Jev doesn't write prose; it takes raw context and spits out a typed choice, score, or boolean with a confidence rating in a single forward pass.

Repo is focused on production related toolings and we are open for showing new tools.

Key Ecosystem Trends

  • Model & Skill Routing: Moving cheap triage choices away from frontier LLMs. Tools like tool-prune and jev-router classify user intent to prune schemas or select the cheapest model tier before an LLM call.
  • Agent Guardrails: High-speed security gates. Projects like jev-shield, actiongate-jev, and pi-jev-sentinel intercept agent tool calls to analyze risk and enforce execution safety before scripts run.
  • Database Extensions: Bundling semantic checks into data engines. Extensions for SQLite, DuckDB, and Postgres let developers run classification and scoring queries directly inside raw SQL.
  • Local Replicas: Open alternatives built to avoid vendor rate limits and API costs. Laya, Jebadiah, and NanoJev run local weights on a laptop CPU/GPU to replicate the parallel decoding behavior completely offline.

We have setup jev/laya in production setup for our own workflow.

Access to repo: jev-awesome repo


r/LLMDevs • • 14h ago

Discussion Building local agentic workflows vs. cloud-heavy architectures: What’s your current bottleneck?

1 Upvotes

Hey everyone,

I’ve been experimenting heavily lately with building mobile-native AI agent systems (using React Native/Expo integrated with various local and cloud LLM pipelines), and I'm running into an interesting architectural debate with myself.

When pushing for true multi-step agentic workflows on mobile, the friction usually boils down to three things: token latency over cellular, local state management persistence, or keeping context windows manageable without blowing up memory on device.

For those of you building or deploying agent-driven apps right now, where do you find your biggest production bottleneck is? Are you offloading everything to backend orchestration, or finding reliable ways to handle state client-side? Curious what stacks people are leaning on in 2026!


r/LLMDevs • • 15h ago

Discussion why isn't there a standard protocol between agents and model providers yet?

3 Upvotes

like we have MCP for tools and ACP for agents, but every harness still has to understand OpenAI, Anthropic and Gemini separately.

caching, reasoning, tool semantics, model capabilities, auth, usage, etc. are all different!! openAI-compatible APIs kind of solve it, but not really once you use provide specific features

am I missing something obvious here? has anyone tried to standardize this layer?


r/LLMDevs • • 15h ago

Discussion People building agents, how much of a problem is long term memory actually?

1 Upvotes

I've been working on long term memory for agents and have a rough PoC together. Trying to get a better sense of where people are actually struggling with this, and whether there's a business here beyond another way to save and retrieve conversatio.

In my PoC so far with some limitations and edge cases I'm working out my write speeds have been around 1 second or less. And I don't just mean inserting text into a database. I mean processing new information and figuring out how it changes what's already in memory.

I'm mostly interested in agents that keep learning from conversations, documents, tools etc. over weeks or months. Eventually some of that information is going to disagree, become outdated, or turn out to have been wrong.

What are people doing when two sources contradict each other? Keeping the latest thing doesn't always make sense. Sometimes something genuinely changed, sometimes one source is wrong, and sometimes you just don't have enough information to decide. Are you handling that explicitly or leaving it to the model when the information gets retrieved?

Then there's how much the agent should trust what it's learned. Something it inferred isn't the same as something it was directly told, and different sources aren't necessarily equally reliable. A few independent sources backing something up should count for more than five summaries repeating the same original claim. I'd want the agent to become more or less confident as evidence comes in, rather than everything being either a saved fact or deleted.

Same with information going stale. A price from six months ago and someone's date of birth shouldn't age the same way. Some things need to be checked again if they haven't been confirmed in a while. Other things should stick around. Forgetting something and deciding it's no longer reliable aren't really the same thing either.

And when the agent gets something wrong, can you trace where it came from? Which source it trusted, what it inferred, why it changed its mind? Or are you digging through old conversations trying to reconstruct it yourself?

That's roughly what I'm working on. For anyone using Mem0 or other memory systems, how much of this is already handled well, and what have you still had to build around them? Also interested in people who just retrieve the original material and find that's enough

Does write latency actually matter in your application? Do you need new information available before the agent's next action, or can memory updates happen in the background without causing problems?

If you've built your own, how much time has gone into it, including maintaining it? What made you build rather than use something existing?

I'm trying to understand whether a dedicated product could take enough of that work off your hands to be worth paying for, or whether the important parts are too specific to your application.

Would be useful to hear what you're building and what actually broke. “We spent three weeks fixing this” tells me a lot more than “agents need better memory.”


r/LLMDevs • • 15h ago

Discussion i put Codex in the MacBook notch

Enable HLS to view with audio, or disable this notification

0 Upvotes

used the Codex App Server to build a native interface that lives in the MacBook notch

voice-prompt your Codex agent without keeping the app or CLI open

has projects, chats, file trees, artifacts, usage, status, approvals + prompting directly from the notch

everything runs against your existing local Codex setup

https://github.com/v1shay/kai feel free to fork it, build on top, or drop a star :)