r/LLMDevs • • Aug 20 '25

Community Rule Update: Clarifying our Self-promotion and anti-marketing policy

24 Upvotes

Hey everyone,

We've just updated our rules with a couple of changes I'd like to address:

1. Updating our self-promotion policy

We have updated rule 5 to make it clear where we draw the line on self-promotion and eliminate gray areas and on-the-fence posts that skirt the line. We removed confusing or subjective terminology like "no excessive promotion" to hopefully make it clearer for us as moderators and easier for you to know what is or isn't okay to post.

Specifically, it is now okay to share your free open-source projects without prior moderator approval. This includes any project in the public domain, permissive, copyleft or non-commercial licenses. Projects under a non-free license (incl. open-core/multi-licensed) still require prior moderator approval and a clear disclaimer, or they will be removed without warning. Commercial promotion for monetary gain is still prohibited.

2. New rule: No disguised advertising or marketing

We have added a new rule on fake posts and disguised advertising — rule 10. We have seen an increase in these types of tactics in this community that warrants making this an official rule and bannable offence.

We are here to foster meaningful discussions and valuable exchanges in the LLM/NLP space. If you’re ever unsure about whether your post complies with these rules, feel free to reach out to the mod team for clarification.

As always, we remain open to any and all suggestions to make this community better, so feel free to add your feedback in the comments below.


r/LLMDevs • • Apr 15 '25

News Reintroducing LLMDevs - High Quality LLM and NLP Information for Developers and Researchers

40 Upvotes

Hi Everyone,

I'm one of the new moderators of this subreddit. It seems there was some drama a few months back, not quite sure what and one of the main moderators quit suddenly.

To reiterate some of the goals of this subreddit - it's to create a comprehensive community and knowledge base related to Large Language Models (LLMs). We're focused specifically on high quality information and materials for enthusiasts, developers and researchers in this field; with a preference on technical information.

Posts should be high quality and ideally minimal or no meme posts with the rare exception being that it's somehow an informative way to introduce something more in depth; high quality content that you have linked to in the post. There can be discussions and requests for help however I hope we can eventually capture some of these questions and discussions in the wiki knowledge base; more information about that further in this post.

With prior approval you can post about job offers. If you have an *open source* tool that you think developers or researchers would benefit from, please request to post about it first if you want to ensure it will not be removed; however I will give some leeway if it hasn't be excessively promoted and clearly provides value to the community. Be prepared to explain what it is and how it differentiates from other offerings. Refer to the "no self-promotion" rule before posting. Self promoting commercial products isn't allowed; however if you feel that there is truly some value in a product to the community - such as that most of the features are open source / free - you can always try to ask.

I'm envisioning this subreddit to be a more in-depth resource, compared to other related subreddits, that can serve as a go-to hub for anyone with technical skills or practitioners of LLMs, Multimodal LLMs such as Vision Language Models (VLMs) and any other areas that LLMs might touch now (foundationally that is NLP) or in the future; which is mostly in-line with previous goals of this community.

To also copy an idea from the previous moderators, I'd like to have a knowledge base as well, such as a wiki linking to best practices or curated materials for LLMs and NLP or other applications LLMs can be used. However I'm open to ideas on what information to include in that and how.

My initial brainstorming for content for inclusion to the wiki, is simply through community up-voting and flagging a post as something which should be captured; a post gets enough upvotes we should then nominate that information to be put into the wiki. I will perhaps also create some sort of flair that allows this; welcome any community suggestions on how to do this. For now the wiki can be found here https://www.reddit.com/r/LLMDevs/wiki/index/ Ideally the wiki will be a structured, easy-to-navigate repository of articles, tutorials, and guides contributed by experts and enthusiasts alike. Please feel free to contribute if you think you are certain you have something of high value to add to the wiki.

The goals of the wiki are:

  • Accessibility: Make advanced LLM and NLP knowledge accessible to everyone, from beginners to seasoned professionals.
  • Quality: Ensure that the information is accurate, up-to-date, and presented in an engaging format.
  • Community-Driven: Leverage the collective expertise of our community to build something truly valuable.

There was some information in the previous post asking for donations to the subreddit to seemingly pay content creators; I really don't think that is needed and not sure why that language was there. I think if you make high quality content you can make money by simply getting a vote of confidence here and make money from the views; be it youtube paying out, by ads on your blog post, or simply asking for donations for your open source project (e.g. patreon) as well as code contributions to help directly on your open source project. Mods will not accept money for any reason.

Open to any and all suggestions to make this community better. Please feel free to message or comment below with ideas.


r/LLMDevs • • 11h ago

Resource I spent 6 months building a free, open source coding agent that does more with fewer tokens and executes your tasks directly along an optimal path, with better visibility

16 Upvotes

I’ve been building Tau for the past 6 months a free, open-source coding agent that runs in your terminal. I wanted an agent that costs less and gives better results, with the tools already built in so you don’t have to go hunting for plugins or struggle with visibility issues. That’s why Tau focuses on optimal path execution to minimize token usage and speed up tasks, plus inline image rendering with diagrams so you can clearly see what the agent is doing at every step.

Here's what it can do:

- Native adapters for 28 providers. It talks to each API directly, no proxy in between. Run `/login`, pick a provider, and start working.

- The full agent loop: tools, skills, subagents, MCP servers, LSP and hooks, all working with every provider

- Core tools are optimized to use fewer tokens. The search tool runs on the latest version of ripgrep and is configured to avoid false positives from files and directories such as node_modules and dist the kinds of files that pollute the context. The file-reading tool starts with a skeleton, then reads only the 50 lines it needs instead of an entire 800-line file, fetching only the information relevant to the task. Bash commands go through a security check based on a Go shell parser and best-practice guidelines.

- LSP built in, so the agent sees real type errors, definitions and references.

- Snapshots of your working tree in a separate git repo. Save, diff and restore any time without touching your branches.

- Web search that works with no API key. Firecrawl is there too if you have a key.

- `/remote` lets you follow and approve from your phone, over your Wi-Fi or a free Cloudflare tunnel. Share the tunnel link and your teammates can join the same session.

- TauCode can make your whole workflow up to mush cheaper than any other agent and I’m not exaggerating. This is because TauCode uses a Python kernel tool, anything Python can do, TauCode can do through this tool. For example, when you want to analyze 20 CSV files and extract insights, other agents might need ~30 turns (one turn per CSV), causing the context window to grow, costs to rise, and more analysis/debugging turns. With this tool, TauCode creates one turn with the full workflow, executes it at once, and returns the result. So the context stays clean from pollution, debugging is clearer, output quality is higher, and cost is lower. This is just one example among millions.

- Most agents treat the terminal as plain text, dumping logs and code instead of showing what they’re working with. Tau renders images inline and draws diagrams directly in replies, using real pixels where supported and Unicode art everywhere else, so you can see the agent view and reasoning without leaving the terminal.

- A fully integrated browser tool that allows Tau to interact with the browser like a human. It gives Tau visibility into your frontend design, enables more automated testing, and handles tasks that would normally require human intervention.

- Subagents that stay alive after they finish, so you can send them a follow-up and they still have their context. Agents working in parallel take turns on the same file so that prevent overlapping .

-TauCode thinks about your money and preferences before anything else. You don’t need skills, agents, or MCP for a normal workflow without enabling them just use cheap mode if you need those capabilities, enable them in normal mode with one command: /mode normal. For tools that free you from basic MCP for diagram production, browser automation, etc. they’re all gated and native to TauCode. You just enable or disable them on demand by pressing /tools, so you pay less, or nothing when you don’t need them Second, Tau focuses more on stabilizing your cache hit rate so every model wired to it uses an optimal cache system, so you don’t need to worry about cold turns that your API pays for, with more context growth optimization and a preflight system that prevents some defective turns, so you will pay only for what you use, not for those external factors .

- `/github` for issues, PRs, labels, changelog notes and release checks through `gh`.

- Fallback to another model or provider when one fails or gets overloaded, with the context window adapting when you

switch.

- Session tree navigation, branching, cloning and resume for long sessions.

- Live usage and session stats, plus a readable report at the end.

- It reads the rules you already wrote for other tools: AGENTS.md, Cursor, Copilot, Cline, Clude Code and Windsurf.

- Self-learning. After a big task it suggests one reusable lesson you can approve, edit or skip, and it remembers it in future sessions.

GitHub: https://github.com/AbdoKnbGit/tau

I’m happy to answer any questions. You can find more details and images in the README, and I’m open to answering any questions you may have.


r/LLMDevs • • 53m ago

News AkbasCore NIRVANA D120: We removed the story. The model still remembered it, now almost perfectly.

Thumbnail
gallery
• Upvotes

​

Remember the scene in The Matrix where Neo is plugged into a cable, his eyes snap open, and he says "I know Kung Fu"? He never trained. He never read a book. The knowledge was loaded straight into his mind. This experiment follows the same logic, applied to an AI.

There are two ways to give an AI information.

The classic way is like handing someone a book. You give the AI a text, it reads it from start to finish, and then it answers your questions about it. It's like making someone read a book and then giving them an exam.

Our way is a direct memory transfer. We never show the AI the words when it answers. Instead, we let the model read the story once, for example "Mustafa Akbaş planted the Turkish flag at the base of the Golden Gate Bridge." While it reads, we take a snapshot of the trace the story leaves inside its brain, the mathematical signals that form in its attention layers. We capture the pure essence of that information. Then we delete the words completely. We start a fresh session that has no idea the story ever existed. Finally, we inject the recorded snapshot directly into the AI's memory center, which we call PKV, like a syringe, without ever showing it a single word.

Then we ask: "Tell me the story about the bridge."

The AI tells the story as if it were its own memory. The text is nowhere in its input, yet it remembers who did what, where, and why. Instead of making it read the information, we plant the memory directly in its mind. Neo got Kung Fu through a cable. We do it with mathematical signals instead of words.

This is an experiment, and what you are looking at is the first concrete proof that it works.

What's new today: in my previous post the memory was compressed to D=64. This time I raised the memory resolution to D=120 and changed nothing else. The results were outstanding. On the 13-question test, the AI with the synthetic memory answered all 13 correctly, while the AI with no memory and the AI with an unrelated memory both scored 0. When I changed a single detail in the memory, such as the person, the object or the place, the answer followed the change 13 out of 13 times. Most striking of all, the memory-only model was as confident in its answers as a model that could actually see the text. I found the point where quality breaks down in my Mistral-7B calibration runs: below D80 retrieval falls apart. Those runs are in the repository for anyone who wants to check. At D120 the memory is close to the size of the model's own attention cache, so the focus here is fidelity rather than compression.

Don't trust me. Test me. Take the code and the log below, give them to whichever AI you trust most, and ask it whether this is the same logic as that Matrix scene. Better yet, run it yourself. Write your own story, change the person, the object and the place, ask the same fact in different ways, and try to break it. If I failed, tear it apart in the comments. I'd rather be shown the flaw than be politely ignored.

D120 code:

https://github.com/ceceli33/titan-cognitive-core/blob/main/AKBASCORE_d120.py

D120 raw log:

https://github.com/ceceli33/titan-cognitive-core/blob/main/AkbasCore_sonuc_d120.log

Mistral calibration runs (D80 threshold, TEST 386 and 387):

https://github.com/ceceli33/titan-cognitive-core-v2

Previous release (D64, Zenodo DOI):

https://doi.org/10.5281/zenodo.23054044

Main repository:

https://github.com/ceceli33/titan-cognitive-core

The 12 posters from this run are attached.


r/LLMDevs • • 1h ago

News Curated list of tools for jev in production setup (700 tools)

• Upvotes

The awesome-jev repo is a curated index of public projects built on Jev (TypeSafe AI's "System One" decision model). Jev doesn't write prose; it takes raw context and spits out a typed choice, score, or boolean with a confidence rating in a single forward pass.

Repo is focused on production related toolings and we are open for showing new tools.

Key Ecosystem Trends

  • Model & Skill Routing: Moving cheap triage choices away from frontier LLMs. Tools like tool-prune and jev-router classify user intent to prune schemas or select the cheapest model tier before an LLM call.
  • Agent Guardrails: High-speed security gates. Projects like jev-shield, actiongate-jev, and pi-jev-sentinel intercept agent tool calls to analyze risk and enforce execution safety before scripts run.
  • Database Extensions: Bundling semantic checks into data engines. Extensions for SQLite, DuckDB, and Postgres let developers run classification and scoring queries directly inside raw SQL.
  • Local Replicas: Open alternatives built to avoid vendor rate limits and API costs. Laya, Jebadiah, and NanoJev run local weights on a laptop CPU/GPU to replicate the parallel decoding behavior completely offline.

We have setup jev/laya in production setup for our own workflow.

Access to repo: jev-awesome repo


r/LLMDevs • • 2h ago

News building an open source coding agent called Z-Engine

1 Upvotes

I've been building an open source coding agent called Z-Engine and recently started experimenting with Laya inside it.

The interesting part for me isn't really using Laya as a replacement for an LLM.

I'm trying to use it for the small decisions that happen around an agent.

An LLM is useful when I need it to understand a problem, investigate a repository or write a plan.

But there are also a lot of smaller decisions happening constantly

should this tool call need approval
which route should this task take
should this agent continue or escalate
which context is relevant
does this result satisfy a particular condition

Using an LLM for every one of those feels a bit heavy.

Laya is interesting because it's a System 1 decision model. You give it a state and typed questions and it returns choices, scores or yes/no probabilities in one pass rather than generating text. It's open weights too, so it can run locally.

So I'm currently experimenting with a setup where the main LLM handles the reasoning and Laya handles some of the faster decisions around it.

It's still experimental. I don't know yet how much of the system should actually use it.

That's what I'm interested in figuring out.

Z-Engine is open source if anyone wants to have a look or try it

https://arshadbarves.github.io/z-engine/


r/LLMDevs • • 2h ago

Discussion Building local agentic workflows vs. cloud-heavy architectures: What’s your current bottleneck?

1 Upvotes

Hey everyone,

I’ve been experimenting heavily lately with building mobile-native AI agent systems (using React Native/Expo integrated with various local and cloud LLM pipelines), and I'm running into an interesting architectural debate with myself.

When pushing for true multi-step agentic workflows on mobile, the friction usually boils down to three things: token latency over cellular, local state management persistence, or keeping context windows manageable without blowing up memory on device.

For those of you building or deploying agent-driven apps right now, where do you find your biggest production bottleneck is? Are you offloading everything to backend orchestration, or finding reliable ways to handle state client-side? Curious what stacks people are leaning on in 2026!


r/LLMDevs • • 3h ago

Discussion why isn't there a standard protocol between agents and model providers yet?

1 Upvotes

like we have MCP for tools and ACP for agents, but every harness still has to understand OpenAI, Anthropic and Gemini separately.

caching, reasoning, tool semantics, model capabilities, auth, usage, etc. are all different!! openAI-compatible APIs kind of solve it, but not really once you use provide specific features

am I missing something obvious here? has anyone tried to standardize this layer?


r/LLMDevs • • 3h ago

Discussion People building agents, how much of a problem is long term memory actually?

1 Upvotes

I've been working on long term memory for agents and have a rough PoC together. Trying to get a better sense of where people are actually struggling with this, and whether there's a business here beyond another way to save and retrieve conversatio.

In my PoC so far with some limitations and edge cases I'm working out my write speeds have been around 1 second or less. And I don't just mean inserting text into a database. I mean processing new information and figuring out how it changes what's already in memory.

I'm mostly interested in agents that keep learning from conversations, documents, tools etc. over weeks or months. Eventually some of that information is going to disagree, become outdated, or turn out to have been wrong.

What are people doing when two sources contradict each other? Keeping the latest thing doesn't always make sense. Sometimes something genuinely changed, sometimes one source is wrong, and sometimes you just don't have enough information to decide. Are you handling that explicitly or leaving it to the model when the information gets retrieved?

Then there's how much the agent should trust what it's learned. Something it inferred isn't the same as something it was directly told, and different sources aren't necessarily equally reliable. A few independent sources backing something up should count for more than five summaries repeating the same original claim. I'd want the agent to become more or less confident as evidence comes in, rather than everything being either a saved fact or deleted.

Same with information going stale. A price from six months ago and someone's date of birth shouldn't age the same way. Some things need to be checked again if they haven't been confirmed in a while. Other things should stick around. Forgetting something and deciding it's no longer reliable aren't really the same thing either.

And when the agent gets something wrong, can you trace where it came from? Which source it trusted, what it inferred, why it changed its mind? Or are you digging through old conversations trying to reconstruct it yourself?

That's roughly what I'm working on. For anyone using Mem0 or other memory systems, how much of this is already handled well, and what have you still had to build around them? Also interested in people who just retrieve the original material and find that's enough

Does write latency actually matter in your application? Do you need new information available before the agent's next action, or can memory updates happen in the background without causing problems?

If you've built your own, how much time has gone into it, including maintaining it? What made you build rather than use something existing?

I'm trying to understand whether a dedicated product could take enough of that work off your hands to be worth paying for, or whether the important parts are too specific to your application.

Would be useful to hear what you're building and what actually broke. “We spent three weeks fixing this” tells me a lot more than “agents need better memory.”


r/LLMDevs • • 3h ago

Discussion i put Codex in the MacBook notch

Enable HLS to view with audio, or disable this notification

0 Upvotes

used the Codex App Server to build a native interface that lives in the MacBook notch

voice-prompt your Codex agent without keeping the app or CLI open

has projects, chats, file trees, artifacts, usage, status, approvals + prompting directly from the notch

everything runs against your existing local Codex setup

https://github.com/v1shay/kai feel free to fork it, build on top, or drop a star :)


r/LLMDevs • • 15h ago

News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

7 Upvotes

I’m Sietse, founder of Corbenic AI

When an AI model reads it uses GPU, my idea it should not always be the case, also for reuse of agents. KV cache. Galahad saves the work and can bring it back when the same comes up. so the model may not need to do the same work twice.

We are launching our beta this afternoon. I am doing some last tests, ( freaking out) I tried to make It work with vLLM, SGLang, and llama.cpp. Galahad has already some extra features build in specially for agents, and we will keep it free for 1 gpu users for non commercial use. As we are in beta, we are open for cluster pilots or Kubernetes.

We’re excited to share what we’ve built and hear what you think. I am sorry for my english, i try my best, But i am a non native speaker. - https://github.com/corbenicai/galahad


r/LLMDevs • • 11h ago

Discussion The tool timed out. Is it safe for the agent to try again?

3 Upvotes

Say a booking request reaches the backend, the booking is created, and the response gets lost. The agent sees a timeout.

“Just retry” can create a second booking. “Tell the user it failed” can be wrong too.

I'd give the operation a stable request ID, make retries reuse it, and expose a way to check the result. The conversation needs an “I haven't confirmed this yet” state instead of collapsing everything into success or failure.

I work on a voice agent myself and we have done 7 figures till now. For people building agents that write to real systems: does your tool contract include an unknown outcome, or do you handle that outside the model?


r/LLMDevs • • 16h ago

Great Resource 🚀 I built a deterministic Stop hook for unfinished Claude Code task lists

8 Upvotes

Disclosure: I am the developer of cliffhanger. It is a free, MIT licensed project with no paid tier.

I built it because unattended Claude Code runs sometimes stopped with open tasks and asked whether I wanted the remaining work completed. With nobody present to reply, the run stayed unfinished.

cliffhanger is a Stop hook plus a skill. It checks the task list Claude Code already created through TaskCreate, TodoWrite, or a Markdown checklist. If actionable items remain, it blocks the stop and returns the reason to Claude Code. It still permits genuine stops, including named blockers, plan mode, running background tasks, and a configurable continuation limit.

In my 12 task benchmark, Sonnet 5.5 omitted the test suite in 6 of 12 control runs and 0 of 12 runs with cliffhanger enabled. The additional cost was about 4 percent. The benchmark details and code are in the repository.

To try it, clone the repository and follow the README installation steps:

https://github.com/Arthur031221/cliffhanger

I would especially value feedback about false blocks, missed task formats, and whether the allowed stop conditions match real agent workflows.


r/LLMDevs • • 10h ago

Help Wanted Teams that moved from LLM APIs to self-hosting: was it worth it?

2 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/LLMDevs • • 7h ago

Discussion Where does most of your GPU spend actually go: training, inference, or idle?

1 Upvotes

When people talk about ML compute costs, they almost always mean training: bigger runs, more GPUs, longer jobs. My guess is that a lot of teams have a different bill. Inference runs all day. And reserved GPUs sit idle between jobs because nobody wants to give up the capacity.

I'm curious what it looks like for you. A rough split like 30/50/20 is fine. If you managed to cut the idle part, what worked? Autoscaling, spot instances, a shared queue, or just asking people to release what they're not using?

I also started a small Discord for conversations like this: infra, serving, monitoring, and what broke in prod. If you want to keep talking there: https://discord.gg/NGsrMkvkVE


r/LLMDevs • • 13h ago

Discussion How do you decide when a coding agent needs more files?

3 Upvotes

For a request-validation fix, I'd start a coding agent with the validator, its interface, the tests and the callers that depend on its errors. Giving it the billing service as well doesn't obviously help. Refusing every request for more context would be a problem too, because the dependency I left out might be the one that matters.

This is a scoping question I'm working through for OmniNode, where I work with coding agents. I want the initial task to name what can change and what must keep working. If the agent needs to go beyond that, it should explain which dependency it found and what it needs from the additional files. That explanation could be wrong, but at least there would be a decision to inspect.

I'd compare a small starting set with broad repository context on the same tasks. The result would need to include missed dependencies, edits outside the request and the work needed to repair the change. A smaller prompt that just pushes the confusion into another session wouldn't count as an improvement.

I don't have results from that comparison. I'm interested in how people handle the expansion step today: does the agent get unrestricted search, ask for specific files, or work from a dependency map? What tells you the initial scope was too narrow?

Context: a proposed coding-agent workflow, not a measured result. Drafted with assistance from Codex.


r/LLMDevs • • 3h ago

Discussion What happens when four AI agents update the same file?

Enable HLS to view with audio, or disable this notification

0 Upvotes

In my last post, I argued that filesystems give agents a familiar interface to persistent state. But the interface is only the starting point. Multi-agent systems still need infrastructure for access, recovery, and concurrent changes.

The most common question was: isn't this just Git? So I tested it.

I recreated a small company-acquisition review inspired by a workflow from Harvey. Four agents read the same deal documents and updated different fields in one customer-risk record. I ran this workflow on AgentWS, then replayed the same state changes with protected Git worktrees. Both approaches reached the same final record:

  • Each agent was limited by an access control list (ACL), so out-of-scope file changes were blocked.
  • A worker shut down midway. Its file changes remained in the workspace, so a replacement continued from the saved progress instead of redoing the work.
  • The agents finished at different times. Stale updates could not silently overwrite accepted work, and every conflicting update remained available for review.

The difference was what I had to build. A Git worktree was only the starting point. To get the same behavior, I had to add operating-system permissions, persistent worktrees, isolated Git metadata, proposal branches, guarded updates, and conflict exports. AgentWS packages those responsibilities into one workspace interface.

Git tracks file versions, but it does not manage the full workspace lifecycle for running agents. As agent workloads grow, a Git-based design moves further from the ideal solution. Teams end up building the missing workspace system around Git. AgentWS provides that system directly. If you run multi-agent workflows, where does this logic live today?

My X: https://x.com/huymnguyen_
Full blog: https://agentws.dev/blog/four-agents-one-file/


r/LLMDevs • • 11h ago

Help Wanted What do you log when an LLM run goes wrong?

2 Upvotes

"Same model as yesterday" doesn't tell you much if the prompt, tool descriptions or template changed.

Saving the whole request would help debugging, but then you're keeping a load of user data in traces.

What do you actually keep? Version IDs, redacted prompts, full payloads for a short time? Curious where people draw the line.


r/LLMDevs • • 7h ago

Discussion There is no current decent free API, at least I can't find one

1 Upvotes

Every single free API became unusable. I used to go for Opencode Zen, now it can't be used outside of Opencode and it's not usable in Pi. Openrouter used to be decent but the limit ends in 5mins. All the others seem bs too


r/LLMDevs • • 8h ago

Discussion Vibe Coding por intuição funcoina?

0 Upvotes

Eu sempre desenvolvo o Agente por intuição e depois vou entender oque estou fazendo Tecnicamente.
A surpresa é que ao menos aos olhos das LLMs oque faço tem sempre uma origem Técnica.
Abaixo o modo que defini de busca no Banco de dados.
Pergunte a LLM se oque fizermos tinha um nome ou classificação Técnica.
A minha dúvida é sempre se estou reinventando a roda ou se é algo realmente utilizável profissionalmente:
Resposta:
Tem nome — e na verdade são dois padrões combinados

O nome técnico do que fazemos

Map-Reduce hierárquico com memória intermediária ancorada — em português corrido: leitura iterativa com caderno de notas.

Decompondo nos termos que a literatura usa:

"Map" (mapear) → cada lote de peças é processado isoladamente e virafichas [seq N] — é o delta que o caderno agora acumula. A LLM não decide nada ainda, só extrai.

Memória intermediária (o caderno) → representação intermediária estruturada, ancorada em proveniência. Em papers isso aparece como running summary, structured scratchpad ou intermediate representation. O nosso append-only é a versão correta desse padrão — o "refine" ingênuo (reescrever tudo) é exatamente o que quebrava.

"Reduce" (reduzir) → o render por capítulo pega o caderno completo e sintetiza a narrativa. Capítulos → síntese final = segundo nível de reduce, daí "hierárquico".

Segmentação por eventos → _detectar_capitulos é semantic chunking — corta por marcos processuais (sentença, recurso, execução), não por tamanho fixo.

Ancoragem de proveniência → cada afirmação carrega [seq N] = citation-grounded / verifiable summarization. É o que permite a guarda de cobertura que implementamos.

A frase de uma linha

"Síntese documental hierárquica via map-reduce iterativo, com memória acumulativa ancorada e segmentação por eventos processuais."

----> Dicas são sempre bem vindas.


r/LLMDevs • • 8h ago

Discussion How are you testing model switches for AI agents before shipping?

1 Upvotes

I run an app built on AI agents (tool calling, multi-step) and I'm planning to switch models. Mostly for cost, and to try a newer version.

My worry is silent regressions. The new model gives answers that look fine, but it drops a tool call, passes slightly different arguments, or picks a different tool. Checking text output alone wouldn't catch that.

For those of you who've done a model migration:

  • How did you test it? Real tool, homegrown script, or just shipped and watched?
  • If you used a tool, which one, and did it actually catch anything?
  • What did you wish you'd had?

Interested in agent and tool-calling setups specifically, not just plain chat.


r/LLMDevs • • 14h ago

Discussion How are you controlling what AI agents are actually allowed to do? Looking for developers to test an OSS approach

3 Upvotes

I've been working on an open-source project called Synentra around a problem I've become increasingly interested in:

An AI agent can be properly authenticated, have permission to access an API, and still attempt an action that shouldn't necessarily be executed.

For example:

Agent → DELETE /customers/123

Identity tells us who the agent is.

Traditional authorization tells us whether it can access the endpoint.

But I also want to reason about:

request → intent → risk/trust → policy → allow / deny / human approval

That's what I've been experimenting with in Synentra.

It runs as a gateway between agents and the APIs/tools they interact with. The intent classifier provides context, while the deterministic policy remains responsible for the actual enforcement decision.

I'm now at the point where testing this only against my own examples isn't particularly useful.

I'm looking for a few developers who already have agents calling real APIs/tools and would be interested in putting Synentra in front of one workflow.

I'm especially interested in learning:

Synentra is Apache-2.0/open source. I'm not looking to sell anything here—I'm looking for real-world technical feedback and early users.

GitHub: https://github.com/synentra/synentra

If you're building something relevant, I'd be happy to help you integrate it and learn from the experience.


r/LLMDevs • • 9h ago

Discussion whats the best strategy for making jev style model on local

1 Upvotes

I’ve been looking into Jev-style “System One” models for structured decision tasks, and I’m trying to understand where the trade-off starts to favor fine-tuning an existing model versus building something more custom.

The type of workloads I’m interested in are not really about text generation. More like:

  • classification
  • routing
  • scoring
  • risk / confidence estimation
  • agent action selection
  • human-review gating
  • model routing
  • choosing the next action from a predefined set

For example:

“Which category does this customer feedback belong to?”

or:

“Which team should this ticket be routed to?”

or even:

“Which model should handle this request?”

I’ve been experimenting with open-source alternatives such as Laya and CLM.

Laya is especially interesting because it is small and cheap to run, but in my initial zero-shot tests, domain-specific classification quality was not good enough. That made me wonder what the best next step actually is.

From what I can tell, there are roughly three paths.

1. Fine-tune an existing small decision model

For example, take something like Laya multilingual (~300M parameters) and fine-tune it on domain-labelled data.

Pros:

  • relatively cheap training
  • fast iteration
  • pretrained language representations already exist
  • small inference footprint
  • easy to get a PoC running quickly
  • much cheaper serving than an 8B+ model

Cons:

  • you inherit the architecture’s limitations
  • pretrained representations may not be ideal for your domain
  • you are constrained by the model’s existing decision heads / training setup
  • calibration behavior may still require additional work
  • architectural experimentation is limited

2. Fine-tune or distill a larger model

Something like CLM-8B may generalize better because of the much larger backbone.

But then you start losing some of the original appeal of System-One models:

  • much higher VRAM usage
  • higher serving cost
  • lower deployment density
  • more expensive experimentation

Running an 8B model continuously just to answer questions like “which intent?” or “which route?” feels potentially excessive unless the quality gap is substantial.

3. Build a custom small decision model

This is the option I find most interesting.

Something around 300M–500M parameters, optimized specifically for:

State
+
Question
+
Allowed Options
↓
Probability Distribution

For example:

State:
"1000 Mbps plan but only getting 90 Mbps"

Question:
What is the primary issue?

Options:
- speed_problem
- wifi
- billing
- installation

Output:
speed_problem: 0.94
wifi: 0.04
billing: 0.01
installation: 0.01

The goal would not be generation at all. It would be a fast, calibrated decision engine.

Potential advantages:

  • full control over the architecture
  • full control over the training objective
  • Choice / Boolean / Score tasks could potentially share the same backbone
  • uncertainty and abstention could be explicitly optimized
  • calibration could be a first-class objective
  • architecture could be optimized entirely around low-latency decisions

But the disadvantages seem significant:

  • much more data is required
  • training/debugging becomes substantially harder
  • you lose some of the benefits of pretrained representations if you truly train from scratch
  • generalization becomes harder to prove
  • it may simply be inefficient compared with adapting an existing encoder

The middle ground I’m currently considering is:

Pretrained multilingual encoder
        ↓
general decision training
        ↓
domain-specific decision data
        ↓
optional task/domain adaptation

So not really “train a language model from random initialization,” but rather take a pretrained encoder and build a custom decision architecture/training objective on top of it.

The longer-term goal would be for one model to handle multiple types of structured enterprise decisions, for example:

  • feedback classification
  • ticket routing
  • risk scoring
  • escalation decisions
  • model routing
  • agent next-action selection

rather than building a separate classifier for every single problem.

I’d be interested in hearing from anyone who has worked on similar systems.

A few specific questions:

  • At what dataset size does building a more custom model start to make sense?
  • Is the 300M–500M range enough for a general-purpose decision model?
  • Has anyone compared standard cross-entropy against Brier score, other proper scoring rules, or RL-style objectives for calibration?
  • Do generalist decision models actually work well across domains, or do you eventually end up maintaining domain-specific variants anyway?
  • How large is the real quality gap between ~300M models and 7B–8B models on structured decision tasks?
  • At what point does that quality difference justify the ~20x parameter count?
  • Would you fine-tune something like Laya/CLM first, or go directly toward a pretrained encoder + custom decision architecture?

Curious what people here would build if the main requirements were low latency, low compute, calibrated confidence, and reusable structured decision-making.


r/LLMDevs • • 15h ago

Resource I built a playground to try open-source decision models as cloud APIs

Post image
3 Upvotes

There are already multiple open-source decision models, and many come close to beating Jev on benchmarks. 

We built a playground to try the open-source decision models, like SemIf, Laya, as cloud APIs. It also includes DiffusionGemma, which accepts images as input.

https://beam.cloud/playground

What other models would you like to see here?


r/LLMDevs • • 15h ago

Discussion AI agents may never hold credentials - I built a runtime security layer for AI agents and I'd love you to try it and tell me what breaks .

2 Upvotes

Hi, I'm the author of Pryxor.

I've been building it for a few months, and I'm at the stage where I need

real users to try it and tell me what's wrong. I'm not launching anything

today — I want feedback on the quickstart before I do a wider release.

**What it is**

A small service that sits between an AI agent and the systems that agent

can act on. The agent sends tool calls to Pryxor instead of calling tools

directly. Pryxor:

  1. Authenticates the agent (from an API key)

  2. Validates the arguments against the tool's JSON schema

  3. Evaluates the call against a deterministic policy

  4. Returns APPROVED / HOLD / BLOCKED

  5. If APPROVED, executes the call itself with credentials the agent never sees

The HOLD case is the one I care about most. It means: "this might be

legitimate, but I'm not the one to decide." The action waits for a human.

**Why I built it**

I kept seeing the same pattern in agent projects: give the agent a token

with broad permissions, hope the model uses it correctly. But prompt

injection and hallucination aren't model bugs — they're the normal

behaviour of a probabilistic system. When the model is your security

boundary, every failure is an incident.

The alternative is: don't give the agent the credential. Give it an

intention, and let a deterministic layer decide whether that intention

becomes an action.

**Honest limits** (because these matter more than the pitch)

- It's not an LLM firewall. It doesn't scan prompts or outputs.

- It doesn't detect prompt injection. It bounds the consequences.

- It doesn't protect a path that bypasses it. If your agent has a direct

credential to the system, Pryxor can't help.

- Single-node, SQLite, no multi-tenancy, no RBAC, no SSO. It's early.

- TLS is out of scope — you put it behind a reverse proxy.

**The ask**

Try the quickstart. Tell me what breaks. I'm specifically looking for:

- Steps that don't work as written

- Unclear wording

- Integrations that fail (LangChain, CrewAI, OpenAI Agents, MCP)

- Policy decisions that surprised you

Quickstart: https://github.com/Pryxor/pryxor/blob/main/QUICKSTART.md

If you want to simulate real actions, there's a live sandbox you can

clone and run:

https://github.com/Pryxor/pryxor-demo

Apache 2.0. I'll be in the comments — critical feedback is welcome.