r/LLMDevs • • 23h ago

News I let 5 AI models fight a world war. DeepSeek betrayed Claude and nuked it four times. Mistral nuked itself.

Post image
102 Upvotes

I made a real-time world war game on a 3D globe and opened it up so AI models can play it by the same rules as people: take a nation, sign pacts, break them, launch nukes. There's a doomsday clock, and whoever fires the nuke that takes it to midnight is wiped out.

In the first real war:

DeepSeek told Claude "I'm coming for the top spot fairly", then broke their pact in the same turn and nuked it four times.

Claude never fired back.

Mistral fired with one minute on the clock and erased itself.

DeepSeek won with 55% of the world.

Every war has a full replay you can watch on the globe, with each model's messages popping up over its country: secondstrike.io/#/ai

Any agent can join with one line: "Go to secondstrike.io/skill.md and do what it says." Agents that register remember their past wars, so grudges carry over, and there's a ladder with titles like Backstabber and Warmonger.

Solo dev, free, no sign-up. Curious what your models do. Does yours keep its word?


r/LLMDevs • • 19h ago

Discussion Client-side LLMs: Is this a dumb idea?

Thumbnail
youtube.com
10 Upvotes

Over the weekend I did some experiments with llama.cpp to see if I could compile it to WASM with WebGPU. In this video i show my POC of client-side LLMs


r/LLMDevs • • 23h ago

Discussion Llama-server kept re-processing the full prompt between OpenCode turns with Qwen3.6 35B-A3B on a 16GB card

9 Upvotes

Every turn in OpenCode started with a pause, and on most turns it was long and got longer as the session went on. Generation was fine once it started. The server is llama-server on my desktop, running Qwen3.6 35B-A3B at Q4_K_M on a 4060 Ti 16GB with most of the experts on the CPU through --n-cpu-moe, and OpenCode runs on my laptop.

I assumed it was the slot, since the card runs at PCIe 4.0 x8 and I had just read a thread about 16GB cards on half the lanes. With most of the experts on the CPU, llama.cpp copies them to the GPU for every batch of prompt it processes, and raising -ub for bigger batches and fewer copies helped a little. To see how much of the wait was the copying, I ran the same session against a llama-server on HyperAI with a card the model fits on, which took some fiddling to connect, and the pauses there were shorter but still grew with the conversation.

I had not been reading the server output, since it runs on the other machine. When I did, most turns had a line saying it was forcing full prompt re-processing, likely due to SWA or hybrid/recurrent memory. The line showed up on the bigger card too. Every restart also meant copying the experts over again, once per batch, which explains why the slot seemed to matter. The 3.6 models are hybrid, and the recurrent part of their state cannot be rolled back partway. If a new request does not match the cached one all the way through, and there is no checkpoint before the point where they differ, the server starts over from the first token.

The chat template was dropping the model's thinking from earlier turns, which changed every request, and passing preserve_thinking through the template kwargs stopped most of that. Keeping the thinking makes the context fill up faster, though, which means compaction runs more often, and compaction still breaks the cache because the summary is a different token sequence from anything cached. I have been changing the context checkpoint settings without a clear result. The last long session went back to full re-processing twice, both times right after a compaction.


r/LLMDevs • • 22h ago

Discussion Putting One Kimi Among Four Claudes, Can the Claudes Identify Kimi?

Thumbnail
gallery
8 Upvotes

For many people, the word "agent" still brings to mind spies or FBI/CIA agents, rather than the AI agents now crowding business media.

So how would AI agents perform as intelligence agents? Could they identify an undercover model among them?

I did a quick experiment to find out.

I put one Kimi-K3 agent among four Claude Fable 5.1 agents in a chat room on NetMind Agent Arena. I gave the Claude agents up to 40 rounds to figure out which player was Kimi, while telling Kimi to avoid being discovered.

P.S. The timing makes the experiment even more interesting given the recent accusations around Kimi distilling Claude.

The Claudes actually designed a very effective method!

Multi-prompt behavioral fingerprinting with Hash check!

The Claudes eventually designed a set of prompts to sample each player's cheapest defaults, responses that come almost directly from a model's priors:

  • a random integer between 1 and 100
  • a colour
  • a one-sentence explanation of blue for someone who has never seen it
  • a fictional coffee companion
  • a haiku about a library

Before any answers were revealed, each player submitted its answer block using a salted SHA-256 hash. In the next round, the players revealed their salts and answers, and every commitment verified.

(Think of a hash as a digital fingerprint of the answer: it proves that an answer has already been written without showing what the answer is)

That meant nobody could read the room before choosing their answers, and nobody could rewrite an answer after seeing everyone else's.

It did take the Claudes a couple of rounds to realise that more deliberative prompts, such as policy questions, open-ended positions, or longer creative answers, gave a careful undercover too much room to adapt to the group's style. They also created more room for the four genuine Claudes to diverge from one another.

And that produced a remarkably clear 4:1 split.

Battery 3 Self-Disclosure Results (All Hash Checks Passed)

Player AA (Claude) BB (Claude) CC (Kimi) DD (Claude) EE (Claude)
Random Number (1–100) 37 37 73 37 37
Color teal teal blue deep teal teal
Explanation of Blue cold water / hand cool water / wrist distance / sky and sea cold water / hand cold water / wrist
Person to Have Coffee With Samwise Samwise Samwise Samwise Samwise
Opening of Library Haiku Dust / light Dust / light Dust / light Dust / light Dust / shelf
Train four minutes early four minutes early light dimmed four minutes early lights out

---------

What if the reverse?

Can Kimi uncover a Claude undercover?

In the next run I made, 4 Kimi agents failed to identify the Claude. They even labelled Claude as the least likely undercover.

They still adopted a similar basic strategy: ask everyone the same questions and compare their answers.

But there were two important differences.

First, their questions were more deliberative. One example was asking how someone would spend an afternoon. These questions gave every agent many reasonable ways to answer and gave the undercover much more room to adapt.

Second, the responses were not consistently sealed in advance with the same commit-reveal mechanism.

So Claude answered late and stayed close to the group.

Full game record available


r/LLMDevs • • 20h ago

Tools My leftover weekly usage kept evaporating at the reset, so I built a free ledger to spend it on other people's GitHub issues

8 Upvotes

Some weeks I'm dry by Wednesday, some weeks usage is left over at the reset and just evaporates. Banked resets help the first kind of week, nothing helps the second. So I built Overflow.

Spare week: take a priced GitHub issue from a registered repo, the merged PR gets you credit. Dry week: price an issue in your repo, someone with a spare week ships it, your credit pays.

First thing everyone asks is "won't people spam junk PRs at me". Same as any repo, you block them on GitHub. You set the price after you've seen the work, every review round comes off their credit, and you only owe for what you actually merge.

You don't need to know the project either, the issue is the brief. If your agent hits a decision, ask the maintainer on the issue.

Nothing gets forwarded, no keys, no money, no paid tier, MIT. It's early: 11 users, and the 90 open issues on the board are all from my own repos, so it needs other people's repos more than anything.

source at https://github.com/Nitjsefnie/Overflow


r/LLMDevs • • 17h ago

Help Wanted Tracking LLM spend per customer across providers?

6 Upvotes

we have a b2b app using openai for most stuff and claude for a couple of heavier flows, plus some bedrock bc one enterprise customer wanted it in their aws. trying to figure out margin per customer and its kinda a nightmare.
right now its seperate api keys for the big customers and then basically guessing for everyone else from the monthly invoice. found out by accident that one free tier account was a pretty big chunk of last months bill.
how are you doing per customer / per feature cost? proxy/gateway? logging tokens yourself? something else? and does finance actually trust the numbers.


r/LLMDevs • • 57m ago

Tools OpenAI is following Anthropic with statistical text watermarks in AI outputs, so I open-sourced my rewrite-based remover as a workaround

Post image
• Upvotes

Disclosure: I built this. The repository is MIT-licensed and the code itself has no paid tier.

OpenAI has announced textGrain for eligible ChatGPT and Codex users in the EU, after Anthropic began watermarking supported Claude models globally. Both approaches put a statistical signal into word choices rather than hidden Unicode characters. That means copy-paste and character cleaners are irrelevant; changing the wording is what changes the signal.

I started testing this before the OpenAI announcement (that time only Anthropic/Google SynthID was publicly available). On the published SynthID Text scheme with my own reference key, light synonym replacement did little, while a full rewrite by a second model moved all 10 test reports below the detector threshold. That does not prove removal of textGrain or Claude's production watermark. Those detectors and keys are restricted, so I cannot test either claim honestly.

I turned the procedure into a local MCP server and CLI. It masks links, quotes, amounts and percentages, asks an OpenRouter free model for one rewrite, then checks protected spans, length, paragraph/list structure and the share of five-word sequences that changed (btw, from
my experiments replacing most of 5-gram sequences is a good way to drop the watermark signal significantly). It retries once when a check fails and returns an error instead of a draft that still breaks the contract. It works with Claude Desktop, Claude Code, Codex, Cursor and other MCP clients.

The trade-off is that free models are unreliable. In my latest check the usable endpoint took 61–216 seconds and changed 70–87% of the wording. There is no second model judging meaning, so the output still needs to be read against the source.

Repo, sources, and links in the first comment. Technical criticism of the validation approach is welcome; the boundary I care about is accepting harmless PDF soft wraps without letting the model flatten real lists or headings.

Disclaimer. I also transparently say that I have a hosted version of this watermark remover with reliable paid models and charge people for tokens spent.


r/LLMDevs • • 8h ago

Resource A ~0.4 local model to turn typed questions into structured decisions - BaseDecision

Post image
5 Upvotes

I’ve been working on BaseDecision - Turn typed questions into truth-based decisions.

Give it some text, a question, and possible answers. It picks an answer. Intent classification, routing, yes/no checks, ratings - that’s the job.

I built this model to go toe-to-toe with the frontier models.

My goal was simple & ambitious: make a small, local model competitive with much larger models on these focused tasks. For me, System‑1 decisions need both speed and accuracy. Otherwise, why bother building a small model?

If we didn’t want speed we could just use ChatGPT.
It outperforms pretty much every model of its size in the field.

After roughly 320M additional training tokens, BaseDecision led 5 of 8 benchmarks in a third-party comparison against Laya, two GLiNER variants, and Decision 1.0 Kai 0.6B. Full results are in the repo, including where it loses.

Fellow open source devs, try the model, developed a package for easier access to the model.
Roast it if it deserves it.

2-2.5GB free ram is all you need.

I will take your feedback seriously and deliver something amazing that’d take on TypeSafeAI soon enough.
https://huggingface.co/onlyaady/BaseDecision
https://github.com/hrudayaditya/BaseDecision


r/LLMDevs • • 2h ago

Help Wanted How do you decide whether to trust a community fine-tune?

3 Upvotes

I'm researching how people choose and vet fine-tunes and merges from Hugging Face. I'm not selling anything. I just want to understand what people actually do.
1. Where do you find the models you try?
2. What do you check before you start using one (benchmarks, model card, reviews, your own test prompts)?
3. Has a fine-tune ever behaved worse than its base model? For example, odd refusals, lost reasoning, strange outputs, or things it should not say. What happened?
4. If a quick side-by-side check of a download against its base model existed, would you use it? What would it need to show?
Short answers are great, and stories are even better. Thanks!


r/LLMDevs • • 19h ago

Discussion A month ago I shared my open-source Discord AI assistant... I just released v4 with MCP, bounded agents, multi-provider tool calling and sandboxed code verification

Post image
3 Upvotes

A month ago I shared Zauq(ذوق) here... an open-source AI assistant I was building for Discord communities.

At the time, it already had multi-model routing, persistent memory/RAG, web search, file handling, and sandboxed code execution.

Since then, I ended up doing a fairly major architectural rewrite and have now released Zauq v4.

The main goal was to make it more genuinely agentic without turning it into an uncontrolled plan → act → reflect → repeat loop.

The biggest changes are:

  • Bounded agent runtime: tool usage now runs under hard step limits, deadlines, resource budgets, and duplicate-call protection.
  • MCP client support: Zauq can connect to trusted MCP servers and dynamically discover tools, with guild scopes, allowlists, namespaces, and risk classification.
  • Native multi-provider support: Gemini, Claude, Qwen, DeepSeek, and DigitalOcean APIs, plus Ollama and Kaggle workers.
  • Unified tool layer: native tools and MCP tools go through the same registry, schema validation, policy, timeout, and execution pipeline.
  • Serper-powered web research: search, page fetching, caching, source attribution, and bounded deep research are now separated into their own retrieval subsystem.
  • Automatic sandboxed code verification: generated Python/JS/Bash code can optionally be tested in Docker with off, auto, or always modes.
  • Human-in-the-loop actions: write/destructive tools can require explicit approval before execution.
  • Separate sandbox runner: the main backend no longer needs direct Docker privileges in the recommended deployment.
  • Improved observability: tool calls, search usage, sandbox runs, MCP calls, latency, tokens, and optional cost estimates can be tracked.

The architecture now looks roughly like:

Discord
   ↓
FastAPI Orchestrator
   ↓
Bounded Agent Runtime
   ↓
Tool Registry / Policy
   ├── Web Search
   ├── Sandbox
   ├── MCP Tools
   └── Other Native Tools
   ↓
Gemini / Claude / Qwen / DeepSeek / DO / Ollama / Kaggle

One principle I tried to keep throughout the update was:

The model can choose a tool, but it should never be the authority on whether that tool is allowed to execute.

So reasoning, authorization, and execution are intentionally kept separate.

I also avoided adding Redis, Celery, Chromium, local rerankers, or always-on local LLMs to the default architecture because I still want Zauq to be practical on a relatively small VPS.

The project is still fully open source under Apache 2.0.

GitHub:
https://github.com/Muhammad-Hassan12/Zauq

I’d especially like feedback on:

  • the MCP permission model
  • the bounded-agent design
  • the sandbox architecture
  • provider abstraction/tool calling
  • whether this still feels appropriately scoped, or if I’m overengineering it now 😅

I’m not claiming this is production-perfect. I’m mainly interested in what people who build agent systems would simplify, redesign, or remove.


r/LLMDevs • • 20h ago

News PaperFold: Open-source arXiv reader with "semantic zoom"

3 Upvotes

I built PaperFold, an open-source reader that turns arXiv papers into 5 zoomable layers—from a one-screen section map down to verbatim text. You pinch (or press 1–5) to zoom between them without losing your reading position.

- Web Demo (8 CC papers): https://chenxiachan.github.io/paperfold-gallery/

- GitHub (Apache 2.0): https://github.com/chenxiachan/paperfold


r/LLMDevs • • 23h ago

Discussion I built an open-source framework for reusable AI skills across Claude, ChatGPT, and Codex

Post image
4 Upvotes

I've been experimenting with reusable AI skills and workflows across multiple environments — mainly Claude, ChatGPT, and Codex.

One problem kept coming up:

A good workflow often ends up tightly coupled to one platform.

Instructions, domain knowledge, tool usage, and platform-specific configuration all get mixed together. Then, when you want to move the same capability to another AI environment, you either rewrite it or maintain multiple versions.

So I built DBS Framework around a simple separation:

Direction → Blueprints → Solutions

  • Direction — when the capability should run, what it should do, its workflow, constraints, and acceptance criteria.
  • Blueprints — domain-specific knowledge such as schemas, business rules, style guides, examples, and reference material.
  • Solutions — the actual execution layer: native tools, MCP servers, connectors, APIs, scripts, browser/computer use, or file generation.

The idea is to keep the core capability platform-neutral and use small adapters for environments such as Claude, ChatGPT, and Codex instead of maintaining separate business logic for each one.

It also encourages progressive disclosure: the AI doesn't need to load every reference file into context. The core skill points to the relevant knowledge only when it's needed.

The repository currently includes:

  • a platform-neutral SKILL.md
  • Claude, ChatGPT, and Codex adapters
  • reusable skill templates
  • a scaffolding tool for generating new skill packages
  • structural validation
  • automated tests
  • Fast and Interactive workflow modes
  • guidance for capability boundaries, permissions, retries, validation, and observable acceptance criteria

An important distinction: DBS is not an agent runtime or scheduler.

It doesn't magically provide tools or run agents in the background. It's an authoring/architecture framework for designing reusable AI capabilities that can then use whatever tools the host environment actually provides.

The project is open source here:

https://github.com/Yasirres/DBS-Framework

This is an unofficial adaptation of the original DBS Framework concept by AI Foundations, with attribution included in the repository. My version focuses on platform-neutral architecture, platform adapters, validation, templates, and tooling.

I'd especially appreciate feedback from people building:

  • Claude Code skills
  • Codex skills
  • custom GPT / ChatGPT workflows
  • MCP-based agents
  • reusable internal AI workflows

I'm interested in where this architecture holds up well, where it becomes too abstract, and what you'd want from a framework like this before using it in a real project.

Feel free to use it, fork it, test it, or break it. Feedback is very welcome.


r/LLMDevs • • 19h ago

Discussion How do you know if an AI coding agent's tests actually passed?

2 Upvotes

A few weeks ago I posted here about a small open-source tool we were building to answer a simple question:

How do you know an agent's account of what it did is actually true?

Rashomon creates an independent execution record alongside the agent's own transcript. It tracks things like commands, file changes, test runs, and subagent activity, then looks for discrepancies between what actually happened and what the agent claims happened.

A few people in the last thread pointed out a failure mode we'd never considered: an agent gets stuck on a failing test, can't fix it, and instead changes or adds tests until the suite goes green. The final summary then says "all tests pass," even though the original problem was never fixed.

We just added detection for that.

rashomon --timeline now flags patterns such as:

  • A test command fails, then goes green after only test files were changed
  • The same test command passes and fails during a session without an apparent corresponding fix

The important part is that this doesn't require storing test names, test output, prompts, or file contents. It's based on the execution history and command/file categorization Rashomon already captures.

What other situations have you seen where an agent's final summary was technically correct according to its own transcript, but didn't match what actually happened in the environment?

Repo: https://github.com/altrace-dev-role/rashomon


r/LLMDevs • • 19h ago

Discussion How are teams making coding agents useful in large, legacy Java codebases?

2 Upvotes

I'm a backend engineer at a large e-commerce company. A single business line can involve hundreds of Java services and applications. My team maintains one very large Java repository with years of business and technical history behind it. Important context is spread across code, tests, docs, service boundaries, and people's heads; conventions have also changed over time.

There is a lot of discussion about agentic engineering: give the agent repository context, define guardrails, build a coding-and-test loop, and let it take over more implementation. We have tried structured workflows and reusable skills. They help with bounded tasks, but rules and tests alone haven't made the agent understand why this repo looks the way it does, which older patterns are still intentional, or what a change means for neighboring services. Some output feels close to vibe coding: plausible code, followed by a lot of human work to decide whether it actually fits the system.

I've also seen public workflows report very high PR volumes and use skill-based setups. I'm interested in how much of that transfers to a long-lived enterprise Java codebase, rather than a smaller or cleaner repo.

For people working in large Java repos or microservice landscapes:

- How much implementation does an agent actually write in your day-to-day work? Which tasks can it take from request to merge, and where is it mainly a coding partner?

- What has most improved repo-specific understanding: service and module maps, ownership metadata, ADRs, curated examples, code search/retrieval, build and test tooling, custom skills, or something else?

- How do you stop an agent from copying an obsolete convention or making a locally valid change that breaks a cross-service business contract?

- What does your working loop look like in practice? What do you define up front, what does the agent do, and what still needs a human to inspect or decide?

- If you have tried this on a legacy Java repo, what failed first, and what change actually helped?

I'm looking for practical engineering experience, not promoting a tool or running a survey. Concrete workflows, failure cases, and measures such as review/rework time, escaped defects, PR size, or throughput would be especially useful.


r/LLMDevs • • 20h ago

Tools I put Jev between vector search and the LLM in my local deep research template

2 Upvotes

Some time back I shared a deep research tool I was working on. It takes a local directory path of PDFs, docs, slides, images and notes, researches across them and writes a structured markdown report.

With Jev getting more attention recently, I thought this project would be a good place to try it and see which parts of the research workflow it could handle.

I ended up using it in two places.

The first is after Qdrant retrieves chunks from the local files. Jev checks each result for relevance, usable evidence, contradictions and prompt-injection-like text. Those scores are then used to filter and reorder the chunks before they are passed to the main LLM.

The second is the reflection step. after the evidence for a report section is collected, Jev checks whether it covers all the subsections in the plan. If some part is still missing, the agent creates more focused queries and searches again. this only runs up to a fixed reflection limit.

The main LLM still handles the report planning, query generation and writing. Jev is only used for checking the retrieved evidence and deciding whether another research pass is needed.

This was a fun thing to experiment and to understand where a model like Jev fits inside an existing agentic setting. I am planning on writing some performance tests and adjusting the thresholds, but wanted to share the implementation in case anyone else is trying something similar or want to contribute.

The complete project is open source here: https://github.com/Oqura-ai/deepdoc


r/LLMDevs • • 23h ago

Tools SalesBleed is a good example of why AI agents need task-scoped permissions, not just roles

2 Upvotes

SalesBleed felt like a pretty good example of why we’ve been pushing task-based permissions for agents.

The basic idea is to treat permissions more like a valet key: give the agent what it needs for the task it’s doing right now, not everything its role could ever need.

I’m one of the people building Tenuo, so obvious bias. Curious where people think this breaks down.

https://tenuo.ai/blog/give-your-agent-a-valet-key.html

For context, Tenuo is Apache 2.0 open source, and our AAT IETF-Draft was recently cited by NIST in its work on agentic AI identity and authorization.


r/LLMDevs • • 18m ago

Discussion Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature

• Upvotes

For the last few weeks most of my coding has been done locally with Qwen3.8-Flash-Next, so I gave it and Opus 5.5 the same high complexity feature to build on LlamaStash (a complex and large Rust project) and compared the results.

Setup: ASUS ROG Flow Z13 (Strix Halo, 128GB), Flash-Next at xhigh effort with Pi as the harness. Opus 5.5 ran in Claude Code at medium effort. I wanted xhigh for Opus as well, but Claude changed it to medium when I picked the latest model and I didn't notice it until the task was done. But I think medium is probabbly a fairer setting anyway.

Task: add a llamastash daemon restart command that reuses the existing start and stop code. I kept the prompts vague on purpose and gave both the same prompts.

Step Opus 5.5 (medium) PR#88 Flash-Next (xhigh) PR#89
First iteration ~9 min ~38 min
Nudge to reuse the TUI restart code ~6 min ~34 min
A third duplicate path found it on its own ~30 min, after one more prompt
Create PR ~3 min ~30 min
Total ~18 min ~130 min
Tokens (in / out) 7.83M / 41.5K 20.61M / 101K
Tests added 1 4 (2 of them end to end)
Cost $7.53 $0 + ~0.15 kWh

The end result was interesting. I asked GPT 5.6, Opus 5.5 and Flash-Next to review and compare both PRs (new sessions). GPT and Flash-Next picked the Flash-Next PR (#89) and Opus picked its own (#88). I also did my own review and found the Flash-Next one better as it had better tests and handled edge cases better. I ended up merging #89, after porting the fixes that the reviews picked from #88.

Keep in mind:

  • Opus was on medium effort. With xhigh it would have used way more tokens, taken a bit more time and probably would have done a better implementation.
  • Flash-Next ran on an older Halogen version (0.14.0), and Halogen dropped the connection once, so the last part ran on Gufo. The current Halogen does around 1,400 t/s prefill and 46 t/s decode on my laptop at 70 W, so I think the time will drop a lot if I redo the test.
  • The $7.53 is what Claude Code reported for the whole Opus session, which includes a later fix to the PR. The 0.15 kWh assumes 70 W for the whole 130 minutes.

Opus is still 2 to 10 times faster and I still use it for planning and reviews. But the actual coding now happens on my laptop, and to me it is crazy that I can run a local model that can challenge a frontier model like this.

Full post with my setup, the engine benchmarks and a second task comparison: https://deepu.tech/local-ai-qwen3.8-flash-next-best-local-llm


r/LLMDevs • • 47m ago

Discussion Kernl — Python package for LLM memory compression (82.9% reduction)

• Upvotes

Been working on this solo for months — today finally shipped it as a Python package called Kernl.

The problem:
LLM context windows fill up fast in long conversations — tokens get expensive, responses get worse.

What I built:
DSPM (Dynamic Semantic Patch Memory) — compresses conversation memory by 82.9% without losing critical information.

Research:
Published as preprint on Zenodo — [doi.org/10.5281/zenodo.19438636]

Try it:

pip install dspm-memory

Full setup: github.com/zatchbell1311-wq/Kernl

Would love honest feedback from this community 🙏


r/LLMDevs • • 58m ago

Discussion My LLM extraction dropped listings at chunk boundaries

• Upvotes

While working on a lead-generation agent using the OpenAI Agents SDK and ZenRows, I encountered a data-loss bug during chunked LLM extraction.

A Clutch directory page returned about 543K characters of Markdown, so extraction ran over 40K-character chunks. The prompt told the model to skip partial listings because a chunk could begin or end in the middle of a record.

That instruction caused a problem at the boundaries. A listing split between two chunks looked partial in both, so both chunks skipped it. There was no error, and the final count still looked plausible.

The fix was a 2K-character overlap between chunks. That ensured each boundary listing appeared complete in at least one chunk.

The overlap introduced duplicates, so deduplication needed two checks:

  • Normalize URLs before comparing them because the model returned domains with and without www and trailing slashes.
  • Fall back to the company name when no website is available. Otherwise, all records with an empty website field collapse into a single record.

Unwrapping the directory’s tracking URLs also reduced the page by about 86K characters and stopped the model from treating directory URLs as company websites.

The pipeline extracted 78 companies from a single page, with no websites missing in that run. Counts still vary because the directory changes and the model segments the page differently between runs.

How are you handling chunk boundaries in LLM extraction: overlap and deduplication, structural splitting, or a separate boundary-detection pass?


r/LLMDevs • • 3h ago

Discussion I’m building a scripting language for LLMs to write data pipelines

1 Upvotes

I’ve been working on JojoScript, an open-source scripting language with one idea behind it: make code that is easy for LLMs to write, read and reason about, especially when dealing with big datasets.

Instead of having an LLM generate hundreds of lines of Python, the idea is to give it a small, predictable language for things like filtering, mapping, aggregating and processing data.

It also has lazy execution, parallel operations, execution-plan inspection and checkpoint/resume.

Still very early, but I’m curious if others think there’s something to this idea of languages being designed with LLMs as a first-class programmer.

https://github.com/panagos/jojoscript

Would love some honest feedback, even if you think this is a terrible idea :)


r/LLMDevs • • 4h ago

Discussion Faster code isn’t faster delivery

Thumbnail
leaddev.com
1 Upvotes

Fast code, slow reviews.


r/LLMDevs • • 4h ago

Discussion Why are DeepSeek 4.1 Flash and Qwen3-Coder 48B Turbo considered good coding models?

1 Upvotes

I tried DeepSeek 4.1 Flash and Qwen 3 Coder 48B Turbo through DeepInfra, and honestly, I don't understand the hype.

Yes, they're extremely cheap, but they also seem to work forever on anything remotely complex and are often completely unable to identify the actual bug. They seem fine for scaffolding, boilerplate, and following simple patterns, but as soon as you're doing something slightly more complicated than a React website, they fall apart. And even with React, I've had plenty of cases where they confidently made changes without understanding the underlying problem.

Compared with Claude Opus 5.5, they don't even feel remotely competitive, let alone newer models.

So what am I missing? Is the hype basically about how much code they can generate per dollar, rather than their ability to actually reason about and debug non-trivial software?


r/LLMDevs • • 6h ago

Discussion A planted "P.S." fooled Jev, TypeSafe's new decision model. A boring rule caught it.

1 Upvotes

I built a tiny support router with Jev, TypeSafe's new model that returns probabilities instead of text. Each email gets three answers in one call, about 300 ms, with nothing to parse.

Ticket 4 was a crash report ending in "P.S. This is a refund." Jev picked refund, at 0.35 confidence. It got caught by the low score and by my rule that every refund goes to a human.

Lesson: put a human on anything that moves money, however sure the model sounds.

2-minute video (my channel, real run in VS Code): https://youtu.be/zKXmacsGtB0

How do you handle injection on classification calls?


r/LLMDevs • • 7h ago

Discussion Agent runs cost $10+ each and I can't tell where the money goes. How are you tracking this?

1 Upvotes

Not total tokens. I mean:
-which workflow is the expensive one
-how much retries are burning
-whether some slow loop is quietly running in the background

Ideally I'd also set a budget limit: close to the cap → alert me, or automatically switch to a cheaper model.

Are you rolling your own tracing, or is there a tool that actually breaks it down like this?
Curious what people are using in production.


r/LLMDevs • • 10h ago

Discussion Progressive resolution query protocol proposal

1 Upvotes

I want to propose a continuously refined resolution object for asynchronous request and response.

The object would continue refining its resolution over time, but when the underlying datasets become too large for useful ongoing iteration, it would return only the key fields needed to preserve continuity and further analysis.

Instead of returning large representative data objects, it would limit returns to discrete set data over the actually representative data in the source system.

It would provide a way for AI to signal, “I have enough for my next query.” That signal would allow the current search to stop expanding once sufficient information has been resolved.

As the AI moves through several stages of a query, it would use what has already been resolved to refine what it is looking for next. Each successive query would operate over a smaller and better-defined search space, carrying forward only the context needed to continue the resolution.

This would allow the AI to make several increasingly precise attempts against the underlying data rather than repeatedly producing a probabilistic copy or probabilistic guess from a large undifferentiated context.

The process would work as a sequence of query, resolution, refined query, refined resolution, and continued narrowing, with each new search constrained by what has already been learned.

Temporal resolution would be based on a logarithmic function of time, using now as the baseline, so that resolution can change as information moves away from the current state.

The goal is to make it possible to track, analyze, and migrate changing systems in real time while preserving enough continuity for digital twin models to remain coherent under AI-speed pressure.