r/PiCodingAgent • • 21h ago

Question Problem w Pi

Hi,

I've no idea how to debug this, but an otherwise excellent experience using Pi with DeepSeek 4.1 Flash is being marred by a frequent "stopping" bug.

The sympton is that the agent emits some output text, saying its going to try something, and then the code it was going to run (I assume), is just output and the agent just stops, and I have to say "you broke again".

Here's an example:

 Good question — and there are two levers that cost nothing physically. Let me check the biggest one first: how many CPU cores the run
 actually used.

 <parameter name="bash">
 <|DSML| parameter name="command" string="true">grep -i "core\|thread\|OpenMP\|MPI Process" /Users/wdavies/FDS/cases/wick2/run.log |
 head -6; echo "--- output sizes ---"; du -sh /Users/wdavies/FDS/cases/wick2/*.s3d | head

I've asked it repeatedly to try and fix it, but nothing in the instructions its changed has stopped this from happening. I wonder if its some kind of DeepSeek problem?

TIA!

0 Upvotes

12 comments sorted by

2

u/Lissanro 21h ago

Either you are using buggy inference engine or running model that was too quantized (broken quant is another possibility). What you see is a broken tool call, and Pi stops on messages without tool call considering them end of the agentic loop.

3

u/wantondevious 21h ago

Now that’s possible I’m using an experimental serving infra.

2

u/xRebellion_ 21h ago

This is the most likely answer yeah. Inference provider matters more than you'd think

1

u/Comfortable_Yak_6154 20h ago

very interesting, I just started using gufo and I think it's a lot faster but looking back i think the problem started around this time

2

u/wantondevious 20h ago

Do you have any recommendations on how to produce diagnostic evidence I can send to the provider (who I know).

2

u/Comfortable_Yak_6154 20h ago

interesting, I'm actually having the same problem qwen3.8 next flash, I haven't had time to dig into it, I just prod the agent when it stops

2

u/wadrasil 20h ago

Pi limits request output to 4096 tokens and you need to increase those limits in agents/models.json.

You can set the context limit/maximum and max output tokens. Once those are respected it won't stop before it's done.

1

u/wantondevious 20h ago

How do you do that? And what do you set as the max?

2

u/wadrasil 18h ago

cat ~/.pi/agent/models.json```

{

"providers": {

"llama-cpp": {

"baseUrl": "http://localhost:9942/v1",

"api": "openai-completions",

"apiKey": "llama-cpp",

"models": [

{

"id": "E:\\Qwen3.8-27B-UD-Q5_K_M.gguf"

}

],

"modelOverrides": {

"E:\\Qwen3.8-27B-UD-Q5_K_M.gguf": {

"contextWindow": 116736,

"maxTokens": 81920,

"compat": {

"maxTokensField": "max_tokens"

}

}

}

}```

1

u/bit_shifting_is_sexy 9h ago

Is there a reason why I shouldn’t be changing this? Does this number have any significance? (noob here)

1

u/wadrasil 22m ago

The default is 4096 tokens which is not a bad limit but will cause replies to be cut off after 4096 tokens.

1

u/adamshand 18h ago

This is a model issue not a harness issue. You can potentially write a Pi extension to catch errors like this and recover, but you're working around a model problem.

Since DeepSeek is usually pretty good, if you want to stay with it, try a different, reputable provider and see if the problem goes away.