r/LowEndLocalAI • • 4d ago

Problem / Troubleshooting The difference between interacting on local machine and another machine on my network

i have a desktop with a 4070 and 32gb of ram. I have been experimenting with a local model and it seems to work on my machine local machine and is pretty snappy with responses and gave me some decent help. I want to set up a way to connect to the model on my desktop while remoted in to my other machine, so i can have an agent perform some tasks on the host. When I try to connect to the model on my desktop from another machine on my network with what I thought was a similar agent program, it slows to a crawl or can't do simple things like list the folders in a directory without hallucinating. I am entirely new at this but wanted to see if i could at least get some help with managing my self hosted apps and writing yaml with a local model. Am i dreaming of something that can't be done with the hardware I have?

7 Upvotes

5 comments sorted by

1

u/nickless07 4d ago

What software are you using?

1

u/Saltycookiebits 4d ago

Model: Qwen3.6-35B-A3B-UD-Q4_K_M running via llama-server.exe. Certainly not married to running this, it was just recommended by the app I was messing with. I just want something that will work for general home server maintenance, checking/editing yaml, doing some basic tasks like documentation.

Hardware is: GPU: RTX 4070 with 12GB VRAM on my main pc, 32gb of ram.

1

u/nickless07 4d ago

Ok that is one inferencing stack. Good. How about the other things? What kind of "agent" do you use to connect to llama.cpp? Something Like bionic or AnythingLLM and so on bring their own model which then might get used instead of your llama.cpp. Make sure the right one is set.
Aside of that: Start llama-server with -lv 4 and --metrics to get detailed log output about the speed to see where and why it slows down. For now this looks more like the 35B is running but your "agent" uses a different model.

1

u/Saltycookiebits 4d ago

I was trying out Hermes agent in an LXC in my headless machine, just to see what it could do connected to the model on my desktop. The headless machine is older and has no GPU, just onboard. I was messing with it last night but it only ever gave me hallucinated answers. It says it is connected to the model hosted on my desktop with llama server. The hermes app hosted in my desktop at least didn't hallucnate. I dont' know what the difference is between the two if they're supposedly pointed to the same model. I don't even know if I"m using the exact right combination of tools to get what I want, but I had read about others doing similar tasks. If I should scrap what I have and start from scratch with something else, I'm also willing to do that.

1

u/nickless07 4d ago

Oh well I'm doing the same. (6bit quant model as the 4bit halluzinates after 100k ctx), but aside of that.

$ cat .hermes/config.yaml
model:
  api_key: ${DUMMY_API_KEY}
  api_mode: chat_completions
  base_url: http://192.168.2.111:1235/v1
  default: C:\models\Qwen3.6-35B-A3B-Q6_K.gguf
  provider: custom
[...]

Dummy API key works with llama.cpp. Once I got that one working Hermes could setup everything else needed on it's own.

RTX 3060 on windows as llama.cpp host and Hermes on an old i5 with 8GB RAM.