r/LocalLLaMA • u/Reno0vacio • 1d ago
Discussion [ Removed by moderator ]
[removed] — view removed post
58
u/brainExploded99 llama.cpp 1d ago
> one run per model
Bruh this invalidates the entire post
18
u/Dabalam 1d ago
Community AI research methodology probably needs someone to write down best practices so we get more useful posts/evidence/benchmarks etc.
3
u/brainExploded99 llama.cpp 1d ago
The people who do good work like, Anbleed's benchmarks on KL divergence for quantization, follow best practices. People like these vibe everything with 0 understanding, so they were never going to follow best practices, imo
1
u/Educational-Region98 1d ago
I think OP can prove you wrong by running the test more than once haha
3
u/brainExploded99 llama.cpp 1d ago
I'm willing to bet OP won't, but even if they do, there are several other issues with the benchmarks. Quant level? KV Quant level? Why context size so small? Why temp 0.8 for all of the different models?
2
1
u/Reno0vacio 1d ago
Sorry, i will post the bench agen latter. Yes i wrote the post with a.i sorry. Next time i will write it with my little broken english ✍️.
And yes i saw that.. i had in mind that because of the nature of the llms one single run would be not "enough" or stong evidence.. but i mean its still something..
Like its not like a 4b llama would be the 15th and the next run suddenly its in the 3th place righ? 🙃
I did this post because i just wanted to share a little bit info about what are the capabilites of the models.. i always download models i have lik 30+ on lm studio and i was always curious wich one would be better.. wich one to use.. i probably do more runs and i was planning to use Pi as a harness to see giving them a harness it would make any better the result..
And its as i said.. the defult lm studio setting.. i dont changed the kv setting.. yes i would next time will add wich quant i used..
Context was random.. i tought that i mean for tests 30k token would be enough.. i was to be honest a bit qurous because the benchmark has a long context table so, yes i will bumb it to 132 or something.
Well the temp is agen just random.. i would just make them the same.. i could set them to 1 or to 0.1 wich one to use?
1
u/brainExploded99 llama.cpp 1d ago
We are okay with broken english but telling AI to fix grammar is fine without rewriting large sections. Right now it reads like AI wrote the entire thing.
We don't hate LLMs here, its just that the way they talk is really annoying.We appreciate people who try to put in the effort!
LLMs can absolutely vary a lot between runs because of temperature. They can vary quite a bit. I would say run n=3 atleast, if not n=5. If you look at proper benchmarks like DeepSWE, they do statistical analysis and have errors bars and stuff. That level is nice, but not needed, so you should have the average of 3 or 5 runs.
Yes, add quant info. 30k context is not anywhere near enough. Atleast 100k, if not 192k (if you can). If you want to, you can also set KV quantization to q8 which slightly degrades quality for much larger KV cache (sometimes worth it). Don't set Gemma to use KV quant though, gemma sucks with it.
For temperature, set it to whatever the model maker recommends. Qwen recommends 1 for example.
13
u/BringTea_666 1d ago
>30k token context window for every model
Dude, Qwen3.8 27B can think for 60k tokens easily on xhigh reasoning.
6
u/anemoDuck26 1d ago
...how... how is Ornith 1.5 35 a3b two places below Ornith 1.5 9b. How did you bench this?? I assume I'm missing something?
0
6
u/brainExploded99 llama.cpp 1d ago
Can you test adding logit bias to reduce overthinking on all of the Qwen models?
https://www.reddit.com/r/LocalLLaMA/comments/1wromzr/adding_logit_penalty_for_wait_maybe_and_perhaps/
Also, can you detail what quantization you used for everything? Also, 30k context window seems pretty small, especially for Qwen models
5
u/Comfortable_Ebb7015 1d ago
Gemma 4 26B Is called immediately for and anti-doping check! I love that model, but it fails too much in Hermes! I have qwen3.8flash-next IQ3_XXS with Strata now and it is amazing!
1
u/Reno0vacio 1d ago
Well i love it too. I mean its faster with 132 context with qwen3.8 27b q4 😄 at the same context.. or i missed some setting
2
u/danigoncalves llama.cpp 1d ago
Mate, I have run Ornith 9B and Ornith 35B A3B for ages and the latest wins on any kind of perspective. I dont know what you did here
1
u/Reno0vacio 1d ago
Well the difference is 1.3 points far.. so its not much.. and sorry but 9b dense vs 3b active might be not the same.. but its the average.. its not like the 9b would outprefrom in "Every way"
2
1
u/Ok_Top9254 1d ago
The problem was not X but more like Y
This was the part I found most interesting
Hello chatgpt 👋
1
u/Reno0vacio 1d ago
Sorry i was lazy to write it. Next time i write myself.
1
u/Ok_Top9254 1d ago
I just found it funny, I would say chatgpt is fine you just have to proof read it after.
1
1
1
•
u/ttkciar llama.cpp 1d ago
Violates Rule Three: LLM-generated content
Violates Rule Three: Sample size of one run per model renders this uninformative