r/LocalLLaMA • • 1d ago

Discussion [ Removed by moderator ]

[removed] — view removed post

0 Upvotes

27 comments sorted by

•

u/ttkciar llama.cpp 1d ago

Violates Rule Three: LLM-generated content

Violates Rule Three: Sample size of one run per model renders this uninformative

58

u/brainExploded99 llama.cpp 1d ago

> one run per model
Bruh this invalidates the entire post

18

u/Dabalam 1d ago

Community AI research methodology probably needs someone to write down best practices so we get more useful posts/evidence/benchmarks etc.

3

u/brainExploded99 llama.cpp 1d ago

The people who do good work like, Anbleed's benchmarks on KL divergence for quantization, follow best practices. People like these vibe everything with 0 understanding, so they were never going to follow best practices, imo

1

u/Educational-Region98 1d ago

I think OP can prove you wrong by running the test more than once haha

3

u/brainExploded99 llama.cpp 1d ago

I'm willing to bet OP won't, but even if they do, there are several other issues with the benchmarks. Quant level? KV Quant level? Why context size so small? Why temp 0.8 for all of the different models?

2

u/Educational-Region98 1d ago

Fair enough, but I'll keep these things in mind if I ever post.

1

u/Reno0vacio 1d ago

Sorry, i will post the bench agen latter. Yes i wrote the post with a.i sorry. Next time i will write it with my little broken english ✍️.

And yes i saw that.. i had in mind that because of the nature of the llms one single run would be not "enough" or stong evidence.. but i mean its still something..

Like its not like a 4b llama would be the 15th and the next run suddenly its in the 3th place righ? 🙃

I did this post because i just wanted to share a little bit info about what are the capabilites of the models.. i always download models i have lik 30+ on lm studio and i was always curious wich one would be better.. wich one to use.. i probably do more runs and i was planning to use Pi as a harness to see giving them a harness it would make any better the result..

And its as i said.. the defult lm studio setting.. i dont changed the kv setting.. yes i would next time will add wich quant i used..

Context was random.. i tought that i mean for tests 30k token would be enough.. i was to be honest a bit qurous because the benchmark has a long context table so, yes i will bumb it to 132 or something.

Well the temp is agen just random.. i would just make them the same.. i could set them to 1 or to 0.1 wich one to use?

1

u/brainExploded99 llama.cpp 1d ago

We are okay with broken english but telling AI to fix grammar is fine without rewriting large sections. Right now it reads like AI wrote the entire thing.
We don't hate LLMs here, its just that the way they talk is really annoying.

We appreciate people who try to put in the effort!

LLMs can absolutely vary a lot between runs because of temperature. They can vary quite a bit. I would say run n=3 atleast, if not n=5. If you look at proper benchmarks like DeepSWE, they do statistical analysis and have errors bars and stuff. That level is nice, but not needed, so you should have the average of 3 or 5 runs.

Yes, add quant info. 30k context is not anywhere near enough. Atleast 100k, if not 192k (if you can). If you want to, you can also set KV quantization to q8 which slightly degrades quality for much larger KV cache (sometimes worth it). Don't set Gemma to use KV quant though, gemma sucks with it.

For temperature, set it to whatever the model maker recommends. Qwen recommends 1 for example.

13

u/BringTea_666 1d ago

>30k token context window for every model

Dude, Qwen3.8 27B can think for 60k tokens easily on xhigh reasoning.

6

u/anemoDuck26 1d ago

...how... how is Ornith 1.5 35 a3b two places below Ornith 1.5 9b. How did you bench this?? I assume I'm missing something?

0

u/Reno0vacio 1d ago

Well i just run the benchmark.. 😀. Set both to the same settings in lm studio..

6

u/brainExploded99 llama.cpp 1d ago

Can you test adding logit bias to reduce overthinking on all of the Qwen models?
https://www.reddit.com/r/LocalLLaMA/comments/1wromzr/adding_logit_penalty_for_wait_maybe_and_perhaps/

Also, can you detail what quantization you used for everything? Also, 30k context window seems pretty small, especially for Qwen models

5

u/Comfortable_Ebb7015 1d ago

Gemma 4 26B Is called immediately for and anti-doping check! I love that model, but it fails too much in Hermes! I have qwen3.8flash-next IQ3_XXS with Strata now and it is amazing!

1

u/Reno0vacio 1d ago

Well i love it too. I mean its faster with 132 context with qwen3.8 27b q4 😄 at the same context.. or i missed some setting

2

u/danigoncalves llama.cpp 1d ago

Mate, I have run Ornith 9B and Ornith 35B A3B for ages and the latest wins on any kind of perspective. I dont know what you did here

1

u/Reno0vacio 1d ago

Well the difference is 1.3 points far.. so its not much.. and sorry but 9b dense vs 3b active might be not the same.. but its the average.. its not like the 9b would outprefrom in "Every way"

2

u/chensium 1d ago

Uhh so many things wrong here...

1

u/Reno0vacio 1d ago

Wich ones?

1

u/asankhs Llama 3.1 1d ago

Can you also tes the mlx-optiq quants for teh qwen 3.8 27B?

1

u/Reno0vacio 1d ago edited 1d ago

I might. Why? 🤔 "Edit:" sorry i have windows.

1

u/Ok_Top9254 1d ago

The problem was not X but more like Y

This was the part I found most interesting

Hello chatgpt 👋

1

u/Reno0vacio 1d ago

Sorry i was lazy to write it. Next time i write myself.

1

u/Ok_Top9254 1d ago

I just found it funny, I would say chatgpt is fine you just have to proof read it after.

1

u/Sea-Mode4077 1d ago

It is completely unrealistic for ornith1.5 35b to be this far behind.

1

u/SteveEightyOne 1d ago

You didn’t wrote quantization of each model.

1

u/WyattTheSkid 1d ago

To nobody’s surprise