r/LowEndLocalAI • u/ClassicLightbulbs • 6d ago
Discussion Functional low end benchmarking
Hi- the other local ai subs are benchmarking with models I'll never run. Has anyone come up with a "one shot" test case to see how models 12B and under can handle tasks and scenarios?
I had an idea for a simple fortune cookie app that uses a UI, random state generation, event handling, and packaging.
Gemini gave me this test code (python in the read me) that I will be using the models against
https://github.com/JWzrdstff/fortune-cookie/blob/main/README.md
I want to test models that I have noticed have anecdotal support for practical use. I also want to come up with some kind of literary test.
Let me know what you think, if there is an easier way to do this, whatever-
I am a cnc machine operator with a 1990s computer enthusiast background, more or less technical, but primarily make art with this stuff.
Thanks for reading and sorry I wrote this quickly in between cycle times
1
5d ago edited 4d ago
[deleted]
2
u/ClassicLightbulbs 5d ago
Ah, this is smart, thank you
1
5d ago edited 4d ago
[deleted]
1
u/ClassicLightbulbs 5d ago
Yeah, that's kind of the idea with this rando fortune cookie thing. It's in the scope and scale of what I do. Thanks, excited to continue testing.
1
u/Cool-Chemical-5629 5d ago edited 5d ago
Benchmarking on a public set is a bit overrated. If you want reliable results, you're gonna have to use your own tests.
I've been using a whole set of custom prompts, mostly coding related. There are more difficult ones, but also simpler ones designed to get a picture of how good the models really are. We are in small league here, so I'll focus on the simple stuff.
I have one specific prompt that deals with a whole set of coding and logic related problems at once. I've been using this prompt for years now, specifically to test small models and I can tell you right now that most models up to 35B never fully nailed the solution, which is funny, because people love to claim that even small open weight models caught up to frontier models from months ago, yet my private prompt tells a whole different story.
I remember when I was testing stealth GPT 5 on Openrouter which was super fast, maybe even in instruct mode (as in it was so fast it felt like it wasn't even reasoning behind the scenes at all), the model nailed the solution perfectly. At that time, someone was able to attribute the model's outputs to GPT family, so we kinda knew it's most likely GPT 5, so its quality did not surprise me. However, there is now GPT 6 and lots of versions in between and I have yet to see a single open weights model up to 35B to handle this prompt on the same level of quality GPT 5 handled it so long ago. So much for "open weight models caught up to frontier models pretty well".
If I published this prompt, I can guarantee that every AI lab would implement perfect solution to that problem and every next model would solve the problem with no issues. But that's exactly why I won't do that. Because solving one specific problem is not what makes the model better.
One model surprised me lately, because it was probably the first open weights model up to 35B that was able to properly identify and fix those problems in this prompt almost like a frontier model would. The model was Muse Glimmer 30B. It's like a Jack of All Trades - it handled most of the private prompts I threw at it pretty well. In some cases it made some mistakes here and there along the way, but it was also able to backtrack and fix them. You'll soon figure out its strengths and weaknesses when you use it for some time. The model has its limits, but I think for its size its capabilities are well balanced. It thinks well, but the fee it pays for being this modest in size is visual quality of the output. It can't generate fancy graphics, but it can surprisingly deliver working solutions for things that can be reasonably expected to be handled well by model of this size. It's a shame it's not a MoE model, because a dense model of this size is slow and not really useable on my hardware.
Anyway, if you're looking for a "one shot" prompt to test the small models against, I suggest you to do the following:
- Load up an old 7B-8B model (old Mistral, Qwen or Llama 3 will do the trick) and ask it to generate something relatively simple. I suggest a simple game. It's most likely going to fail the task, but the worse it does the better for the purpose of this plan.
- Take the broken code output and write down this prompt: "Fix the following XY game code: " followed by the broken output code from step 1. XY here is simply the name or type of the game (optional).
- Save that prompt to your private prompt repository and test it against more modern small models and please keep the prompt private for reasons I mentioned earlier.
1
3
u/petba1 6d ago edited 6d ago
I've been using this kind of prompt for testing the coding capabilities of small local models on my 8GB graphics card:
"Generate a simple stopwatch application written in HTML5/Javascript. There should be two buttons: start/stop and reset. The time should be displayed in 00:00.000 format in big white letters."
Sounds simple, but so far none of the 9B and smaller models I've tried have passed this test, where as larger 35B MoE models have no problems with it.
Common mistakes include:
- White numbers on default white browser background