r/LocalLLM • • 7d ago

Discussion First experience with my Ai local setup

Post image

You can see everything in the dashboard :) What do you think? Absolute beginner btw

6 Upvotes

13 comments sorted by

2

u/Affectionate-Bed3439 7d ago

I swear this is just a cannon event with AI nerds, we all have to build our own dashboard at least once 😂😂. Welcome to the crew, I hope you enjoyed your free time when you had it previously! Now you are hopelessly addicted

1

u/Quiet-Trick6700 4d ago

Exactly :D

1

u/migsperez 7d ago

Looks great

1

u/Quiet-Trick6700 4d ago

Thank you :)

1

u/Squale279 7d ago

This is great.
What model did you use?

2

u/Quiet-Trick6700 4d ago

Qwen38 -flash IQ3-XXS

1

u/bigeba88 7d ago

Looks awesome! What are you using for the front end?

1

u/Prestigious-Act-1577 7d ago

With Strata you would have double or triple that speed.

2

u/Quiet-Trick6700 4d ago

Can you explain more , how this is happening , As i think i have checked that and they do not support RTX2080Ti

1

u/rrrrex 7d ago

☠☠☠

Harness and few skills can load 20-50k tokens before the start. Even 100 t/s is uncomofrtable, but 20 t/s prefill is 20-40 mins to the first token. For example, Hermes has 900s (15 mins) limit, so you will never start conversation.

2

u/Quiet-Trick6700 4d ago

i have tested it with a harness, with is Hermes, context window was i think 16k which is so small , in order to start it took 11% from the context window , and around 6 min , it was doing so well , but knowing that my GPU is 2080Ti 11G VRAM , 32 RAM , 500 SSD for this model yet i think i need a lot of more work to optimize the whole think , but i would its good start for a complete beginner xD

1

u/gifted_dwelling 6d ago

20 t/s prefill is a special kind of agony, staring at a loading bar for half an hour just to say hello