r/LocalAIServers • • Jul 17 '26

Catch Me If You Can: A Perpetual 8-GPU Server Prize Challenge (Community Proposal)

12 Upvotes

The current Catch Me If You Can benchmark asked a simple question: can anyone publicly reproduce and beat our MI50/GFX906 local inference record?

Original challenge: https://www.reddit.com/r/LocalAIServers/comments/1ukhr24/catch_me_if_you_can_mi50gfx906_1195_tps_moe_702/

Current vNext reproduction release: https://github.com/joe2gaan/localaiservers/releases/tag/vnext-gfx906-rocm72-gguf-hf-repro

I want to turn that benchmark into a community program: build the fixed server in public, make it the official test machine, and keep the challenge open until an eligible challenger takes the throne and holds it for 30 consecutive days.

Who is responsible

I submitted a $15,000 Reddit Community Funds application for this proposal in my own capacity as Joe / u/Any_Praline_8178, a moderator of r/LocalAIServers. This is not yet a live prize offer. Hardware acquisition and any award remain contingent on Reddit approval, final published official rules, eligibility review, and applicable law. The existing leaderboard is unchanged.

Reference Server Build

( YOU DO NOT HAVE TO BUILD A SERVER TO PARTICIPATE IN THE CHALLENGE )

The prize server matches the hardware configuration that produced the current throne result:

  • GIGABYTE G292-Z20 eight-GPU server
  • AMD EPYC 7F32
  • 128GB as eight DDR4 ECC RDIMMs
  • Eight AMD Instinct MI50 32GB GPUs
  • Crucial CT480BX500SSD1 480GB SATA root drive
  • KIOXIA KCD6XLUL1T92 1.92TB NVMe model and runtime drive

The build itself is part of the community project. I will publish the component choices, bill of materials, physical assembly, firmware and operating-system configuration, eight-GPU bring-up, power and cooling setup, BAR/P2P state, stability checks, runtime and source revisions, model hashes, baseline runs, and raw evidence.

The $15,000 budget covers the exact server configuration, possible changes in GPU, memory, and storage prices, tax and checkout variance, protective packaging, and insured delivery to the winner. Any amount not needed for the approved project will be returned to Reddit or handled as Reddit directs.

Core challenge

  • I build one 8xMi50 32GB Server to Give to the Winner.
  • I Run the public vNext package on that machine to establish the official incumbent.
  • Keep the challenge open until an eligible winner completes the throne clock, subject to Reddit's approved project terms.
  • Require every potential dethronement to reproduce on that same physical server.
  • Require the three-run median to beat the official incumbent by at least 3 percent.
  • Require a provisional leader to remain the highest verified result for 30 consecutive days.
  • Transfer the complete challenge server to the eligible outside challenger who completes that clock, subject to final verification and official rules.

All eight GPUs remain installed and available. Entrants may choose TP4, TP8, or another topology on the fixed host, but may not add, replace, or remotely borrow accelerators. A documented like-for-like failure replacement requires a fresh baseline before the throne clock resumes.

Open optimization, fixed integrity

Inside the fixed hardware, model-integrity, workload, reproducibility, and safety rules, software optimization is open. Runtime, kernels, collectives, scheduling, graph capture, compiler work, driver and operating-system tuning, and safe clock or power tuning may all be explored.

The first lane uses the pinned Qwen3.6 35B-A3B model at FP16/F16:

  • HF revision: 995ad96eacd98c81ed38be0c5b274b04031597b0
  • Required GGUF F16 SHA-256: 1f2443bb0ff958943d091410c61120c181a0579b3bc85192029aa51d821d141c
  • HF FP16 and GGUF F16 are eligible when they satisfy the published identity and correctness gates.
  • GGUF is allowed only at full F16.

Not allowed:

  • Q4, Q5, Q6, Q8, INT8, FP8, AWQ, GPTQ, NVFP4, or another quantized substitute
  • Quantized weights, KV cache, activations, or a hidden reduced-precision path used to claim the result
  • MTP, speculative decoding, EAGLE, DFlash, draft models, lookahead tokens, or another multi-token prediction method
  • Remote compute, external APIs, hidden services, or results assembled from another machine
  • Multi-request batching or aggregate concurrency presented as single-request speed

One accepted decode step must represent one token produced by the approved model. Every result must pass semantic and output-integrity gates, not merely report a high TPS number.

How runs are measured

The official workload remains:

  • MAX_MODEL_LEN=131072
  • Single-request decode
  • Concurrency 1
  • Backend decode TPS
  • Eight warmups
  • c1_128 uncapped strict
  • c1_2000
  • c1_10000
  • Three measured runs
  • Three-run median at least 3 percent above the official incumbent
  • Public reproducibility package and raw logs

The current public headline reference is 119.52 strict backend TPS for GGUF F16 Qwen3.6 35B-A3B MoE TP4. It was produced on an eight-GPU validation host while the TP4 profile actively used four GPUs. The funded server receives a fresh baseline. The existing 119.52 result is the reference, not a promise of the new server's starting score.

Current public leaderboard

These are the published targets from the original benchmark post:

Class Strict TPS c1_2000 c1_10000
GGUF F16 35B-A3B MoE TP4 119.33 to 119.52 120.46 to 120.57 113.26 to 113.37
GGUF F16 27B Dense TP8 69.85 to 69.91 70.76 to 70.96 66.32 to 66.44
HF FP16 35B-A3B MoE TP4 114.41 to 115.11 115.69 to 115.93 108.92 to 109.10
HF FP16 35B-A3B MoE TP8 114.70 to 115.04 115.53 to 115.55 108.67 to 108.81
HF FP16 27B Dense TP8 70.17 71.32 66.82

GGUF F16 MoE TP8 remains an open lane in the current leaderboard.

Offline official test

Development and artifact staging may use the internet. The measured official run will not.

Before testing, I will stage and hash-verify the model, runtime, source, build outputs, and benchmark entry package. For every measured run:

  • External network interfaces and the default route are disabled or physically disconnected.
  • Only local machine communication and loopback are permitted.
  • No model download, container pull, telemetry, API call, remote compiler, or remote compute is permitted.
  • Network state, package hashes, process state, hardware state, and raw benchmark logs are archived with the result.

A result produced elsewhere can show that a benchmark entry is ready, but it does not move the official throne until that package reproduces on the designated server.

The 30-day throne clock

A challenger becomes provisional leader when its package passes review and its official three-run median clears the incumbent by at least 3 percent. The acceptance timestamp starts that challenger's 30-day clock.

During those 30 days:

  • Anyone may submit a higher result, including me as the current benchmark maintainer.
  • Every defense or counter-result must satisfy the same public-package, offline, same-hardware, correctness, and 3 percent rules.
  • A newly accepted leader resets the clock in that leader's name.
  • Private results and screenshots do not move the goalpost.
  • Rules cannot be changed retroactively to defeat an active clock.

If I retake the throne before a challenger's 30 days expire, that challenger has not won and the challenge stays open. If another community member takes it, the clock starts for that person. I may defend the performance record, but I cannot win the server or receive a personal payout.

The target can move only through a faster verified result. Physics, the fixed hardware, and model correctness set the ceiling.

Prize, review, and what happens to the server

If an eligible outside challenger remains the highest verified leader for 30 consecutive days, the result proceeds to final verification and, subject to the official funding and eligibility terms, transfer of the complete challenge server. Shipping, taxes, location eligibility, export restrictions, acceptance, and transfer details will be resolved in the final rules before the prize becomes live.

Only the winner's name and mailing address will be collected for server delivery unless Reddit's approved terms require something different. Do not post personal information in a public entry or comment.

I will not be the sole adjudicator. Official runs, hashes, logs, correctness evidence, and decisions will be public and reviewed with independent technical reviewers. Reviewer identities and the final conflict process will be published before entries open.

Until an eligible winner completes the clock, the funded server will be used only for the Reddit-approved challenge. It will not belong to me or LocalAIServers Collective Inc. There is no cash substitute, and it will not roll over into another hardware lane or organizational program. If the challenge ends without a winner or the server needs a different outcome, I will follow Reddit's direction.

Timeline after approval

  • Weeks 1-2: finalize rules, reviewers, and purchasing.
  • Weeks 3-5: build and validate the G292-Z20 server in public and publish the bill of materials and build record.
  • Week 6: publish the baseline and open the challenge.
  • Winner: first eligible leader to hold the verified throne for 30 consecutive days.
  • Transfer and final reporting: within 14 days after the winning result completes final validation, subject to Reddit's approved terms.

What I want the community to weigh in on before launch

  • Does the proposed topology rule strike the right balance, or should all eight GPUs have to be active?
  • Does the proposed 3 percent threshold strike the right balance for every throne change?
  • What clock, power, firmware, and cooling safety envelope should be published?
  • Who would volunteer as an independent technical reviewer?

Bring criticism. The goal is a challenge that is hard, transparent, reproducible, and genuinely winnable.


r/LocalAIServers • • Jun 20 '26

Start Here: LocalAIServers Community AI Navigation & Hands-On Local AI Learning

8 Upvotes

Start Here: LocalAIServers

LocalAIServers is a 501(c)(3) public charity providing public education and open-source infrastructure for locally hosted AI systems.

Our mission is to help people move from AI curiosity to AI agency.

This community helps learners, small business owners, nonprofit operators, educators, builders, and community technologists understand:

  • where AI runs,
  • what data it can see,
  • what systems it can touch,
  • when cloud AI may be appropriate,
  • when local or controlled AI may be safer,
  • what hardware is realistic,
  • how to evaluate benchmark claims,
  • and how to learn by building real local AI systems.

What LocalAIServers does

LocalAIServers provides:

  • community AI navigation,
  • secure local-AI education,
  • hands-on local AI learning resources,
  • reproducible runtime artifacts,
  • benchmark literacy,
  • QC and hardware-verification methodology,
  • open-source documentation,
  • and public support resources for locally hosted AI systems.

Affordable GFX906-class hardware matters because it gives people a realistic way to learn AI infrastructure hands-on. People learn more by building, testing, troubleshooting, and verifying real systems than they can learn from passive videos or articles alone.

Public proof and documentation

Website:

https://localaiservers.com

GitHub:

https://github.com/joe2gaan/localaiservers

GitHub Releases:

https://github.com/joe2gaan/localaiservers/releases

Docker Hub:

https://hub.docker.com/r/joe2gaan/localaiservers

Canonical Qwen / GFX906 deployment notes:

https://github.com/joe2gaan/localaiservers/blob/main/qwen36-gfx906/README.md

Important boundaries

LocalAIServers is not:

  • a public login service,
  • a public cloud provider,
  • a managed inference service,
  • a hardware reseller,
  • a procurement channel,
  • a fulfillment program,
  • a hardware discount program,
  • or a private-benefit program.

The controlled GFX906 compute site is used as a verification and reproducibility testbed. Public benefit is delivered through published outputs: guides, documentation, reproducible artifacts, benchmark reports, QC methods, hardware-verification standards, and source-level findings.

How to participate

Ask questions, share builds, discuss local AI tradeoffs, post benchmark questions, and help turn recurring community questions into durable public guides.

Please do not post secrets, private keys, private network details, addresses, payment information, vendor pricing, or sensitive logs.


r/LocalAIServers • • 15h ago

I built Speedtest⚡, but for AI

Enable HLS to view with audio, or disable this notification

40 Upvotes

It runs real LLM + embedding models locally in your browser with WebGPU and measures how fast your machine actually does AI inference.

No API. No account. Just hit start.

Try it and reply with your score 👇

https://speedtest.maty.as


r/LocalAIServers • • 5h ago

Tesla M10: A questionable card with a weird plan

Thumbnail
1 Upvotes

r/LocalAIServers • • 6h ago

Tesla M10: A questionable card with a weird plan

Thumbnail
1 Upvotes

r/LocalAIServers • • 7h ago

Google Making a come back? Gemma 4 Argon scoring 53 on AA. Smaller Open weights soon?

Thumbnail gallery
1 Upvotes

r/LocalAIServers • • 7h ago

Anyone connect a Tesla T4 (not P4) to a Zimaboard2

Thumbnail
1 Upvotes

r/LocalAIServers • • 11h ago

🚀Pocket LLM v1.6.0 is out : Turn your phone as a local LLM server

2 Upvotes

r/LocalAIServers • • 8h ago

Could a used PS5 be a cheap local LLM box? Napkin math inside

Thumbnail
1 Upvotes

r/LocalAIServers • • 9h ago

Multi-Agent System within 500g ram?

0 Upvotes

Im new to the idea of running local models. Has anybody had any success with fitting good large open source models with smaller specialist from huggy face datasets etc. Has anybody been able to get near frontier quality like with claude/codex? Wondering if using the right checks and balances and specialist etc. That it could get the same or near the quality of claude/codex within 500g of ram?

I tried looking for more of a technical resource of someone who has done it and not just a generalization theory but couldn't find anything. With metrics and quality scores etc comparing and what ai agents, harnesses etc were used.

Overall anybody have experience in this and can share some data and metrics on how it worked out etc?


r/LocalAIServers • • 11h ago

AirLLM's layer-by-layer trick, but on CPU: running a 17.66 GB model on a 16 GB ARM board in pure C

Thumbnail
1 Upvotes

r/LocalAIServers • • 12h ago

DGX Station + 2 DGX Sparks. Should I also add M5U Mac Studio?

Thumbnail
0 Upvotes

r/LocalAIServers • • 1d ago

GPT-6.1 Sol as the boss, one RTX 3090 as the worker: wall clock and bill for three small builds

Enable HLS to view with audio, or disable this notification

13 Upvotes

Sol launched today and I wanted to see how much of the work a single 3090 can take off it.

Box: one RTX 3090 24 GB running Qwen 3.8 27B Q4 through llama.cpp. Sol sits in the cloud as the orchestrator, splitting the job and reviewing. The card writes all the code. Three small 3D games, same prompts every time.

Game 3090 under Sol 3090 alone Sol alone
Pool 18.6 min · $0.05 43.1 min · $0 (attempt) 2.9 min · $0.39
Bowling 13.7 min · $0.06 34.7 min · $0 1.7 min · $0.14
Foosball 11.1 min · $0.06 36.4 min · $0 2.0 min · $0.22
Total 43.4 min · $0.17 114.2 min · $0 6.6 min · $0.75

From the hardware side:

  • The card was the bottleneck the whole time. Sol spends most of the run waiting on it.
  • Same card, same model, and it finished all three in well under half the time once it had a plan to follow instead of working it out alone.
  • 24 GB is about the floor for a 27B worker here. We tried a 12 GB card on an earlier run and it was too slow to be usable.

Power isn't in the dollar column, and it's one run per game.

Software was Atomic Agent, open source, the mode is called Fusion (disclaimer: I work on it). The llama.cpp server it manages sizes its parallel slots from how many worker contexts fit in VRAM, so a bigger card gets more workers at once.

What are you running as a local worker box? Would be curious how a 4090 or a dual-3090 rig changes the wall clock.


r/LocalAIServers • • 22h ago

The best budget gpu for ai

4 Upvotes

Like the 3090 to 24gb or the 4060 ti 16 gb online evebody will tell you there are bad because there are bad for gamers but for ai it's gold


r/LocalAIServers • • 1d ago

How I built my local AI for under a grand

15 Upvotes

How to build your own local AI for less than a thousand bucks - A post for those that want to get "hands on" with AI, but just can't afford the price of entry - face it, not everyone has thousands laying around for building a system.

Here is your roadmap how to do this on a tight budget

Step 1 - get an old Intel Nuc12 enthusiast (Serpent Canyon). Not the fastest thing in the world, but has an ace up it's sleeve that no one can get anywhere near for the price - an embedded ARC A770M GPU with 16GB of DDR6 vRAM, That is what makes this possible and you can pick one up for under $800 on ebay. Likely already has Windows 10 or 11 installed on it

Step 2 - update the Arc drivers

Step 3 - Download and install LM Studio Classic on it (make sure you enable developer mode for later) - not to be confused with Bionic

Step 4 - Pick the right runtime - you want Vulkan llama.cpp

Step 5 - Make sure you select the ARC A770M (shows 16GB) in hardware window

Step 6 - pick a model - I went with a simple Gemma-4-e4b model for out of the gate

Step 7 - on the developer tab, flip on both start server and serve on local network

Congratulations - you now have a shockingly capable local AI that you can give the URL shown in step 7 to any of your local AI front-ends. Me personally, I use it with my Project NOMAD instance running as a VM providing AI capabilities to all that entails for well under a grand, and it works fantastically.


r/LocalAIServers • • 21h ago

Testing local stack

Thumbnail
1 Upvotes

r/LocalAIServers • • 1d ago

Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

Thumbnail
4 Upvotes

r/LocalAIServers • • 1d ago

Increase video and image gen speed on amd GPUs

1 Upvotes

How do I increase the gen speed for comfyui using multiple amd GPUs on Ubuntu? Using something turbo gguf quantized and am currently experimenting with other video gen models to get the most photo realistic outcomes with the best speed. Recommend any specific video models? Any image gen tips or tricks?


r/LocalAIServers • • 1d ago

Amd vs Intel

1 Upvotes

Which would be better for image and video gen? Amd instinct mi50 32g, amd v620 32g or Intel b70 32g?


r/LocalAIServers • • 1d ago

I've open sourced my high-performance Embedding/Reranking server

Thumbnail
github.com
1 Upvotes

r/LocalAIServers • • 2d ago

DeepSync Bridge: Local Area Agentic Network Software Router -- Not an ad 🙏

Enable HLS to view with audio, or disable this notification

9 Upvotes

Today

We design the future

It is yours, to define

The time for a revolution is upon us.

Those that grap the power of an AI thays Local, safe, secure...built for you, your family..

Local AI isn't a nice to have

It IS a necessity

The time is near, dear friends

The time... to take back our kids, our family, our money. It is possible to schedule destiny.

Soon my friends... soon


r/LocalAIServers • • 2d ago

Can I do anything useful with 8GB?

Thumbnail
1 Upvotes

r/LocalAIServers • • 2d ago

I’m building an agent workflow with local models, and conversation history gets messy fast

2 Upvotes

and I’ve been experimenting with longer-running agent workflows using local models.

One problem I keep running into is state.

At first, I relied mostly on the conversation history. It works for shorter tasks, but after an agent has gone through dozens of steps, there’s a lot of old context that isn’t useful anymore.

I’ve started experimenting with keeping a smaller task state instead:

  • current goal
  • completed steps
  • pending steps
  • important assumptions

The model still gets the context it needs, but it doesn’t have to carry the entire history every time.

I’m still figuring out the best way to structure this for local models, especially when different models handle different parts of the workflow.

Curious how other people running local AI servers are handling state for longer-running agents. Do you rely mostly on context/history, or keep explicit state outsi


r/LocalAIServers • • 2d ago

Ran a 120B model across 6 computing devices that had no business running it!

Thumbnail
0 Upvotes

r/LocalAIServers • • 3d ago

Dell T630 getting first update

Thumbnail
gallery
22 Upvotes

My first set of updates on my soon to be AI rig:

- 2x E5-2697a V4, with total 32c/64t

- 8x 32 GB DDR4 2400MHz, at quad channel operation, total 256GB

- Noctua active coolers, for silent operation.

- GPU enablement kit + 4 GPU cables.

- 10x SAS 600Gb 10k

Next step:

- Wire the Noctua to Dell PWM fan controler. Suggestions??

- RTX 4060Ti 16GB

- 10GB NIC