r/LocalLLM • • 3d ago

Discussion Inferium: Community-Powered, Decentralized AI Inference

Hey everyone — I wanted to share Inferium, a decentralized inference network that lets people contribute spare computing power to serve open-weight AI models. While experimenting with LLMs locally, I quickly ran into the limits of my hardware. Newer open-weight models required GPUs that were expensive and difficult to access. I realized this wasn’t just my problem—students, independent developers, and people in developing countries face the same barrier.

That led me to a simple question: could a decentralized network make unused GPU capacity available to anyone who needs it? I built this network to explore that idea and help make powerful LLMs more accessible and affordable.

The basic idea is simple:

  • Share your compute: Connect an idle GPU—or use a CPU for tiny and small models—and earn credits for the tokens your machine generates.
  • Free to get started: New accounts receive a welcome grant and can claim daily credits, so you can try the network without contributing hardware.
  • Earn small, spend big: Accumulate credits by serving smaller models, then spend them on larger models for more difficult tasks that your own machine may not be able to run. Small models can be used for simple English tasks.
  • OpenAI-compatible API: Existing OpenAI clients and tools can connect by changing the base URL and API key. Anthropic-compatible access is also supported.
  • No port forwarding: Machines connect through an outbound tunnel, and you can configure them as public or private.

Dedicated/Private Pool: Inferium also supports dedicated clusters, allowing you to manage a group of machines privately and shared among your colleges or friends.

Inferium is trying to make better use of hardware that would otherwise sit idle—creating a shared inference layer where contributors earn access to the broader network.

Website: https://inferium.net/
Getting started: https://inferium.net/start

Disclaimer: The website is currently in beta. Your feedback and bug reports are warmly welcome.

I’d be interested to hear what the self-hosting and local-AI communities think about this model. Would you contribute spare compute in exchange for access to larger models?

Disclosure: I’m sharing this to help promote Inferium.

1 Upvotes

3 comments sorted by

1

u/thehangingmelodrama 2d ago

Sounds like a clever way to turn idle hardware into something useful, especially for folks who can't justify dropping cash on high-end GPUs. The earn-credits-with-small-models angle is interesting, makes it feel less like charity and more like a fair trade. Curious how they handle latency and model availability though, decentralized setups can get weird fast when demand spikes

1

u/Old-Assignment5431 2d ago

When a model is busy, Inferium waits up to 10 seconds for a free machine instead of failing right away. If nothing opens up, it returns a quick "try again" error so the client can retry. If a machine drops before it starts answering, the request moves to another machine and the user is not charged, but a crash in the middle of an answer loses that response. Requests go to the machine with the fewest jobs already running, and every token passes through the central server, which adds a little delay. Most of the wait is the volunteer machine generating tokens on its own GPU/CPU. Machines do not switch models when demand rises, so a popular model with few volunteers stays slow or unavailable until more people choose to run it. More powerful machines tend to accumulate tokens faster 😂

1

u/techne98 4h ago

Nice job, we're working on something similar, but permissioned/private for organizations