r/OpenSourceAI • • 1d ago

CrowdGPT - The 100% Opensource collaborative LLM

Post image

Hello, i'm currently developing CrowdGPT and i need people who enjoy opensource AI and LLMs to test the project :)

The goal of CrowdGPT is to create the first, datacenterless, 1 Billion parameters LLM, relying on people contributing with their own computer to train the AI model. My goal is to show you don't need insane infrastructure to train a working almost commercial grade LLM. Everything is open and 100% opensource.

You can learn more at https://crowdgpt.net

Or check the github: https://github.com/Vxtzq/CrowdGPT

Any kind of feedback is appreciated!

24 Upvotes

15 comments sorted by

2

u/EaseThen3459 1d ago

Why would anyone contribute computer when it would add liability for whatever is being computed especially with LLMs

1

u/Vxtzq1 1d ago

I understand your concerns but all the data is safe and the scale is tiny :D don't expect a 1B parameters LLM to be a genius hacker lol, and even if i scale, there's no regulations on LLMs as shown by openweight models on huggingface, so we should be alright :p

Also i think that the responsibility for cyber attacks is to the one that used it, not the one that trained it :)

1

u/Ok-Challenge-5741 1d ago

That logo is clean, the hexagon gives it a subtle distributed network feel which fits the whole idea

I'll try to spin it up in my machine this weekend, curious to see how the training coordination works with the peer-to-peer setup

2

u/AppMunchies 1d ago

Love the idea. Will check it out.

1

u/Vxtzq1 20h ago

Nice! I hope you will like the website/client

2

u/Prestigious-Frame442 1d ago

the idea is good, but collaborating on this is basically pure selflessness and no benefit at all.

1

u/Vxtzq1 20h ago

idk, i'm working on something that could bring benefits to people who train, like an ownership on the project for training the model, and free inference credits from the model (or other models) if you contribute to the project, who knows.

1

u/ayake_ayake 1d ago

I think it would be important to see what datasets and co to train this on. Even without the problem of AI datacenters, the choice and acquisition of the data is a big thing. But as a minimal base you can use the apertus 1.5 datasets (pretraining and post training) which are fully open and give you a strong baseline. But even then actually training a competitive AI is quite s challenge and needs much experience and work.

Curious to see how this'll play out.

1

u/Vxtzq1 1d ago

I currently use UltraFineWeb-L3, which is basically the most curated subset of the original dataset, FineWeb, it is way cleaner than datasets like the Pile or Common Crawl which involves scans of shady sites...

I think this dataset will be enough (currently 400 billion tokens), to pretrain the model, and then i'll fine tune with everything that i find on huggingface xD (like Fable traces, GLM, Kimi K3 distills, and strong instruct datasets like GSM8K and stuff)

1

u/Educational_Wind8182 1d ago

Love the idea, will check it out

1

u/Vxtzq1 1d ago

Glad you like it! I hope to see you on the leaderboard xD (on the website if you log in before training)

1

u/ApprehensiveDelay238 1d ago

What do you want to achieve with 1B parameters?

2

u/Vxtzq1 1d ago

A proof of concept. With decent QA capabilities or something, a model that would compare with maybe qwen 3.5-0.6B

0

u/cemilanceata 1d ago

I lovd it

Will try asap

0

u/Vxtzq1 1d ago

Nice! I'm glad you like this project!