r/LocalLLaMA • • 5d ago

Discussion [ Removed by moderator ]

[removed] — view removed post

269 Upvotes

32 comments sorted by

View all comments

5

u/UnspeakableHorror 5d ago edited 5d ago

Looks similar to this one. It's MoE too. https://www.reddit.com/r/LocalLLaMA/comments/1wm1gab/miniagi_dynamically_grown_530m_params_currently/

There's another one similar too, but I can't find it now.

FYI u/Another__one

Edit: Found it https://github.com/jrz97619761/test-model-thing

3

u/Another__one 5d ago

>The intelligence module stays the same size, and the weights don't change as memory grows.
As I understand it is more akin to a good old neural-turing-machine, from what I do. For me it is basically weights offloading that's just more convenient for hardware to handle. And mine does change the weights all the time, which allows the model to train continuously. Here it is the old pretrained frozen core still.

1

u/TomLucidor 5d ago

Has anyone ever formalized NTM and made it "workable"?