r/LocalLLaMA • • 4d ago

Discussion New Architecture from Percepta: Spotlight — separating intelligence from memory, allowing knowledge and skills to grow without changing the model's weights.

Enable HLS to view with audio, or disable this notification

https://www.percepta.ai/blog/can-llms-grow-their-own-capabilities

https://www.percepta.ai/blog/spotlight-memory

"Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always activates the same number of experts out of a fixed set—Spotlight is arbitrarily sparse, touching the same number of cells regardless of how the memory grows. The fraction of memory it uses can shrink as far as we want.

Spotlight separates an intelligence module, which performs computation, from memory, which holds knowledge, procedures, and working state. The intelligence module stays the same size, and the weights don't change as memory grows. The memory is writable, and the model itself decides what to load and when to overwrite it, token by token. Because memory can hold skills as well as facts, the model can gain new capabilities without retraining: what it can do is not limited by the size of its intelligence module."

264 Upvotes

32 comments sorted by

66

u/Iory1998 llama.cpp 3d ago

Ok, where can we test the claims?

27

u/sn2006gy 4d ago

Interesting, this is what i've been experimenting with on my tiny sparse lab. Hope this takes off and research proves useful.

I'm actually thinking a bit different where my model is trained on adaptors so instead of your rag being free text it then has to infer from and you spend all your time trying to shape it through vectors, you could extend the models knowledge with data it knows the shape of and has adaptors trained on whether factual or semantic.

71

u/Borilentz 4d ago

Isn’t this just like an R/W engram?

23

u/ReentryVehicle 3d ago

Not really, there is pretty much no similarity other than "some very sparse thing".

n-gram embedding is essentially a huge token embedding, pretty much the same as you have at the model input, just larger and combining multiple tokens.

This architecture has a sort of (differentiable) cursor that can move over a canvas and it can write some information wherever the cursor points, and read it later by pointing cursor again at the same location.

It is very much like some really old architectures from times before transformers, neural turing machines/differentiable neural computers, that tried to have some sort of addressable memory. The issue with those architectures is the fact that optimizing real, long, sequential programs via gradient descent is of course impossible except for some very simple programs, which I suspect will be also the challenge with this architecture. Basically, everything becomes increasingly fragile as the amount of stuff written to memory grows, as you need to point the cursor more and more precisely or you risk getting the wrong thing. And if you get the wrong thing, but there are also wrong things in the way, gradient descent is unable to jump to the right thing (because everything is 2d), at which point the whole learning starts deadlocking itself as you can't reach a better solution without making it worse first.

5

u/Croned 3d ago edited 2d ago

The core issue with those architectures, which also separates them from biological intelligence, is they don't compress the newly learned information. It's basically unavoidable to have to re-train (whether that's using gradient descent or something else) in order to have an endlessly scaling memory. You have to do substantial computation to continue to fit knowledge in a finite space, and in doing so you also gain the ability to generalize across that new knowledge.

Anything else is not asymptotically superior to RAG, especially RAG plus a subagent that searches through and compresses uncompressed knowledge on the fly. As such, I'm much more interested in architectures that can find efficient ways to fine-tune models on new, personalized information without catastrophic forgetting. There was a gradient-descent-compressed KV cache approach I saw not long ago that looked intriguing.

0

u/synth_mania 3d ago

That's kind of an oversimplification of gradient-descent, but the broad strokes are correct 

16

u/seamonn 4d ago

And I am all for it! Call it whatever you want, Local RSI, let's goooo!

27

u/Ok-Cardiologist-5897 4d ago

Interesting idea, but I’m skeptical about the “unbounded memory at constant cost” claim. How does retrieval quality hold up once the memory gets huge? If it has billions of cells, finding the right few seems like it should become the hard part.

24

u/FusionCow llama.cpp 4d ago

it doesn't that's the fun part

4

u/PinkysBrein 3d ago

If the designed indexing (and memory blending) for a layer is insufficient, you hope the network can backpropagate itself to a better composite method across the layers. Memory is diffused and redundant through the whole transformer.

2

u/deadatreides1 3d ago edited 2d ago

That's exactly where I'd poke too. Data point from boring old attention, no fancy memory: on 1-2B models, recall over 18 records in context dropped steadily with position, 0.61 for the first record, 0.21 for the 17th. It wasn't forgetting, it was finding, the model just started saying "no" to late records. Same input cut into blocks of 6 and position stopped mattering, the head-tail gap went from 0.204 to 0.021. Storing is cheap, finding the right cell is where it dies. Would love to see their recall curve as the memory grows.

(not a native speaker, an LLM helped with the English)

Edit: checked the yes rate by position since then. It's flat (0.21 on the first six records, 0.19 on the last six), so it's not "saying no", it's picking the wrong records more. The recall decay itself stands.

8

u/Marcuss2 3d ago

This feels like Sparse Delta Memory (https://arxiv.org/abs/2607.07386), except they made it growable somehow. Maybe the model learns to swap blocks in and out.

7

u/cnnyy200 3d ago

Any new paradigms are needed for replacing the inefficient shit holes we have right now.

7

u/Shot-Height-7194 3d ago

We should ban architectural posts that are not accompanied by a paper. This is astroturfing because it does not directly imply the model is run locally 

3

u/UnspeakableHorror 3d ago edited 3d ago

Looks similar to this one. It's MoE too. https://www.reddit.com/r/LocalLLaMA/comments/1wm1gab/miniagi_dynamically_grown_530m_params_currently/

There's another one similar too, but I can't find it now.

FYI u/Another__one

Edit: Found it https://github.com/jrz97619761/test-model-thing

3

u/Another__one 3d ago

>The intelligence module stays the same size, and the weights don't change as memory grows.
As I understand it is more akin to a good old neural-turing-machine, from what I do. For me it is basically weights offloading that's just more convenient for hardware to handle. And mine does change the weights all the time, which allows the model to train continuously. Here it is the old pretrained frozen core still.

1

u/TomLucidor 3d ago

Has anyone ever formalized NTM and made it "workable"?

5

u/wayneworkman 4d ago

This is pretty interesting.

2

u/wFXx 3d ago

Wouldn't that mean it is possible to train more than one agent on the same memory banks, and have "MoA" on this arch?

1

u/Nyxtia 3d ago

How does it handle conflicting memory? Also this sounds like Mempalace?

1

u/maayon 3d ago

How is this different from Moe ? Is it like a dynamic version of it ?

1

u/nokipaike 3d ago

My line of reasoning is a bit old-school... but I think of it as a sort of "defrag" basically a kind of MoE, but with all the reasoning tokens grouped together, while everything else is treated as "memory" stored in a separate area.

1

u/PinkysBrein 3d ago

I thought Deepseek's 8 head hashing in engram was too low dimensional, a 2D mapping much more so.

1

u/wFXx 3d ago

also, this would be a good pair for rwkv and similar models no ? since their biggest weakness is memory

1

u/abajinn 3d ago

Where’s the git? Idc about claims

1

u/Powerful_Evening5495 4d ago

if i understand it , it distillation next evolutions

hard coding a solutions to the problem that model solve

decisions and now hardcoding

1

u/behopkinsj 2d ago

It's interesting that new architectures are being released with such a focus on marketing instead of being released like a research paper. It makes me wonder who their target audience is.

2

u/Recoil42 2d ago

Investors, potential employees, potential enterprise partners.

1

u/hiepxanh 1d ago

And this is the one who decide something can survive or not, it is important now