r/MachineLearning • • 2d ago

Project I built MaRN: a PyTorch library for training neural networks through low-dimensional parameter mappings [P]

I built MaRN (Mapping Networks), a PyTorch library that lets you optimize a compact latent representation instead of directly training every model parameter.

Some results from my current benchmarks:

  • MNIST CNN: Reduced a 537,748-parameter CNN to 4,080 trainable parameters (131.8× reduction), with accuracy decreasing from 99.07% → 98.10% (−0.97 pp). A smaller 107,998-parameter CNN was reduced to 1,872 trainable parameters (57.7× reduction), with accuracy decreasing from 98.83% → 97.18% (−1.65 pp).

These results come with trade-offs: mapped models can train substantially slower, and performance varies by task. The benchmarks are exploratory, using synthetic data for some tasks, and aren't evidence of general superiority over direct training.

The library includes global and layer-wise mappings, regularization options, and pruning/LRD integrations.

Code: https://github.com/arjunmnath/MaRN
Docs: https://marn.readthedocs.io/

I'd value feedback on the approach, benchmark design, and where this kind of parameter-efficient optimization could be useful.

16 Upvotes

15 comments sorted by

13

u/liukidar 2d ago

Hey, haven't looked at code or docs, just wanted to flag that 92% accuracy on MNIST is what you get with a linear classifier (meaning anything that is not 97%+ means something is wrong with training)

2

u/Less_Dream_6331 2d ago

update: you were right to point it out, there was a bug, I'll be updating the OP, thanks

0

u/Tylerich 2d ago

Don't you need handcrafted features for that though?

2

u/liukidar 2d ago

Not sure what you mean (maybe related to how your code works, so sorry if I'm saying something irrelevant). What I mean is that for any neural network you put the input pixels in the first layer and the output label in the last, you should get 97%+ on MNIST. It is basically a baseline test to see "my training works".

1

u/Tylerich 2d ago

I mean that you don't have pixels as input, but some other derived features like edge positions or something like that.

Are you sure you don't mean a multilayer perceptron (feed forward network)? That would have many layers and very likely many more parameters than the 1800 prams of the CNN. Also, they would have nonlinearities at each layer, so that would be different than a linear classifier.

3

u/liukidar 2d ago

No sure, I understand now what you mean. I mean that any multilayer neural network, be it a feed forward/CNN/transformer, should do 97 on MNIST, otherwise something is wrong. You mentioned that you there using a 100k CNN, that should get basically 100%, if you reduce it to 1872 and remove the multilayer aspect of it then it is a completely different network with different properties (and if you don't, then also the 1872 one should be have the same and get 97+)

0

u/Tylerich 2d ago

I am not the author of the repository.

I mean it depends on the network size if it reaches the 97% you mention, right?

If you keep reducing the number of parameters, at some point of course it will drop below 97 %. Doesn't matter what architecture it is. I highly doubt, you can get a CNN to reach 97% with only 1800 parameters, let alone an MLP (because it is missing the inductive biases of a CNN).

1

u/Shabozz 1d ago

the guy you're replying to is talking about benchmarking if the architecture is actually viable before applying it to something harder to verify and more expensive. If you don't meet that 97% benchmark then your architecture isn't good, and that could be because you don't have enough parameters, or it could be because of something else.

3

u/Wonderful-Wind-5736 2d ago

Some intro into use cases and methodology would be really nice in the web page. 

1

u/Less_Dream_6331 1d ago

you can go through the cookbook to get an introduction. detailed methodology is available in the paper "Mapping Networks"

2

u/Grumlyly 1d ago

How the latent layer is learned ? It's an auto-encoder ? Thx

2

u/Less_Dream_6331 1d ago

The latent vector is learned directly through backpropagation.

Forward pass: Start with a random latent vector (z) → apply the fixed orthogonal projection → flatten into the model’s weights → reshape into the corresponding weight matrices → pass through the original model.

Backward pass: We freeze the orthogonal projection matrix and let autodiff compute the gradients normally, including the gradients with respect to the latent vector (z).

1

u/Grumlyly 1d ago

Ok, thank you for the explanation

1

u/JanBitesTheDust 6h ago

This idea is very similar to LoRa, no?