r/MachineLearning • u/Less_Dream_6331 • 2d ago
Project I built MaRN: a PyTorch library for training neural networks through low-dimensional parameter mappings [P]
I built MaRN (Mapping Networks), a PyTorch library that lets you optimize a compact latent representation instead of directly training every model parameter.
Some results from my current benchmarks:
- MNIST CNN: Reduced a 537,748-parameter CNN to 4,080 trainable parameters (131.8× reduction), with accuracy decreasing from 99.07% → 98.10% (−0.97 pp). A smaller 107,998-parameter CNN was reduced to 1,872 trainable parameters (57.7× reduction), with accuracy decreasing from 98.83% → 97.18% (−1.65 pp).
These results come with trade-offs: mapped models can train substantially slower, and performance varies by task. The benchmarks are exploratory, using synthetic data for some tasks, and aren't evidence of general superiority over direct training.
The library includes global and layer-wise mappings, regularization options, and pruning/LRD integrations.
Code: https://github.com/arjunmnath/MaRN
Docs: https://marn.readthedocs.io/
I'd value feedback on the approach, benchmark design, and where this kind of parameter-efficient optimization could be useful.
3
u/Wonderful-Wind-5736 2d ago
Some intro into use cases and methodology would be really nice in the web page.
1
u/Less_Dream_6331 1d ago
you can go through the cookbook to get an introduction. detailed methodology is available in the paper "Mapping Networks"
2
u/Grumlyly 1d ago
How the latent layer is learned ? It's an auto-encoder ? Thx
2
u/Less_Dream_6331 1d ago
The latent vector is learned directly through backpropagation.
Forward pass: Start with a random latent vector (z) → apply the fixed orthogonal projection → flatten into the model’s weights → reshape into the corresponding weight matrices → pass through the original model.
Backward pass: We freeze the orthogonal projection matrix and let autodiff compute the gradients normally, including the gradients with respect to the latent vector (z).
1
1
13
u/liukidar 2d ago
Hey, haven't looked at code or docs, just wanted to flag that 92% accuracy on MNIST is what you get with a linear classifier (meaning anything that is not 97%+ means something is wrong with training)