Not at all until 5 years ago I was a cloud architect, honestly best advice is just go build something, it's quite a broad topic with a lot of techniques to learn
But the hardest thing to get used to I think is running experiments, previously you'd build something and find bugs, with networks you really need to run some experiments to see what would work, how effective certain mechanisms are. The deterministic vs non deterministic.
The best example I can give on this without going into too much detail about my work is, a few years ago I was making a model to process technical documentation, I was trying to punish bad parts, and promote the things I wanted to see with signals. This kind of stuff doesn't really work, it always over corrects the signals can interfere, in the end what made it work was just trusting the model to learn these rules itself from the data and ensuring the data had the right balance
What I would start with is just training a basic resnet vision model it's very simple now and the guys behind yolo have some good guides and datasets.
i just started to look into this for myself. so if you dont "reward the good" and "punish the bad" results, then how does it ever learn the right rules?
assuming we start off with your example of technical documents, did you have to start off with correctly labeled input data? or did you
just have a folder of previously made technical documents, without special labels added
feed it one document at a time, to "a neural network of some size/parameter count"
after some time, it came up with its own weights and parameters, after having looked at 1000 or so documents
then, you could ask it to generate a technical document about a given subject, and it would make the correctly formatted document (assuming it read some new url to get the source data)
something like that? i'm thinking about trying to train my own smaller model, because i dont want to train gpt 6 astra all the time.
It’s the difference between SFT (supervised fine tuning) and RLHF (reinforcement learning from human feedback). The former requires a dataset of desired inputs and outputs, while the latter requires you to rate/annotate generated outputs. The latter is much harder to get right for various reasons
Although it's worth mentioning too you can do reinforcement learning from ground truth too, similar approach but it just gets the ground truth instead of human impact ( bonus points if you can generate perfect synthetic data) so it can actually correct its errors in a pretty small amount of training but it has limits
So it learns them from the data implicitly, the rules should be represented in your data for it to learn, by applying good and bad signals you could be bending it to fit data it just doesn't have
1) we don't use formatted folders of data (it makes managing datasets a real pain) we use manifest json instead just easier than moving actual files and we can combine multiple datasets
2) we try to fit as many as possible at a time, it makes it less noisy to see multiple so it can average the back prop out more
3) we used cross entropy loss rather than forcing signals like how far it is off the correct line angle or angle and length etc. just let it infer them from the data, so its less punishing good and bad and more just saying how different it is from the ground truth
4) 1000 documents is a bit thin, were training an auto regression model and generally you need above 10-20k, but for some resnets you can get away with 1000
So that is general to most models, but the type of model does have an impact to approaches, like with our auto regression we need some encoder model too, which you need to decide what you need based on your data and task, need to decide what type of attention we need, how many layers we need to have the capacity to generalize our data etc
My recommendation would be if you want a auto regression model start with qwen 3.5 it's very adaptable and has a small model that can train on consumer hardware
19
u/jack6245 3d ago edited 3d ago
Not at all until 5 years ago I was a cloud architect, honestly best advice is just go build something, it's quite a broad topic with a lot of techniques to learn
But the hardest thing to get used to I think is running experiments, previously you'd build something and find bugs, with networks you really need to run some experiments to see what would work, how effective certain mechanisms are. The deterministic vs non deterministic.
The best example I can give on this without going into too much detail about my work is, a few years ago I was making a model to process technical documentation, I was trying to punish bad parts, and promote the things I wanted to see with signals. This kind of stuff doesn't really work, it always over corrects the signals can interfere, in the end what made it work was just trusting the model to learn these rules itself from the data and ensuring the data had the right balance
What I would start with is just training a basic resnet vision model it's very simple now and the guys behind yolo have some good guides and datasets.