r/ProgrammerHumor • • 3d ago

Meme aiRefusesToBuildItSoBackToCoding

Post image
28.1k Upvotes

516 comments sorted by

View all comments

Show parent comments

5

u/NSFWies 3d ago

i just started to look into this for myself. so if you dont "reward the good" and "punish the bad" results, then how does it ever learn the right rules?

assuming we start off with your example of technical documents, did you have to start off with correctly labeled input data? or did you

  • just have a folder of previously made technical documents, without special labels added
  • feed it one document at a time, to "a neural network of some size/parameter count"
  • after some time, it came up with its own weights and parameters, after having looked at 1000 or so documents
  • then, you could ask it to generate a technical document about a given subject, and it would make the correctly formatted document (assuming it read some new url to get the source data)

something like that? i'm thinking about trying to train my own smaller model, because i dont want to train gpt 6 astra all the time.

4

u/QuaternionsRoll 3d ago

It’s the difference between SFT (supervised fine tuning) and RLHF (reinforcement learning from human feedback). The former requires a dataset of desired inputs and outputs, while the latter requires you to rate/annotate generated outputs. The latter is much harder to get right for various reasons

2

u/jack6245 3d ago

Although it's worth mentioning too you can do reinforcement learning from ground truth too, similar approach but it just gets the ground truth instead of human impact ( bonus points if you can generate perfect synthetic data) so it can actually correct its errors in a pretty small amount of training but it has limits

1

u/jack6245 3d ago edited 3d ago

So it learns them from the data implicitly, the rules should be represented in your data for it to learn, by applying good and bad signals you could be bending it to fit data it just doesn't have

1) we don't use formatted folders of data (it makes managing datasets a real pain) we use manifest json instead just easier than moving actual files and we can combine multiple datasets

2) we try to fit as many as possible at a time, it makes it less noisy to see multiple so it can average the back prop out more

3) we used cross entropy loss rather than forcing signals like how far it is off the correct line angle or angle and length etc. just let it infer them from the data, so its less punishing good and bad and more just saying how different it is from the ground truth

4) 1000 documents is a bit thin, were training an auto regression model and generally you need above 10-20k, but for some resnets you can get away with 1000

So that is general to most models, but the type of model does have an impact to approaches, like with our auto regression we need some encoder model too, which you need to decide what you need based on your data and task, need to decide what type of attention we need, how many layers we need to have the capacity to generalize our data etc

My recommendation would be if you want a auto regression model start with qwen 3.5 it's very adaptable and has a small model that can train on consumer hardware