Day 34

I learned from Andrej Karpathy’s State of GPT talk.
Pretraining → Fine-tuning → Reward Modeling → Reinforcement Learning.
There’s a whole process behind turning a model that learns language patterns into something that can follow instructions and interact with humans.
Day 34 of 1% better.
LinkedIn post: https://lnkd.in/p/gtfKSTHy (opens in a new tab)