Krina Vaghela
← Blog

Day 14

Picture from the day 14 LinkedIn post

Day 14 of 1% Better.

Today, I learned about Training, Pre-training, Fine-tuning, and Post-training.

Here's the simplest way to think about it:

1. 𝐓𝐫𝐚𝐢𝐧𝐢𝐧𝐠
↳ The process of updating a model's weights so it can learn or improve.

2. 𝐏𝐫𝐞-𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠
↳ The model starts from scratch with random weights.
↳ It learns language, patterns, and general knowledge from massive amounts of data.
↳ This is the most compute-intensive stage of building an LLM.(a lot)

3. 𝐅𝐢𝐧𝐞-𝐭𝐮𝐧𝐢𝐧𝐠
↳ A pre-trained model is trained on a specific dataset.
↳ It adapts the model for a particular domain/task.
↳ Since the model already has general knowledge, it requires much less compute.

4. 𝐏𝐨𝐬𝐭-𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠
↳ This stage improves how the model behaves.
↳ It helps the model follow instructions, align with human preferences, and generate safer responses.
↳ Most production AI models go through post-training before they're released.

LinkedIn post: https://lnkd.in/p/gJ-8pBRF (opens in a new tab)