From Completion to Conversation
Chat Template + SFT: Large Models Finally Learn to "Talk"
The gap between a completion machine and a conversational assistant is this training paradigm.
OpenAI agreed on a format, then trained the model specifically on that format, and large models could genuinely hold a conversation
The Evolution from Completion to Conversation
1
Standardize the Conversation Format (Chat Template)
Borrowing from the Jinja template language, define special tokens: <|im_start|> / <|im_end|> to wrap each message
2
Assemble All Messages in the Format
Three roles — system / user / assistant — concatenated into a single flat string and sent to the model
3
SFT (Supervised Fine-Tuning)
Continue training on massive "formatted conversation + high-quality answer" data. The model learns to act as an assistant within this format rather than just completing text.
4
Scale Up
Bigger models, more data: GPT-3 → ChatGPT → GPT-4 — exponential improvement in capability
Chat Template vs. SFT Comparison
<|im_start|>system← system prompt role marker
You are a helpful AI assistant.← System Prompt content
<|im_end|>← end marker
<|im_start|>user← user message begins Who is Zixia Fairy? (紫霞仙子是谁?) <|im_end|>
<|im_start|>assistant← model begins completion here (model completion area)← SFT trains the model how to output here <|im_end|>
<|im_start|>user← user message begins Who is Zixia Fairy? (紫霞仙子是谁?) <|im_end|>
<|im_start|>assistant← model begins completion here (model completion area)← SFT trains the model how to output here <|im_end|>
Before SFT (Base Model)
User: Who is Zixia Fairy? (紫霞仙子是谁?)
紫霞仙子是哥哥,哥哥你喜欢我,你说你喜欢我…
(Continues in the style of training data — nothing like answering a question; Chinese example preserved intentionally)
(Continues in the style of training data — nothing like answering a question; Chinese example preserved intentionally)
After SFT (Chat Model)
User: Who is Zixia Fairy? (紫霞仙子是谁?)
Zixia Fairy (紫霞仙子) is a character from the film A Chinese Odyssey, played by Athena Chu. She is Supreme Treasure's destined love, who sacrificed everything for him.
(Understands what answering means; responds as an assistant)
(Understands what answering means; responds as an assistant)
SFT (Supervised Fine-Tuning) is still Token prediction at its core — the only change is that the training data becomes "formatted conversations + high-quality answers".
The model learns to produce a proper response after <|im_start|>assistant, instead of just continuing the training corpus.
The model learns to produce a proper response after <|im_start|>assistant, instead of just continuing the training corpus.
At this moment, the "completion machine" evolved into a "conversational assistant"