Full-Vocabulary, Top-k, and Sampled Token Distillation
This post briefly discusses the difference between full-vocabulary, top-k, and sampled-token distillation.
Table of Contents
1. Distillation Objective
In On-Policy Distillation or any other distillation/RL methods, the objective to optimize usually involves a KL divergence between the policy model \(\pi_{\theta}\) and the reference model \(\pi^{\ast}\) on next-token distribution, given a prefix sequence of tokens, i.e.,
\[ \mathrm{KL}\left( \pi_{\theta}(\cdot | x_{\lt t}) \parallel \pi^{\ast}(\cdot | x_{\lt t}) \right) \]
Or possibly the reverse KL.
2. Full-Vocabulary 全词表
The full-vocabulary approach is just the definition of KL divergence. In short, it enumerates all the tokens in the vocabulary \(\mathcal{V}\), and then computes the probability.
\[ \mathrm{KL}_{\text{FV}}=\sum_{v \in \mathcal{V}} \pi_{\theta}(x_{t}=v | x_{\lt t}) \log \frac{\pi_{\theta}(x_{t}=v|x_{\lt t})}{\pi^{\ast}(x_{t}=v|x_{\lt t})} \]
3. Top-k
The drawback of full-vocabulary approach is straight-forward: if the model has a very large vocabulary (\(\gt 150,000\) for modern LLMs), the computational cost and communication cost for transferring tokens and corresponding probabilities will grow quickly in a distributed training system. Therefore, we may only sample the top-k tokens of the reference model (or teacher model in the context of OPD) to approximate the KL disvergence:
\[ \mathrm{KL}_{\text{Top-K}} = \sum_{v \in \texttt{TopK}_{\text{ref}}(\mathcal{V})} \pi_{\theta}(x_{t}=v | x_{\lt t}) \log \frac{\pi_{\theta}(x_{t}=v|x_{\lt t})}{\pi^{\ast}(x_{t}=v | x_{\lt t})} \]
Note here, since the teacher only samples \(k\) tokens, the teacher will have to softmax them to make sure the probabilities sum up to \(1\).
4. Sampled-Token
In sampled-token approach, we even use less tokens to approximate the KL divergence, i.e., only the token sampled by the student.
\[ \mathrm{KL}_{\text{ST}} = \pi_{\theta}(x_{t}=x_{\text{student}} | x_{\lt t}) \log \frac{\pi_{\theta}(x_{t}=x_{\text{student}} | x_{\lt t})}{\pi^{\ast}(x_{t}=x_{\text{student}} | x_{\lt t})} \]