Knowledge Distillation: Distilling the Knowledge in a Neural Network
A big model's probabilities for the wrong answers teach a small model how to generalize.
Hinton, Vinyals and Dean (2015) raise the temperature of a large model's softmax until its small probabilities are big enough to train on, then fit a small model to those softened outputs. The small model keeps most of the large model's accuracy and can learn from a fraction of the data.
Explaining the paperDistilling the Knowledge in a Neural NetworkOne loss function with a temperature in it, tested on MNIST digits, on the acoustic model behind Android voice search, and on a 15,000-class image set inside Google.
The surest way to gain a point or two of accuracy is to train several models and average their predictions. The cost arrives at deployment: an ensemble of ten networks needs ten forward passes and ten sets of weights for every query, and a service that answers phone requests in real time cannot pay that. A single very large network has the same problem in a different form.
Rich Caruana and his collaborators had shown that the knowledge in an ensemble can be compressed into one small network that is much easier to deploy.2 This paper, by Geoffrey Hinton, Oriol Vinyals and Jeff Dean at Google, gives that idea a general training rule they call distillation: run the large model (the teacher) with a raised softmax temperature, and train the small model (the student) to reproduce the full probability vector it outputs, not only its top answer.1
The page follows the paper's order: what information a trained classifier holds beyond its answer, the temperature that exposes it, the loss and its gradient, the result that matching logits is a limiting case, and then the experiments, from MNIST to a production speech recognizer to a 15,000-class image problem solved with specialist models.
The cumbersome model and the deployment gap
An ensemble beats its members because each member's errors depend partly on accidents of its training run, such as its random initialization and the order it saw the data. Averaging several members cancels part of those private errors and keeps what they agree on. The paper's speech ensemble is ten copies of one architecture, trained the same way from different random initializations, and the average still beats any single copy.
The paper calls the expensive model the cumbersome model: either a literal ensemble, or one large network trained with a strong regularizer such as dropout. Dropout randomly deletes units during training, which trains an exponentially large family of sub-networks that share weights, and at test time runs the full network once with scaled weights as an approximation to averaging them all.5 The MNIST teacher later on this page is of that kind.
The paper opens with insects, which have a larval form built for feeding and an adult form built for travel and reproduction. Training and deployment have different requirements in the same way. Training can be slow, parallel and enormous, because it runs once; deployment must answer each query quickly and cheaply, many millions of times.
What a classifier knows beyond its top answer
If you equate a model's knowledge with its weights, it is hard to see how a network with different weights, or a different architecture, could receive it. The paper proposes a different view: the knowledge is the learned mapping from inputs to outputs. And the output of a classifier is a full probability distribution over classes. A good model shown a BMW puts almost all its probability on "BMW", but the small remainder is ordered: a garbage truck gets a very small probability and a carrot a far smaller one. That ordering records which classes the model finds similar, and it determines how the model generalizes to inputs it has not seen.
A one-hot training label (probability 1 on the correct class, 0 on all others) carries none of that ordering. Distillation replaces it with a soft target: the teacher's whole output vector for the same input, which the student is trained to reproduce.
For a confident teacher, the useful part of the soft target is very small numbers. On MNIST, one image of a 2 might get probability of being a 3 and of being a 7, and a different 2 the reverse. The ratio says which 2s look like 3s and which look like 7s. In the cross-entropy loss, the gradient on a student logit is , the student's probability minus the target's, so with a target of the upward gradient on the student's class-3 logit is at most per example, so the similarity information in the targets barely reaches the student.
The softmax has a parameter that changes this. A network turns its logits (the unbounded scores it computes for each class) into probabilities by exponentiating and normalizing. Dividing every logit by a temperature first gives equation (1); the name comes from statistical physics, where a hotter system spreads its probability over more states:
At this is the ordinary softmax. A larger shrinks the differences between logits before they are exponentiated, so the distribution gets softer (its entropy rises) and, as , every class approaches for classes. A smaller sharpens it, and as the output becomes one-hot on the top class. The ratio of any two probabilities is , so raising compresses every ratio toward 1 while keeping the order of the classes.
A worked example with the logits used in Figure 1. A teacher looking at one image of a 2 outputs these logits for the digits 0 to 9:
The 3 and the 7 have the next-highest logits after the 2; the 4 has the lowest. The probabilities at three temperatures:
| temperature | p(2) | p(3) | p(7) | p(8) | p(4) | entropy |
|---|---|---|---|---|---|---|
| T = 1 | 95.9% | 2.9% | 1.07% | 0.059% | 0.005% | 0.29 bits |
| T = 4 | 39.6% | 16.5% | 12.9% | 6.2% | 3.4% | 2.73 bits |
| T = 20 | 14.0% | 11.8% | 11.2% | 9.7% | 8.6% | 3.30 bits |
The logit gap between the 3 and the 4 is 6.3, so their probability ratio is at , at , and 1.4 at . At the classes other than the 2 hold 60% of the probability, against 4.1% at , and the maximum entropy for ten classes, reached by the uniform distribution, is , or 3.32 bits. In Figure 1, watch the 3 and the 7 at a middle temperature: they rise well above the dissimilar digits before everything flattens toward 10% at the top of the range.
On the log scale the bars at already have the same order as at : the temperature does not create the similarity structure, it raises the small probabilities to a size where the loss responds to them. The structure comes from a teacher that learned the data. This separates soft targets from label smoothing, a different regularizer that replaces the one-hot label with, for example, 0.9 on the correct class and an equal share of the remaining 0.1 on every other class, the same for every input. Soft targets differ from image to image. The popular name for the information in the wrong-class probabilities, "dark knowledge", comes from a later Hinton talk; the paper calls it a "rich similarity structure over the data".4
The paper gives two reasons this helps the student. A hard label is one class out of , at most bits per example (3.3 bits for ten digits); a high-entropy soft target describes the teacher's whole view of how the input relates to every class. And soft targets have "much less variance in the gradient between training cases": two similar images of a 2 get similar soft targets, so they pull the student's logits in similar directions. With more information per example and less noise between examples, the student can often train on much less data than the teacher needed, and at a much higher learning rate.
How knowledge distillation works
The training procedure has four parts: a trained teacher, a transfer set of inputs, a temperature, and a loss. The transfer set is what the student trains on. It can be unlabeled data, since the teacher provides the targets, or the original training set; the paper found the original training set works well. For each transfer input the frozen teacher produces logits , which become soft targets through equation (1) at temperature . When the teacher is an ensemble, the soft target is the arithmetic or geometric mean of the members' distributions. The student computes its own logits and its own softened probabilities at the same temperature, and the loss is the cross-entropy between the two distributions. Because the teacher's is fixed, this cross-entropy equals the KL divergence plus the teacher's entropy , a constant the student cannot change, so minimizing either one gives the same student:
The high temperature is used only during training. Once trained, the student runs at like any other classifier.
When the true labels are known for some or all of the transfer set, the paper adds them. It tried folding the labels into the soft targets and found it worked better to keep two objectives and take a weighted average: the soft-target cross-entropy above, and an ordinary cross-entropy with the true labels, computed from the same student logits at . The best results came from a "considerably lower" weight on the hard-label term. The paper's reasoning: the student usually cannot match the soft targets exactly, and when it must miss them, missing in the direction of the correct answer helps.
Mixing the two terms needs one correction. The gradient of the soft-target term shrinks as when the temperature goes up (the next section derives this), while the hard term is always computed at . Without a correction, every increase of during tuning would also shift the balance toward the hard labels. The paper multiplies the soft-target loss by , which keeps the relative contribution of the two terms roughly fixed as changes. A training step, with illustrative MNIST sizes:
# one distillation step (illustrative MNIST sizes: B = 128 images, N = 10)
with no_grad():
v = teacher(x) # teacher logits [B, N], frozen
z = student(x) # student logits [B, N]
p = softmax(v / T, dim=1) # soft targets at temperature T
log_q = log_softmax(z / T, dim=1) # student at the SAME temperature
soft = -(p * log_q).sum(1).mean() # cross-entropy with the soft targets
hard = cross_entropy(z, y) # true labels y, temperature 1
loss = T**2 * soft + lam * hard # T^2 keeps the soft term's scale
loss.backward() # gradients reach the student only
# after training: predict with softmax(student(x)), i.e. T = 1With transfer images and classes, v, z, p and log_q are all . The teacher runs under no_grad, so its weights never change and the gradient of loss flows only into the student's weights. The value of lam is not given in the paper beyond being much lower than the soft term's weight.
The gradient, and why the soft loss is multiplied by T²
The paper's one derivation starts from the gradient of the soft-target cross-entropy with respect to a student logit. For a softmax followed by cross-entropy, the derivative with respect to the softmax's input is prediction minus target, because the logarithm in the loss cancels the exponential in the softmax. Here the softmax's input is , so the chain rule adds one factor of to get the derivative with respect to itself:
Gradient descent subtracts this, so a student logit moves down where the student gives the class more probability than the teacher () and up where it gives less. The in (2) is the chain-rule factor only. The second factor of appears when is large compared with the logits, because then itself shrinks.
To see it, use for small . With large relative to the logits, every exponent and is small, and each exponential in (2) can be replaced by its first-order expansion:
Now assume the logits of each transfer case are shifted to have mean zero, so that the student's logits and the teacher's logits each sum to zero. This costs nothing: adding the same constant to every logit leaves the softmax unchanged, so any network's logits can be centered per example without changing its predictions. Both denominators become , the two 1s in the numerators cancel, and two factors of remain:
At high temperature the soft-target gradient falls as , which the multiplier in the code cancels. Figure 2 computes the exact gradient (2) for one transfer case, the teacher logits of the 2 above and a partly trained student's logits for the same image, and plots its size against . Watch the low end and the high end separately. At the gradient has length 0.032, small because both softmaxes put nearly all their mass on the 2. It peaks near at 0.094 for this pair, then falls along the dashed line of equation (4): at the exact value is 0.040 against 0.035 from (4), and at it is 0.00149 against 0.00140. The approximation is poor at (0.56 against 0.032) because the teacher's logits span 9.8, far more than .
In numbers: the hard-label gradient for this student is 0.067. Without the multiplier, the soft term's gradient at is 0.0015, so the hard term is 45 times larger and the soft targets barely train the student. With the multiplier the soft term is 0.64 at and 0.60 at , nine to ten times the hard term at both temperatures, and the weight lam keeps the same meaning across the whole range.
Matching logits is the high-temperature limit of distillation
Equation (4) has a second reading. The expression is the gradient of with respect to , and the factor is a positive constant, which changes the step size but not where the minimum is. So at high temperature, with logits zero-meaned per case, distillation trains the student by least-squares regression of its logits onto the teacher's.
The paper credits logit matching to Caruana and collaborators, citing the 2006 Model Compression paper. That paper used the ensemble to label a large set of unlabeled (partly synthetic) inputs and trained a small network on those predictions. Regression onto the teacher's pre-softmax logits with a squared error appears in Ba and Caruana (2014).3 Distillation covers the whole range: a transfer set labeled by the teacher as in 2006, probability matching at as in speech work the paper cites (see the speech section below), and logit regression in the high-temperature limit.
Between those ends, the temperature sets how strongly the teacher's very negative logits pull on the student's logits. At high every logit gap counts equally, including the "definitely not this class" scores. During the teacher's own training, a class it already rates as very unlikely contributes almost no gradient, so whether its logit ends at −8 or −12 barely changes the teacher's loss, and the paper describes these logits as "almost completely unconstrained" and possibly very noisy. At a moderate , the soft targets for those classes are close to zero for both models, is close to zero, and the student is not pulled toward them. Whether ignoring them helps or hurts is, in the paper's words, an empirical question, since the negative logits may also carry information.
Figure 3 shows six classes of one transfer case, each with the teacher's logit, the student's logit and the negative gradient on that student logit, which is the direction a descent step moves it. The printed cosine compares this vector with pure logit matching, a negative gradient proportional to on every row. Look at the bottom row, the "4", whose teacher logit is far below the others: at low temperature its gradient is close to zero although its logit gap is the largest on the axis.
The MNIST experiments in the next section test the trade-off by shrinking the student. A student too small to absorb everything is better off spending its capacity on the larger logits, and the paper found that the best temperature drops as the student gets smaller.
MNIST: transferring how a model generalizes
The teacher is a network with two hidden layers of 1200 rectified linear units, trained on all 60,000 MNIST training images with dropout, weight constraints, and inputs jittered by up to two pixels in any direction. It makes 67 errors on the 10,000 test images. A smaller network with two hidden layers of 800 units, trained on the hard labels with no regularization, makes 146. The same 800-unit network trained with one added objective, matching the teacher's soft targets at , and no other regularization, makes 74. The transfer set contains no jittered images, so whatever the student learned about shifted digits came through the teacher's soft targets.
Temperature interacts with student size as the previous section predicted. With 300 or more units in each hidden layer, all temperatures above 8 gave similar results. With the student cut to 30 units per layer, temperatures from 2.5 to 4 worked significantly better than higher or lower ones.
The next test removes every 3 from the transfer set, so the student never receives a transfer image of a 3. It still sees 2s, 5s and 8s whose soft targets give class 3 a small probability, because some of those digits look a little like a 3. Without any adjustment, the distilled student makes 206 test errors, 133 of them on the 1,010 test 3s: it already classifies 877 of the 3s correctly.
The paper attributes most of the remaining errors to the output bias of class 3, which is much too low. The bias sets how often the student predicts the class, and the training signal on it is the soft-target gradient (2) summed over the transfer set. Training settles at the point where that sum is zero, so the student's average probability of 3 across the transfer images ends up equal to the teacher's, and the teacher gives 3 little probability on every image because none of them is a 3. Raising the class-3 bias by 3.5 brings the total down to 109 errors, 14 of them on 3s, so 996 of the 1,010 test 3s (98.6%) are classified correctly by a network whose transfer set had no 3 in it. The 3.5 is a tuned value: the paper chose it to optimize performance on the test set.
Figure 4 shows the bias trade-off. A bias change adds the same constant to the class-3 logit of every test image, which slides every image's margin (its 3 logit minus its best other logit) by that amount, for 3s and for all other digits alike. Raising the bias fixes 3s and, past some point, starts turning other digits into false 3s. Look for the minimum of the error curve and what happens to the two tails as the slider moves past it.
The same correction works in the opposite direction. If the transfer set contains only the 7s and 8s of the training set, the student over-predicts those two classes and makes 47.3% test errors. Lowering the biases of 7 and 8 by 7.6 brings this to 13.2%, so a network whose transfer set held only two digit classes classifies 86.8% of the test set correctly across all ten.
Distilling an ensemble of speech recognition models
The second experiment uses a production system, "a slightly outdated version of the acoustic model used by Android voice search". A speech recognizer of this kind slices audio into 10 ms frames, and the acoustic model classifies each frame into one of 14,000 states of a hidden Markov model (HMM, a model of how sounds follow one another in words). A decoder then searches for the word sequence that best balances high-probability states against what a language model considers a likely sentence.
The network sees 26 consecutive frames of 40 Mel-scale filterbank coefficients each, 1,040 input numbers, and predicts the HMM state of the 21st frame. It has 8 hidden layers of 2560 rectified linear units and a 14,000-way softmax. Counting weights: (2.7M) into the first layer, (45.9M) between hidden layers, and (35.8M) into the softmax, 84.4M in all, which matches the paper's "about 85M". Training uses about 2,000 hours of spoken English; at one frame per 10 ms that is 720 million frames, the paper's "about 700M training examples".
Two metrics are reported. Frame accuracy is the fraction of frames whose HMM state the network gets right on its own. Word error rate (WER) is the fraction of words the full recognizer, decoder and language model included, gets wrong, which is what a user experiences. The decoder repairs many frame-level mistakes, so frame accuracy around 60% is compatible with a WER around 11%.
| system (paper Table 1) | test frame accuracy | WER |
|---|---|---|
| baseline, one model | 58.9% | 10.9% |
| ensemble of 10 | 61.1% | 10.7% |
| distilled single model | 60.8% | 10.7% |
The distilled model has the same size as the baseline and is distilled from the ensemble of ten. It recovers 1.9 of the ensemble's 2.2-point frame-accuracy gain, 86%, which the paper reports as "more than 80%", and all of its WER gain. The ensemble improves WER less than frame accuracy (measured on a 23,000-word test set) because frame classification is not the objective the decoder optimizes. The deployed model costs the same per query as the baseline; the cost of training ten models stays offline.
The paper compares this with Li et al. (2014), who also trained a small acoustic model to match a larger model's class probabilities, but at and with a large unlabeled data set. Their best distilled model closed only 28% of the error-rate gap between the small and large models trained on hard labels.6
Soft targets as a regularizer: training on 3% of the data
The same speech model tests soft targets as a regularizer: train the 85M-parameter network on 3% of the data, about 20 million frames. With hard labels it overfits severely: training frame accuracy reaches 67.3%, higher than the 63.4% the full-data model reaches on its own training set, while test accuracy peaks at 44.5% and then drops sharply, so the run has to be early-stopped at the peak. Train the same network on the same 3% with soft targets from the full-data model and it converges to 57.0% test accuracy with no early stopping, about 2 points below the 58.9% obtained from all of the data.
The soft-target run reaches a similar training accuracy (65.4% against 67.3%), so its higher test accuracy does not come from fitting the small set less. Both runs fit it; the soft targets, produced by a model trained on all the data, pass on regularities from the 97% of frames this network never saw. Press Play to watch the two runs diverge.
Specialists for a 15,000-class problem
The last experiment applies the ensemble idea where a full ensemble cannot be trained at all. JFT is an internal Google data set of 100 million labeled images in 15,000 classes. Google's baseline convolutional network for it had trained for about six months with asynchronous stochastic gradient descent on many cores, so training several such networks for an ensemble was out of the question.
The paper's alternative keeps one generalist model trained on everything and adds many specialist models, each trained on a subset of classes that the generalist tends to confuse, such as different types of mushroom or different kinds of bridge. A specialist does not need a 15,000-way softmax: it keeps its own few hundred classes and merges every other class into a single dustbin class. In Figure 6 the generalist spreads its probability across a confusable cluster while the specialist for that cluster commits to one member.
Three choices keep specialists cheap. First, each specialist starts from the generalist's weights, so it inherits the low-level feature detectors and only refines them. Second, it trains on a stream that is half examples from its special classes and half examples drawn at random from everything else. That stream overstates the special classes, so after training the dustbin's logit is increased by the log of the oversampling factor: adding to a logit multiplies that class's unnormalized score by , which undoes a -fold overrepresentation of the special classes relative to the dustbin. Third, the clusters are found without labels. The paper runs online K-means on the columns of the covariance matrix of the generalist's predicted probabilities. Each class's column records how its predicted probability rises and falls with every other class's across images, so two classes the generalist keeps splitting its probability between have similar columns and fall into the same cluster. The paper's Table 2 lists clusters such as bridge, cable-stayed bridge, suspension bridge and viaduct. With the clusters fixed, specialists train independently of one another, in a few days instead of many weeks.
At test time the specialists combine with the generalist in two steps. The generalist names its most probable class . The specialists whose class subsets contain form the active set , which may be empty. The combined prediction is the full distribution over all classes that is closest, in summed KL divergence, to the generalist's distribution and to each active specialist's distribution :
A specialist has probabilities only for its own classes and the dustbin, so when is compared with it, all of 's probability on the specialist's outside classes is summed into one dustbin number. There is no closed-form solution in general. The paper writes as a softmax of free logits at and runs gradient descent on those logits, separately for every test image.
The direction of the KL divergence decides what kind of average comes out. The paper notes that when every model gives one probability per class, the solution is the arithmetic or the geometric mean of the models' distributions depending on the direction. Equation (5) puts each model's distribution first, the forward KL, and minimizing a sum of forward KL divergences over gives the arithmetic mean. If two models give a class 0.6 and a third gives it 0.001, the arithmetic mean is 0.40; the geometric mean, before renormalizing, is 0.071, so one model that rules a class out pulls it near zero for everyone.
With 61 specialists, each covering 300 classes plus the dustbin, top-1 test accuracy on JFT rises from 25.0% to 26.1%, a 4.4% relative improvement (1.1 points). The paper also reports a conditional accuracy, measured only on test images from specialist classes with predictions restricted to those classes, which rises from 43.1% to 45.9%; it is a different, narrower quantity than the headline. Because the specialists' class subsets overlap, many classes are covered by more than one specialist, and the gain grows with that count: +3.4% relative accuracy for classes with one specialist, +7.4% with two, up to +16.6% with nine (and +14.1% for ten or more). The 350,037 test images in classes no specialist covers are unchanged. Summed over the 677,313 test images in the table, the specialists add 6,821 correct top-1 answers, a 1.0-point gain, consistent with the headline.
This setup resembles a mixture of experts, in which a gating network learns which expert should handle each example while the experts train (the design behind sparse models such as Switch Transformers). A learned gate makes training hard to parallelize: each expert's effective training set keeps changing with the others' performance, and the gate has to compare experts on the same example. The specialists here have no learned gate. The routing is the frozen generalist's top-1 class, fixed before any specialist trains, so every specialist can train in parallel with no communication.
What the paper left open, and what came after
The JFT specialists were combined at test time; the paper had not yet distilled them back into one network, and says so in its discussion. Section 6.1 proposes, as work in progress, a fix for specialists' tendency to overfit their enriched training streams: give each specialist a full softmax and train it on soft targets from the generalist for the non-special classes, as the 3% speech experiment suggests would preserve its knowledge of those classes.
The training rule itself became the standard way to compress classifiers and later spread to other model types. In this library, Proxy-KD distills a black-box language model such as GPT-4, whose logits are not available, by first aligning an open proxy model to it and then distilling the proxy's full output distribution into a student. The rCM paper distills a many-step diffusion model into a one- to four-step generator, combining consistency distillation with score distillation.
Questions you might still have
Is raising the temperature the same thing as label smoothing?
No. Label smoothing gives every wrong class the same small probability, for every input. Soft targets come from a trained teacher, so they differ from image to image and put their mass on the classes this particular image resembles: a 2 that looks like a 3 gets a larger 3 probability than a 2 that looks like a 7. Temperature only scales how much of that structure reaches the loss.
Why is the soft-target loss multiplied by T²?
At high temperature the soft-target gradient shrinks as 1/T² (equation 4). Without a correction, raising T from 4 to 20 cuts the soft term’s gradient by about 27 times in the worked example on this page while the hard-label term stays fixed, so the hard labels take over. Multiplying the soft loss by T² cancels the shrinkage and keeps the two terms in roughly the same proportion whatever T you try.
How can the distilled model recognize a digit it never saw?
With every 3 removed from the transfer set, the student still sees 2s, 5s and 8s whose soft targets give class 3 a little probability, because those digits resemble a 3. That is enough for it to get 877 of the 1,010 test 3s right. The remaining errors come mostly from the output bias of class 3, which training left low because no transfer image was a 3. Raising that one bias by 3.5 brings it to 996 of 1,010, 98.6%.
What temperature should I use?
The paper used T = 20 for its 800-unit MNIST student and reports that for students with 300 or more units per layer every temperature above 8 gave similar results. For a student cut to 30 units per layer, temperatures between 2.5 and 4 worked significantly better than higher or lower ones. A small student does better at a moderate temperature that ignores the teacher’s most negative logits.
Does the student need labels, or the same architecture as the teacher?
Neither. The transfer set can be entirely unlabeled, since the teacher supplies the targets; the paper found the original training set works well, especially with a small extra loss on the true labels. The student can be any network with the same output classes: the MNIST student had 800 units per layer against the teacher’s 1200, and the speech student matched the size of one ensemble member, not the ensemble of ten.
Did earlier work already match logits?
Caruana and collaborators showed in 2006 that an ensemble can be compressed into one small network, by labeling a large transfer set with the ensemble’s predictions and training on it. Regressing the student’s pre-softmax logits onto the teacher’s with squared error is the method of Ba and Caruana (2014). Distillation contains logit matching as its high-temperature limit.
Is this where modern LLM and diffusion distillation comes from?
Yes. The temperature softmax and the soft-target loss in this paper are the ancestor of most later distillation methods. Two explainers in this library follow the line forward: Proxy-KD distills a black-box LLM such as GPT-4 through an open proxy model, and rCM distills a many-step diffusion model into one to four steps.
Footnotes & further reading
- The paper: Hinton, Vinyals, Dean, Distilling the Knowledge in a Neural Network (Google; NIPS 2014 Deep Learning Workshop; arXiv March 2015). ↩
- The prior model-compression result: Buciluǎ, Caruana, Niculescu-Mizil, Model Compression (KDD 2006). It trains the small model on the ensemble's predictions over a large transfer set. ↩
- The explicit logit-matching objective: Ba, Caruana, Do Deep Nets Really Need to be Deep? (2014), which regresses the small net's pre-softmax logits onto the teacher's with a squared error. ↩
- The name "dark knowledge" comes from a Hinton talk, Dark Knowledge (TTIC, 2014), not from the paper, which calls it a "rich similarity structure". ↩
- Dropout as an implicit ensemble: Hinton, Srivastava, Krizhevsky, Sutskever, Salakhutdinov, Improving neural networks by preventing co-adaptation of feature detectors (2012). ↩
- Li, Zhao, Huang, Gong, Learning small-size DNN with output-distribution-based criteria (Interspeech 2014). ↩
- Descendants in this library: distilling a black-box LLM through an aligned proxy in Proxy-KD, and few-step diffusion distillation in rCM (score-regularized consistency models).
How could this explainer be improved? Found an error, or something unclear? I read every message.