ViT, explained · Part 1 of 2 · Covers Title, Abstract, §1

The Big Idea: An Image Is Worth 16×16 Words

The title, the abstract and the introduction of the Vision Transformer paper, line by line: what it means to cut a picture into 16×16 patches and read them like words, why convolutional networks ruled computer vision, what inductive bias, locality and translation equivariance are, and the claim that enough data beats built-in assumptions. With a real run of ViT-B/16 and ResNet-50 on the same picture and the quadratic cost of attention worked out by hand.

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby. ICLR 2021, 2020. arXiv:2010.11929

On 22 October 2020, twelve researchers at Google posted a paper with an odd title: An Image is Worth 16×16 Words. It took the Transformer, the network design that had taken over language processing, cut photographs into small squares, and fed the squares to the Transformer as if they were words. On the standard photo benchmark, ImageNet, the model matched or beat the best convolutional networks of the day, once it had been trained on enough pictures. The paper was accepted at ICLR 2021, and the model it introduced, the Vision Transformer (ViT), is today the base of most image models and of the image half of models that read both text and pictures.

This series reads the paper slowly, in the paper's own order. Each piece follows the same pattern:

  1. The exact lines from the paper, as a highlighted screenshot in a teal box like the one below.
  2. A plain-English explanation, with a yellow box for every new word.
  3. A picture, and real code when it helps, with its real output.
  4. Why it matters, and only then the next piece.

The small Paper §1 tag above each heading tells you which section of the paper you are reading. The screenshots come from the paper's second arXiv version (3 June 2021), the one most people read today. We run every experiment on the model weights Google released, through the Hugging Face transformers library, on an ordinary laptop.

The title

Let us take the title apart, one piece at a time.

  • An image. A digital picture. For this paper, a photograph of one main thing: a cat, a car, a bird.
  • 16×16 words. The picture is cut into squares of 16 pixels by 16 pixels. Each square is treated as one "word". A 224×224 picture gives 14×14 = 196 such squares. The paper calls them patches.
  • Transformers. The kind of neural network used. It was introduced in 2017 for translating sentences, and by 2020 it was the standard network for text. The paper uses it with almost no changes.
  • Image recognition. The task: look at a picture and say what is in it. The paper's version of the task is image classification: choose one label out of a fixed list (for ImageNet, one of 1,000 labels).
  • At scale. The key words. The method only works well when the model is trained on very many pictures: 14 million to 300 million. On the "ordinary" 1.3 million pictures of ImageNet, it loses to convolutional networks. This is the paper's main finding, and most of the paper is about it.
2012AlexNetdeep CNN wins theImageNet contest2015ResNetCNNs 100+ layersdeep; still thevision standard2017Transformerattention only;built fortranslation2018BERTpre-train, thenfine-tune: theNLP recipe2020ViTthe Transformer,unchanged, onimage patchescomputer visionlanguage (NLP)this paper: the two lines meet
Two lines of research meet. In vision, AlexNet (2012) and ResNet (2015) made convolutional networks the standard. In language, the Transformer (2017) and BERT (2018) replaced older designs and introduced the pre-train-then-fine-tune recipe. ViT (2020) takes the language design, unchanged, to pictures.
History: the five milestones on this timeline (optional reading)

If you want the back-story behind each dot on the line, read on. If not, skip to the next section; nothing below is needed for the paper.

2012: AlexNet. Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton (University of Toronto) entered a deep convolutional network in the ImageNet contest and won by a large margin: a top-5 error of 15.3%, against 26.2% for the runner-up, which used hand-made features. Three choices made it work: the ReLU activation, which is simply f(x)=max⁡(0,x)f(x) = \max(0, x) and trains much faster than the smooth curves used before; training on two GPUs; and a trick called dropout to fight overfitting. The result convinced the field that learning features from data beats designing them by hand, and it started the deep-learning era in computer vision.

2015: ResNet. Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun (Microsoft Research Asia) found that very deep networks were harder to train than shallow ones, even on the training data. Their fix was the residual connection: instead of making a block of layers learn a new output H(x)H(x), make it learn only the change, so the block computes

y=F(x)+xy = F(x) + x

where xx is the block's input, F(x)F(x) is what the layers compute, and the "+ x+\,x" is a shortcut that passes the input straight through. If a block has nothing useful to add, it can learn F(x)≈0F(x) \approx 0 and do no harm. This let them train 152-layer networks, win ImageNet 2015 with 3.57% top-5 error, and it is why ViT has a "+" after every block too (Part 2, Equations 2 and 3). The ResNets in this paper are direct descendants.

2017: the Transformer. Ashish Vaswani and seven colleagues at Google published Attention Is All You Need, a network for translating sentences that used no recurrence and no convolution, only attention: every word computes how much to look at every other word,

Attention(Q,K,V)=softmax⁡ ⁣(QK⊤dk)V\text{Attention}(Q, K, V) = \operatorname{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right) V

where QQ, KK and VV are three versions of the word vectors (queries, keys and values), QK⊤QK^\top scores every pair of words, dk\sqrt{d_k} keeps the scores from growing too large, and softmax turns each row into weights that add up to 1. Because every word can reach every other word in one step, the design trains fast on GPUs and TPUs, and it scales: bigger models kept getting better. ViT uses this encoder with almost no change.

2018: BERT. Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova (Google) took the Transformer encoder, hid 15% of the words in billions of sentences, and trained the model to guess the hidden words from both sides (the masked language model). The pre-trained model was then fine-tuned, with one small new layer, for eleven different language tasks, and set a new best result on all of them. Two things carry over to ViT directly: the recipe (pre-train once on a huge dataset, fine-tune cheaply per task) and the special [CLS] token whose output vector stands for the whole input. ViT's [class] token is the same idea, and its model sizes (Base and Large) are copied from BERT. The BERT breakdown on this site reads that paper in full.

2020: ViT. Twelve researchers at Google Research (Brain Team), led by Alexey Dosovitskiy and Neil Houlsby, asked what happens if the two lines meet: cut a picture into 16×16 patches, treat each patch as a word, and feed them to the BERT-style encoder. The answer, as the rest of this part explains, is that it loses on 1.3 million pictures and wins on 300 million. That finding, "large scale training trumps inductive bias", is the paper's thesis, and it made the Transformer the common design for text and images alike.

The title's analogy is worth drawing out, because the whole paper rests on it. A language model reads a sentence as a sequence of word tokens, each turned into a vector of numbers. ViT reads a picture as a sequence of patch tokens, each turned into a vector of numbers of the same size. After that step, the Transformer cannot tell whether it is reading a sentence or a photo.

Language: words → tokens → vectorsVision: patches → tokens → vectors"The cat sat on the mat"Thecatsatonthemat6 tokens, each a vector of 768 numbersa picture, cut into 3 × 3p1p2p3p4p5p6p7p8p99 patch tokens, each a vector of 768 numbers
The analogy of the title. Left: a sentence becomes six word tokens, and each token becomes a vector of 768 numbers. Right: a picture is cut into nine patches, each patch becomes a token, and each token becomes a vector of 768 numbers. The same Transformer reads both.

Here is the paper's own drawing of the model, which Part 2 explains block by block. We show it now because it makes the title concrete.

The authors and the footnotes

The order of the names matters less than usual: six of the twelve are marked as equal contributors. When this series says "the authors" it means the team.

There is one footnote on page 1, attached to the last sentence of the abstract.

The abstract, sentence by sentence

The abstract is one paragraph of four sentences. We read it in two halves.

Sentence 1 says that in language the question was settled. "De-facto standard" means "the standard in practice, even if nobody voted on it". By 2020 every strong language model was a Transformer. The second half of the sentence says vision had not followed.

Sentence 2 names the two ways vision had used attention so far. Both keep the convolutional network in charge. To see why that matters, we need the two words.

CNN: a small filter slides over the imageViT: every patch attends to every patchthe window visits every position3 × 3 window: each output sees 9 neighbours (locality)the same 9 weights at every position (weight sharing)far-away pixels meet only after many layersone patch looks at all 49 patches, in the first layerhow much to look at each is learned (attention weights)no built-in idea that neighbours matter more
The two operations side by side. Left: a convolution slides a 3×3 window over the picture; each output sees only its 9 neighbours and the same 9 weights are used at every position. Right: in a Transformer one patch attends to all 49 patches at once, in the very first layer, with weights that are learned rather than fixed by position.

The figure shows the real difference between the two operations. A convolution is local (each output looks at a small neighbourhood) and shares weights (the same filter is used everywhere). Self-attention is global (each token can look at every token) and its weights depend on the content, not on the position. Hold on to those two words, local and global; the introduction will come back to them under the name "inductive bias".

Sentence 3 is the thesis. "Pure transformer" means no convolution layers inside the network (Part 2 will show that the very first step, the patch embedding, can be written as one convolution, but the authors mean no CNN body). "Sequences of image patches" is the title again.

Sentence 4 is the method and the result, and it introduces the two-step recipe that the whole paper follows.

Step 1: pre-train once on a very large labelled image setStep 2: transfer to each small benchmarkImageNet-21k: 14M imagesor JFT-300M: 303M imageswith class labelsthousands of TPU-core-daysViTlearns to seeall the weightsare learned hereViT copystarts pre-trainedImageNet1,000 classes, 1.3M imagesViT copystarts pre-trainedCIFAR-100100 classes, 50k tiny imagesViT copystarts pre-trainedVTAB19 tasks, 1,000 images each+ one new classification layer per benchmark (fine-tuning, Part 3)
The recipe of the paper. Step 1: pre-train ViT once on a very large labelled picture set (ImageNet-21k or JFT-300M), which costs thousands of TPU-core-days. Step 2: for each small benchmark, start from a copy of the pre-trained model, add one new output layer, and fine-tune.

If you have read the BERT breakdown, this recipe is familiar: BERT pre-trains on text once and fine-tunes a copy per task. ViT does the same with pictures, with one difference. BERT's pre-training needs no labels (it hides words and guesses them). ViT's pre-training in this paper is supervised: every pre-training picture comes with a class label.

The abstract also says what the paper does not do. It says "image classification tasks", not detection or segmentation.

image classificationone label for the whole picture"cat"this paperobject detectiona box around every objectcat, remoteleft to later work (Part 6)segmentationa class for every pixelevery pixel labelledleft to later work (Part 6)
What the paper does and does not do. It trains ViT for image classification: one label for the whole picture. Object detection (a box around every object) and segmentation (a class for every pixel) are left to later work, which Part 6 covers.

Let us run it

Before reading the introduction, it helps to see the object the paper is talking about. We loaded the released ViT-B/16 model (google/vit-base-patch16-224: pre-trained on ImageNet-21k, fine-tuned on ImageNet) and gave it one picture, the test photo from the huggingface/cats-image dataset.

The sample picture: two tabby cats lying on a pink sofa, with a remote control between them
The picture we use throughout this part: two tabby cats on a pink sofa, with two remote controls. It is 640×480 pixels; the model resizes it to 224×224.

The code is in vit_part1.py:

python
proc = AutoImageProcessor.from_pretrained('google/vit-base-patch16-224')
vit = ViTForImageClassification.from_pretrained('google/vit-base-patch16-224').eval()
x = proc(images=image, return_tensors='pt').pixel_values            # resized to 224 x 224, scaled to [-1, 1]
C, H, W = x.shape[1:]
P = vit.config.patch_size
N = (H // P) * (W // P)
plain text
ViT-B/16 (google/vit-base-patch16-224)
  image tensor: (1, 3, 224, 224)  = batch of 1, 3 channels, 224 x 224 pixels
  numbers in one image: 3 x 224 x 224 = 150,528
  patch size P = 16; patches per side = 224 // 16 = 14; number of patches N = 14 x 14 = 196
  numbers in one patch: 16 x 16 x 3 = 768
  sequence length with the [class] token: 196 + 1 = 197
  hidden size D = 768, layers = 12, heads = 12
  parameters: 86,567,656
  token matrix after the embedding step: (1, 197, 768)  (1 image, 197 tokens, 768 numbers each)
  output scores: (1, 1000)  (one score per ImageNet class)

Every number in that output is one line of arithmetic, and it is worth doing the arithmetic once by hand. The picture is a block of numbers of shape channels × height × width:

x∈RC×H×W,C⋅H⋅W=3×224×224=150,528\mathbf{x} \in \mathbb{R}^{C \times H \times W}, \qquad C \cdot H \cdot W = 3 \times 224 \times 224 = 150{,}528

where:

  • C=3C = 3 is the number of colour channels (red, green, blue);
  • H=W=224H = W = 224 are the height and width in pixels, after the model's preprocessing resized the photo;
  • 150,528150{,}528 is how many numbers the model receives for one picture.

Cutting it into patches of side P=16P = 16 gives

N=HP⋅WP=22416⋅22416=14×14=196 patches,P2C=16×16×3=768 numbers per patchN = \frac{H}{P} \cdot \frac{W}{P} = \frac{224}{16} \cdot \frac{224}{16} = 14 \times 14 = 196 \ \text{patches}, \qquad P^2 C = 16 \times 16 \times 3 = 768 \ \text{numbers per patch}

and the sequence the Transformer reads has one extra token at the front, the [class] token (explained in Part 2), so its length is

N+1=197N + 1 = 197

The 196 × 768 = 150,528 numbers of the patches are exactly the 150,528 numbers of the picture: nothing is lost or added by cutting, only rearranged. The name ViT-B/16 encodes two of these numbers: "B" for Base (the BERT-base size: 12 layers, 768-wide vectors, 12 attention heads) and "/16" for the patch side. The model has 86.6 million parameters, which the paper rounds to 86M in its Table 1.

224 pixels = 14 patches × 16 pixels14 × 14 = 196 patchesone patch: 16 × 16 pixels× 3 colour channelsthe arithmetic of ViT-B/16the image3 × 224 × 224 = 150,528 numbersone patch16 × 16 × 3 = 768 numberspatches(224 / 16)² = 14 × 14 = 196tokens196 + 1 [class] = 197token matrix197 × 768 numbers
The arithmetic of ViT-B/16. A 224×224 picture is a 14×14 grid of 16×16 patches: 196 patches, each 16 × 16 × 3 = 768 numbers. With one class token the sequence has 197 tokens, each a vector of 768 numbers.

Now the answer. The model outputs 1,000 scores, one per ImageNet class; a softmax turns them into probabilities that add up to 1:

pk=esk∑j=11000esjp_k = \frac{e^{s_k}}{\sum_{j=1}^{1000} e^{s_j}}

where sks_k is the raw score for class kk and pkp_k is its probability. We printed the five largest, and then did the same with a classic convolutional network, ResNet-50 (microsoft/resnet-50), on the same picture:

python
rproc = AutoImageProcessor.from_pretrained('microsoft/resnet-50')
resnet = ResNetForImageClassification.from_pretrained('microsoft/resnet-50').eval()
xr = rproc(images=image, return_tensors='pt').pixel_values
with torch.no_grad():
    ro = resnet(pixel_values=xr, output_hidden_states=True)
p_res = torch.softmax(ro.logits[0], -1)
plain text
  top-5 classes:                       (ViT-B/16)
     0.9374  Egyptian cat
     0.0384  tabby
     0.0144  tiger cat
     0.0033  lynx
     0.0007  Siamese cat

ResNet-50 (microsoft/resnet-50), a convolutional network
  image tensor: (1, 3, 224, 224)
  parameters: 25,557,032
  feature map after stage 0: (64, 56, 56)  (channels, height, width)
  feature map after stage 1: (256, 56, 56)  (channels, height, width)
  feature map after stage 2: (512, 28, 28)  (channels, height, width)
  feature map after stage 3: (1024, 14, 14)  (channels, height, width)
  feature map after stage 4: (2048, 7, 7)  (channels, height, width)
  top-5 classes:
     0.9416  tiger cat
     0.0344  tabby
     0.0016  remote control
     0.0014  Egyptian cat
     0.0008  jinrikisha

both models agree on the top class: False  (ViT: Egyptian cat, ResNet-50: tiger cat)
the same picture (two tabby cats on a pink sofa) through two modelsViT-B/16 (Transformer)Egyptian catEgyptian cat: 0.93740.9374tabbytabby: 0.03840.0384tiger cattiger cat: 0.01440.0144lynxlynx: 0.00330.0033Siamese catSiamese cat: 0.00070.0007ResNet-50 (CNN)tiger cattiger cat: 0.94160.9416tabbytabby: 0.03440.0344remote controlremote control: 0.00160.0016Egyptian catEgyptian cat: 0.00140.0014jinrikishajinrikisha: 0.00080.0008both say "cat" with about 94% confidence; they disagree about the breed (the picture has two cats and a remote control)
The same picture through two models. ViT-B/16 says Egyptian cat with probability 0.937; ResNet-50 says tiger cat with 0.942. Both are sure it is a cat and both put "tabby" second. They disagree about the breed, and ResNet-50 also noticed the remote control.

Three things to notice:

  • Both models are right, in the sense that matters: the picture shows cats, and both put about 94% on a cat class. ImageNet has several house-cat classes (Egyptian cat, tabby, tiger cat, Persian, Siamese), and the picture does not clearly belong to one, so the disagreement is about a fine label, not about the object.
  • ResNet-50 is a different shape of machine. Its printed "feature maps" shrink from 56×56 to 7×7 as the picture passes through five stages, while the number of channels grows from 64 to 2,048. That is what a CNN does: it summarises larger and larger neighbourhoods. ViT keeps its 197 tokens of 768 numbers at every layer; nothing shrinks.
  • ResNet-50 has 25.6 million parameters, ViT-B/16 86.6 million. The paper's fair comparison is against much bigger CNNs (Part 4). This run is only a first look, not a contest.

Transformers took over language

Now the introduction. Its first paragraph is about language, not pictures.

Each of the four citations stands for one idea, and the paper leans on all four. We screenshot the sentences it points to.

parameters of NLP Transformers named in the first paragraph (log scale)Transformer "big", 2017Transformer "big", 2017: 213,000,000 parameters213MBERT-large, 2018BERT-large, 2018: 340,000,000 parameters340MGPT-3, 2020GPT-3, 2020: 175,000,000,000 parameters175BGShard, 2020GShard, 2020: 600,000,000,000 parameters600B100M1B10B100B1TViT-B/16 has 86M parameters, ViT-L/16 307M, ViT-H/14 632M: the sizes of BERT, not of GPT-3
How big the language Transformers in that paragraph are, on a log scale: 213 million parameters for the 2017 Transformer's large version, 340 million for BERT-large, 175 billion for GPT-3 and 600 billion for GShard. The ViT models of this paper (86 million to 632 million) are BERT-sized, not GPT-3-sized.

The paragraph's last sentence is the one to remember: "there is still no sign of saturating performance". In language, the curve of accuracy against model size and data size had not flattened. The rest of the introduction asks whether the same curve exists for pictures, and answers yes, with a catch.

In vision, convolutions still ruled

Let us take the three sentences in turn.

"Convolutional architectures remain dominant." The three citations are the history of CNNs in three steps. LeCun et al. (1989) trained a convolutional network to read handwritten postcodes, the first practical CNN. Krizhevsky et al. (2012), the AlexNet paper, won the 2012 ImageNet contest by a large margin with a deep CNN trained on GPUs, which started the deep-learning era in vision. He et al. (2016), the ResNet paper, made CNNs of 100+ layers trainable.

"Combining CNN-like architectures with self-attention." The second sentence lists what people had tried. Non-local networks (Wang et al., 2018) added an attention layer inside a CNN for video. DETR (Carion et al., 2020) put a Transformer on top of a CNN's output to detect objects. Stand-alone self-attention (Ramachandran et al., 2019) and Axial-DeepLab (Wang et al., 2020a) went further and replaced every convolution with a local form of attention. The abstract's two categories are these: attention with a CNN, and attention replacing parts of a CNN while "keeping their overall structure in place".

1. CNN + attentionNon-local nets (2018)DETR (2020)conv layerconv layerconv layerconv layerattention on toppixelsthe CNN still does the work2. attention inside a CNNStand-alone (2019)Axial-DeepLab (2020)conv layerlocal attentionconv layerlocal attentionconv layerpixelsthe CNN skeleton stays3. pure local attention"specialized attentionpatterns", hard on TPUslocal attentionlocal attentionlocal attentionlocal attentionlocal attentionpixelsefficient in theory, slow in practiceViT: a standard Transformerglobal attention,patches as tokensTransformer layerTransformer layerTransformer layerTransformer layerTransformer layer16 × 16 patches"fewest possible modifications"
The three earlier ways of using attention in vision, next to ViT. (1) A CNN does the work and attention is added on top (Non-local networks, DETR). (2) Some CNN layers are swapped for local attention, but the CNN skeleton stays (stand-alone self-attention, Axial-DeepLab). (3) Every layer is local attention with a special pattern, efficient in theory but hard to run fast. ViT is a standard Transformer with global attention over patch tokens: the "fewest possible modifications".

"Specialized attention patterns." The third sentence is the technical heart of the paragraph. Why did the convolution-free models need special patterns at all? Because plain attention over pixels is impossibly expensive. The paper spells this out in its related work (Part 2): "each pixel attends to every other pixel. With quadratic cost in the number of pixels, this does not scale to realistic input sizes." Let us compute the cost, because the whole design of ViT follows from it.

The quadratic cost of attention over pixels

Self-attention computes, for every token, a score against every token. With nn tokens that is an n×nn \times n table of scores per attention head per layer. The attention series derives the formula; here it is with the shapes written in:

Attention(Q,K,V)=softmax⁡ ⁣(QK⊤d)V\text{Attention}(Q, K, V) = \operatorname{softmax}\!\left(\frac{QK^\top}{\sqrt{d}}\right) V

where:

  • QQ, KK, VV are the queries, keys and values, three n×dn \times d tables made from the nn token vectors (for ViT-B, d=64d = 64 per head);
  • QK⊤QK^\top is the n×nn \times n table of scores: entry (i,j)(i, j) says how much token ii should look at token jj;
  • the softmax turns every row into weights that add up to 1;
  • the result is n×dn \times d: one new vector per token.

The expensive part is the n×nn \times n table. Its size is

attention weights per head per layer=n2,memory in float32=4 n2 bytes\text{attention weights per head per layer} = n^2, \qquad \text{memory in float32} = 4\,n^2 \ \text{bytes}

We computed this table for several choices of what a token is (vit_part1_math.py, part (a)):

python
for name, n in [('32 x 32 CIFAR pixels', 32 * 32), ('224 x 224 ImageNet pixels', 224 * 224),
                ('ViT-B/32 patches + [class] (7x7+1)', 7 * 7 + 1), ('ViT-B/16 patches + [class] (14x14+1)', 14 * 14 + 1),
                ('ViT-B/16 at 384 px (24x24+1)', 24 * 24 + 1)]:
    nn_ = n * n
    bytes_ = nn_ * 4
plain text
    tokens are                                n            n x n  float32 memory
    32 x 32 CIFAR pixels                  1,024        1,048,576         4.19 MB
    224 x 224 ImageNet pixels            50,176    2,517,630,976        10.07 GB
    ViT-B/32 patches + [class] (7x7+1)       50            2,500         10.0 kB
    ViT-B/16 patches + [class] (14x14+1)      197           38,809        155.2 kB
    ViT-B/16 at 384 px (24x24+1)            577          332,929         1.33 MB
    pixels / patch tokens = 50,176 / 197 = 254.7 x more tokens, so 64,872 x more attention weights
    ViT-B/16 has 12 heads x 12 layers = 144 attention tables of 197 x 197 = 5,588,496 weights for one image
attention weights per head per layer = n × n (log scale)ViT-B/32 patches (7x7+1)ViT-B/32 patches (7x7+1): 2,500 weights2,500 (10.0 kB)ViT-B/16 patches (14x14+1)ViT-B/16 patches (14x14+1): 38,809 weights38,809 (155.2 kB)ViT-B/16 at 384 px (24x24+1)ViT-B/16 at 384 px (24x24+1): 332,929 weights332,929 (1.33 MB)32 x 32 CIFAR pixels32 x 32 CIFAR pixels: 1,048,576 weights1,048,576 (4.19 MB)224 x 224 ImageNet pixels224 x 224 ImageNet pixels: 2,517,630,976 weights2,517,630,976 (10.07 GB)10^310^410^510^610^710^810^9pixels of a 224 × 224 image: 254.7 × more tokens than ViT-B/16, so 64,872 × more attention weights
Attention weights per head per layer, on a log scale. Over the 50,176 pixels of a 224×224 picture, one attention table has 2.5 billion entries (10 GB in float32). Over ViT-B/16's 197 patch tokens it has 38,809 entries (155 kB). Even the 1,024 pixels of a tiny CIFAR picture need a million entries.

Read the second row. If every pixel of a 224×224 picture attended to every other pixel, one attention table of one head of one layer would hold 50,1762≈2.550{,}176^2 \approx 2.5 billion numbers, about 10 GB. ViT-B has 144 such tables (12 heads × 12 layers), and training needs several copies of each. It cannot be done. Even for the tiny 32×32 pictures of CIFAR, one table has a million entries. That is why the earlier convolution-free models restricted each pixel to a local neighbourhood, and why those restrictions needed "specialized attention patterns" that ran poorly on accelerators.

Now read the fourth row. With 16×16 patches as tokens, n=197n = 197 and one table has 38,809 entries. The ratio between the two is

50,17621972=(50,176197)2≈254.72≈64,872\frac{50{,}176^2}{197^2} = \left(\frac{50{,}176}{197}\right)^2 \approx 254.7^2 \approx 64{,}872

Patches make attention about 65,000 times cheaper than pixels. That single fact is the engineering reason the title says "16×16 words" rather than "50,176 words". And 197 tokens is a very ordinary sequence length for a Transformer (BERT was trained with 512), so the standard, dense, accelerator-friendly attention can be used unchanged. The paper's related work (Part 2) credits Cordonnier et al. (2020) with the patch idea, using 2×2 patches on small pictures; the ViT authors use bigger patches, bigger pictures and far more data.

"Classic ResNet-like architectures are still state of the art." The paragraph closes with the three results that defined the top of the ImageNet leaderboard in 2020: Mahajan et al. (2018) pre-trained ResNeXt CNNs on 3.5 billion Instagram pictures; Xie et al. (2020), Noisy Student, used 300 million unlabelled pictures with a self-training loop; and Kolesnikov et al. (2020), Big Transfer (BiT), pre-trained very large ResNets on the same ImageNet-21k and JFT-300M datasets that ViT will use. All three are CNNs made better with more data, which is the same lever ViT pulls.

The fewest possible modifications

The four sentences are the whole method. We illustrate each.

"Split an image into patches." The picture is cut into a grid of equal squares and the squares are read off in order, row by row, like text.

an image, cut into a 3 × 3 grid123456789flattenrow by row123456789a sequence of 9 patch tokens, in reading order (left to right, top to bottom)patch 1 is the top-left corner, patch 9 the bottom-right; the order is fixeda position embedding (Part 2) tells the model where each patch came from
An image cut into a 3×3 grid of patches and flattened, row by row, into a sequence of nine patch tokens. Patch 1 is the top-left corner and patch 9 the bottom-right. The order is fixed; in Part 2 a position embedding tells the model where each patch came from.

"The sequence of linear embeddings of these patches." Each patch is 16 × 16 × 3 = 768 raw pixel numbers. A "linear embedding" multiplies those 768 numbers by a learned matrix to give the patch's token vector. In the paper's notation (Equation 1, which Part 2 covers in full) a flattened patch xpi\mathbf{x}_p^i becomes

xpiE,xpi∈R1×768,E∈R768×D\mathbf{x}_p^i \mathbf{E}, \qquad \mathbf{x}_p^i \in \mathbb{R}^{1 \times 768}, \quad \mathbf{E} \in \mathbb{R}^{768 \times D}

where DD is the Transformer's width (D=768D = 768 for ViT-B, so E\mathbf{E} happens to be square, 768 × 768). "Linear" means no non-linearity: it is one matrix multiplication, the simplest possible way to turn a patch into a vector. The output has the same shape as a word embedding in BERT, which is the point.

one patch, 16 × 16 pixels × 3 channelsred, green, blueflatten…768 numbers in a row (16 × 16 × 3)multiply by a learned matrix E (768 × 768): the "linear embedding"…one patch token: a vector of D = 768 numbers, the same shape as a word vector in BERT
One patch becomes one token. The 16×16×3 block of pixels is flattened into a row of 768 numbers, multiplied by a learned 768×768 matrix E (the "linear embedding"), and comes out as a vector of 768 numbers, the same shape as a word vector in BERT.

Our code found this layer in the released model. In the implementation it is stored as a convolution with a 16×16 filter and a stride of 16, which is exactly the same computation as "cut into 16×16 patches and multiply each by E\mathbf{E}", because a stride-16 filter of size 16 touches each pixel exactly once:

plain text
    patch embedding layer: Conv2d(3, 768, kernel_size=(16, 16), stride=(16, 16))
    patch vectors: (196, 768) = 196 patches x 768 numbers, arranged on a 14 x 14 grid

"Treated the same way as tokens (words)." After the embedding, the Transformer sees 196 vectors of 768 numbers. It would see the same thing for a 196-word sentence in BERT. Nothing in the attention layers, the feed-forward layers or the normalisation knows about pixels, rows or columns.

"In supervised fashion." Pre-training is ordinary classification with labels, on a very large labelled set. BERT's label-free pre-training was its big trick; ViT does not need a trick, because Google had hundreds of millions of labelled pictures. (Part 5 shows a first attempt at label-free pre-training for ViT, and Part 6 shows how later papers made it work.)

Mid-sized data, and the inductive-bias explanation

The fourth paragraph starts at the bottom of page 1 and ends at the top of page 2. It reports the first, disappointing result, and explains it.

Four technical words appear here, and the rest of the series uses all of them, so we take them slowly.

The ResNet-50 run above showed locality in action: its feature maps shrink from 56×56 to 7×7 stage by stage, which is how a CNN gradually lets far-away pixels meet. ViT's 197 tokens all meet in the first layer.

Let us write these two properties down, because they are often confused. A picture is a function x[i,j]x[i, j] of its row ii and column jj. A two-dimensional convolution with a filter ww (here 3×3) computes

(x∗w)[i,j]=∑u=02∑v=02w[u,v] x[i+u, j+v](x * w)[i, j] = \sum_{u=0}^{2} \sum_{v=0}^{2} w[u, v] \, x[i + u, \, j + v]

where:

  • x[i+u,j+v]x[i + u, j + v] are the 9 pixels under the filter when its top-left corner sits at (i,j)(i, j);
  • w[u,v]w[u, v] are the 9 filter weights, the same at every position (that is weight sharing);
  • the sum is one output number; sliding over all (i,j)(i, j) gives the filtered picture.

Define the shift operator TtT_{t} that moves a picture tt pixels to the right:

(Tt x)[i,j]=x[i, j−t](T_{t}\, x)[i, j] = x[i, \, j - t]

Then equivariance and invariance of a function ff are

equivariance:f(Tt x)=Tt f(x)invariance:f(Tt x)=f(x)\text{equivariance:}\quad f(T_{t}\, x) = T_{t}\, f(x) \qquad\qquad \text{invariance:}\quad f(T_{t}\, x) = f(x)

In words: an equivariant ff lets the shift pass through (shift then filter equals filter then shift); an invariant ff swallows the shift. A convolution is equivariant because the same weights ww are used at every position: the sum at output position (i,j+t)(i, j + t) of the shifted picture sees exactly the pixels that the sum at (i,j)(i, j) saw in the original.

Equivariance with real numbers. We built a 12×12 toy picture with a bright 5×5 square in it, applied a 3×3 vertical-edge filter, then shifted the picture 2 pixels right and applied the same filter (vit_part1_math.py, part (b)):

python
img = torch.zeros(1, 1, 12, 12)
img[0, 0, 3:8, 2:7] = 1.0                                           # a bright 5 x 5 square at rows 3..7, columns 2..6
k = torch.tensor([[-1., 0., 1.], [-2., 0., 2.], [-1., 0., 1.]]).view(1, 1, 3, 3)   # a 3 x 3 vertical-edge filter (Sobel)
shift = 2
shifted = torch.zeros_like(img)
shifted[..., shift:] = img[..., :-shift]                             # the same square, 2 pixels to the right
out = Fn.conv2d(img, k)                                              # 10 x 10 output (no padding)
out_s = Fn.conv2d(shifted, k)
diff = (out_s[..., shift:] - out[..., :-shift]).abs().max().item()   # move the second output back by 2 and compare
plain text
    toy image (12, 12), bright square at rows 3-7, columns 2-6; filter (3, 3) (vertical edges)
    filter weights: [[-1.0, 0.0, 1.0], [-2.0, 0.0, 2.0], [-1.0, 0.0, 1.0]]
    output (10, 10); strongest response in the original at column 0, in the shifted at column 2
    max |shifted output moved back by 2 - original output| = 0.0
    row 5 of the output, original:    4    4    0    0    0   -4   -4    0    0    0
    row 5 of the output, shifted:     0    0    4    4    0    0    0   -4   -4    0
    the same filter weights (9 numbers) are used at every one of the 100 output positions (weight sharing)

Read the two printed rows. The filter fires +4 where the square's left edge is (dark to bright) and -4 where its right edge is (bright to dark). In the shifted picture the very same values appear, two columns to the right. Moving the second output back by two columns and subtracting the first gives a maximum difference of exactly 0.0. That is equivariance: the feature moved with the object, nothing else changed. The filter never had to learn anything about "the square at column 4" versus "the square at column 6"; one set of 9 weights handles every position.

the same 3 × 3 edge filter, applied to a square and to the square moved 2 pixels rightinput (12 × 12)row 3, column 2: 1row 3, column 3: 1row 3, column 4: 1row 3, column 5: 1row 3, column 6: 1row 4, column 2: 1row 4, column 3: 1row 4, column 4: 1row 4, column 5: 1row 4, column 6: 1row 5, column 2: 1row 5, column 3: 1row 5, column 4: 1row 5, column 5: 1row 5, column 6: 1row 6, column 2: 1row 6, column 3: 1row 6, column 4: 1row 6, column 5: 1row 6, column 6: 1row 7, column 2: 1row 7, column 3: 1row 7, column 4: 1row 7, column 5: 1row 7, column 6: 1convoutput (10 × 10)row 1, column 0: 1row 1, column 1: 1row 1, column 5: -1row 1, column 6: -1row 2, column 0: 3row 2, column 1: 3row 2, column 5: -3row 2, column 6: -3row 3, column 0: 4row 3, column 1: 4row 3, column 5: -4row 3, column 6: -4row 4, column 0: 4row 4, column 1: 4row 4, column 5: -4row 4, column 6: -4row 5, column 0: 4row 5, column 1: 4row 5, column 5: -4row 5, column 6: -4row 6, column 0: 3row 6, column 1: 3row 6, column 5: -3row 6, column 6: -3row 7, column 0: 1row 7, column 1: 1row 7, column 5: -1row 7, column 6: -1input shifted 2 px →row 3, column 4: 1row 3, column 5: 1row 3, column 6: 1row 3, column 7: 1row 3, column 8: 1row 4, column 4: 1row 4, column 5: 1row 4, column 6: 1row 4, column 7: 1row 4, column 8: 1row 5, column 4: 1row 5, column 5: 1row 5, column 6: 1row 5, column 7: 1row 5, column 8: 1row 6, column 4: 1row 6, column 5: 1row 6, column 6: 1row 6, column 7: 1row 6, column 8: 1row 7, column 4: 1row 7, column 5: 1row 7, column 6: 1row 7, column 7: 1row 7, column 8: 1convoutput shifted 2 px →row 1, column 2: 1row 1, column 3: 1row 1, column 7: -1row 1, column 8: -1row 2, column 2: 3row 2, column 3: 3row 2, column 7: -3row 2, column 8: -3row 3, column 2: 4row 3, column 3: 4row 3, column 7: -4row 3, column 8: -4row 4, column 2: 4row 4, column 3: 4row 4, column 7: -4row 4, column 8: -4row 5, column 2: 4row 5, column 3: 4row 5, column 7: -4row 5, column 8: -4row 6, column 2: 3row 6, column 3: 3row 6, column 7: -3row 6, column 8: -3row 7, column 2: 1row 7, column 3: 1row 7, column 7: -1row 7, column 8: -1positive response (dark-to-bright edge)negative response (bright-to-dark edge)max |shifted output moved back by 2 - original output| = 0.0: the feature moved with the object, nothing else changed
Translation equivariance made concrete. Left pair: the toy picture and the edge filter's output, which marks the square's left edge (positive) and right edge (negative). Right pair: the same picture moved 2 pixels to the right, and the output, which has moved by exactly 2 pixels and is otherwise identical. The maximum difference after shifting back is 0.0.

This is why a CNN can learn from fewer examples. A cat in the top-left corner teaches the same filters as a cat in the bottom-right. The network does not need to see cats everywhere; the design guarantees that what it learns in one place applies in every place.

Patch tokens are not shift-equivariant

Now the other side. ViT's first step cuts the picture along a fixed 16-pixel grid. Shift the picture by less than 16 pixels and every patch now contains a different mix of pixels, so every patch vector changes. Shift by exactly 16 and each patch's content moves, whole, into the neighbouring slot. We measured this on the real cat picture and the real ViT-B/16 patch-embedding layer, using cosine similarity to compare vectors:

cos⁡(a,b)=a⋅b∥a∥ ∥b∥\cos(\mathbf{a}, \mathbf{b}) = \frac{\mathbf{a} \cdot \mathbf{b}}{\lVert \mathbf{a} \rVert \, \lVert \mathbf{b} \rVert}

where a\mathbf{a} and b\mathbf{b} are two 768-number patch vectors, the dot is their dot product and ∥⋅∥\lVert \cdot \rVert is the length. The cosine is 1 for identical directions, 0 for unrelated ones, and negative for opposite ones. For each shift kk we compared patch (i,j)(i, j) of the original with patch (i,j+⌊k/16⌋)(i, j + \lfloor k/16 \rfloor) of the shifted picture, that is, we move the token grid back by the whole number of patches the picture moved (0 for shifts under 16, 1 for a shift of 16):

python
embed = vit.vit.embeddings.patch_embeddings                          # one 16 x 16 convolution with stride 16 = the linear patch projection
with torch.no_grad():
    ek = embed(shift_right(x, k))[0].view(14, 14, -1)
back = k // P                                                    # whole patches the picture moved
a, b = grid[:, :14 - back], ek[:, back:]                        # compare patch (i, j) of the original with patch (i, j + back) of the shifted
cos = Fn.cosine_similarity(a.reshape(-1, 768), b.reshape(-1, 768), dim=-1)
plain text
     shift k  grid moved back  mean cosine  min cosine   ViT-B/16 top-1             ResNet-50 top-1
           0                0        1.000       1.000   Egyptian cat 0.937         tiger cat 0.942
           1                0        0.935       0.643   Egyptian cat 0.888         tiger cat 0.939
           4                0        0.649      -0.219   Egyptian cat 0.861         tiger cat 0.847
           8                0        0.507      -0.353   Egyptian cat 0.939         tiger cat 0.826
          12                0        0.457      -0.446   Egyptian cat 0.847         tiger cat 0.607
          16                1        1.000       1.000   Egyptian cat 0.908         tiger cat 0.737
    for k = 16 compared at the SAME grid position (no moving back): mean cosine = 0.428  (each slot now holds its left neighbour)
0.4000.5000.6000.7000.8000.9001.00001481216patch cosine, ViT, 0: 1.000patch cosine, ViT, 1: 0.935patch cosine, ViT, 4: 0.649patch cosine, ViT, 8: 0.507patch cosine, ViT, 12: 0.457patch cosine, ViT, 16: 1.000P(Egyptian cat), ViT, 0: 0.937P(Egyptian cat), ViT, 1: 0.888P(Egyptian cat), ViT, 4: 0.861P(Egyptian cat), ViT, 8: 0.939P(Egyptian cat), ViT, 12: 0.847P(Egyptian cat), ViT, 16: 0.908P(tiger cat), ResNet, 0: 0.942P(tiger cat), ResNet, 1: 0.939P(tiger cat), ResNet, 4: 0.847P(tiger cat), ResNet, 8: 0.826P(tiger cat), ResNet, 12: 0.607P(tiger cat), ResNet, 16: 0.737patch cosine, ViT: 1.000P(Egyptian cat), ViT: 0.908P(tiger cat), ResNet: 0.737shift of the picture to the right, in pixelscosine similarity / probability
What a shift does to ViT's patch tokens, and to both models' answers. Green: the mean cosine similarity between the original and shifted patch vectors falls from 1.000 to 0.457 as the shift grows to 12 pixels, then jumps back to 1.000 at 16 pixels, when whole patches move one slot. Blue and orange: the probability of each model's top class moves around but the class never changes.

Two lessons sit in this table.

  • The patch tokens are not equivariant. A 1-pixel shift, invisible to a person, already changes the average patch vector's cosine to 0.935 and the worst one to 0.643. By 12 pixels the mean is 0.457 and some vectors point in the opposite direction. Only a shift of exactly one patch (16 pixels) restores a cosine of 1.000, and then only because we moved the token grid back by one slot. Compared at the same slot, the 16-pixel shift gives 0.428. A convolution has none of this: its output moved by exactly 2 pixels and changed by exactly 0.0.
  • ViT still gets the answer right. At every shift its top class stays "Egyptian cat" with probability between 0.847 and 0.939. The model was never told that shifts do not matter; it learned it, from millions of pictures in which cats sat in different places. That is the paper's argument in one table: what the CNN has by design, the Transformer can learn from data, if there is enough data. ResNet-50 also keeps "tiger cat" throughout, as expected for a CNN.

An honesty note on the ResNet column: its confidence drops to 0.607 at a 12-pixel shift and 0.737 at 16. A real ResNet is not perfectly equivariant either, because its strided layers and pooling sample the picture on coarse grids, and our shift fills the strip that enters on the left by repeating the edge column, which adds a small artificial band. The toy filter above is exactly equivariant; a whole ResNet is only approximately so. The qualitative point stands: both networks keep their answer under shifts that scramble every ViT patch vector.

Large scale training trumps inductive bias

The two datasets named here are the paper's fuel.

"14M-300M images" refers to these two: 14 million for ImageNet-21k, 303 million for JFT-300M. Against them, the 1.3 million pictures of ImageNet are "mid-sized", which is a startling thing to call the dataset that defined the field for a decade.

amount of pre-training data (more →)accuracy on the target taskResNet (CNN)ViTsmall data: the CNN's built-inassumptions help; ViT lagslarge data: ViT learns theassumptions, and more, from dataa sketch of the paper's claim, not measured data (the real numbers follow)
A sketch of the paper's claim. With little pre-training data, a convolutional network's built-in assumptions help, and it beats ViT. With a lot of data, ViT learns those assumptions, and more, from the pictures themselves, and overtakes. This is a drawing of the idea, not measured data; the measured numbers follow.

The paper proves the sketch in Section 4.3, which Part 4 reads in full. Two sentences from it are worth seeing now, because they are the evidence behind "the picture changes".

The exact numbers are in the paper's Table 5 (appendix), which Part 4 redraws in full. Here is the table, and the ImageNet rows for the two 16-pixel-patch models as a chart.

74%77%80%83%86%89%1.314303ViT-B/16, 1.3: 77.91%ViT-B/16, 14: 83.97%ViT-B/16, 303: 84.15%ViT-L/16, 1.3: 76.53%ViT-L/16, 14: 85.15%ViT-L/16, 303: 87.12%ViT-L/16: 87.12%ViT-B/16: 84.15%pre-training images (millions, log scale): ImageNet 1.3M, ImageNet-21k 14M, JFT-300M 303MImageNet top-1 accuracy (%)
The ImageNet rows of Table 5 for ViT-B/16 and ViT-L/16, against the size of the pre-training set on a log scale. With 1.3 million pictures the smaller model wins (77.91 against 76.53). With 14 million the larger model is ahead (85.15 against 83.97), and with 303 million it is far ahead (87.12 against 84.15).

Now the four headline numbers. They all belong to the paper's biggest model, ViT-H/14 (632 million parameters, 14-pixel patches), pre-trained on JFT-300M.

top-1 accuracy of the best model (ViT-H/14 pre-trained on JFT-300M)ImageNetImageNet: 88.55%88.55%ImageNet-ReaLImageNet-ReaL: 90.72%90.72%CIFAR-100CIFAR-100: 94.55%94.55%VTAB (19 tasks)VTAB (19 tasks): 77.63%77.63%ImageNet-ReaL: ImageNet with cleaner labels. VTAB: the average over 19 small tasks with 1,000 training images each.
The four headline results of the introduction: 88.55% top-1 on ImageNet, 90.72% on ImageNet-ReaL (ImageNet with cleaner labels), 94.55% on CIFAR-100 and 77.63% averaged over the 19 tasks of VTAB. All from ViT-H/14 pre-trained on JFT-300M.
BenchmarkWhat it testsViT-H/14 (JFT)Note
ImageNet1,000-class photo classification, 50,000 test pictures88.55%the standard number everyone compares
ImageNet-ReaLthe same pictures with corrected labels (Beyer et al., 2020)90.72%higher because fewer "wrong" labels in the test set
CIFAR-100100 classes of tiny 32×32 pictures94.55%shows transfer to a very different kind of picture
VTAB (19 tasks)19 varied tasks with 1,000 training examples each77.63%shows transfer with very little data

These numbers are compared with the best CNNs in Table 2 of the paper (Part 4). The short version: ViT-H/14 beat the Big Transfer ResNet and Noisy Student on every one of these benchmarks, and did so with less pre-training compute. Part 4 goes through every row and every error bar.

Who's who: the papers the introduction is talking to

The introduction cites sixteen earlier works in five paragraphs. Here they are in one place, so the rest of the series can refer back to them. Each is listed in full in the references at the end of this part.

Short namePaperYearWhat it is, in one lineRole in the ViT paper
TransformerVaswani et al., Attention Is All You Need2017the attention-only network for translationViT is its encoder, unchanged
BERTDevlin et al., BERT2018pre-train a Transformer encoder, fine-tune per taskthe recipe, the model sizes and the class token
GPT-3Brown et al., Language Models are Few-Shot Learners2020a 175-billion-parameter language model"over 100B parameters"
GShardLepikhin et al., GShard2020a 600-billion-parameter translation model"over 100B parameters"
LeNetLeCun et al., Backpropagation Applied to Handwritten Zip Code Recognition1989the first practical CNN"convolutional architectures remain dominant"
AlexNetKrizhevsky et al., ImageNet Classification with Deep CNNs2012the deep CNN that won ImageNet 2012start of the deep-learning era in vision
ResNetHe et al., Deep Residual Learning2015CNNs of 100+ layersthe "ResNet-like architectures" ViT must beat
Non-localWang et al., Non-local Neural Networks2018attention blocks inside a CNN, for videoCNN + attention
DETRCarion et al., End-to-End Object Detection with Transformers2020a Transformer on top of a CNN, for detectionCNN + attention
Stand-aloneRamachandran et al., Stand-Alone Self-Attention in Vision Models2019local attention replacing every convolution"replacing the convolutions entirely"
Axial-DeepLabWang et al. (2020a), Axial-DeepLab2020attention along rows, then along columns"replacing the convolutions entirely"
Instagram pre-trainingMahajan et al., Exploring the Limits of Weakly Supervised Pretraining2018CNNs pre-trained on 3.5 billion Instagram picturesResNets are still state of the art
Noisy StudentXie et al., Self-training with Noisy Student2020CNN self-training with 300M unlabelled picturesResNets are still state of the art
BiTKolesnikov et al., Big Transfer2020very large ResNets pre-trained on ImageNet-21k and JFTthe main baseline, by the same team
ImageNetDeng et al., ImageNet: A Large-Scale Hierarchical Image Database2009the 14-million-picture collection behind ImageNet-1k and 21kthe public pre-training set
JFTSun et al., Revisiting Unreasonable Effectiveness of Data2017Google's 300-million-picture internal datasetthe private pre-training set

Next, in Part 2: the related work the paper builds on (local attention, Sparse Transformers, iGPT, and the 2×2-patch model of Cordonnier et al.), and then the model itself: Figure 1 block by block, Equations 1 to 4 with the real numbers of ViT-B/16 (the patch embedding, the class token, the position embeddings, one encoder layer), the Appendix A formulas, and where the 86 million parameters live.

Run it yourself

The two scripts behind this part are code/papers/vit/vit_part1.py (the cat picture through ViT-B/16 and ResNet-50) and code/papers/vit/vit_part1_math.py (the quadratic cost table, the equivariance check on a toy convolution, and the patch-shift experiment). Both run on a laptop CPU in under a minute each and download google/vit-base-patch16-224 (about 350 MB) and microsoft/resnet-50 (about 100 MB) the first time.

bash
pip install torch transformers datasets pillow
python vit_part1.py         # shapes, patch arithmetic, top-5 of both models; writes results/part1.json
python vit_part1_math.py    # quadratic cost, equivariance, patch shifts; writes results/part1_math.json
Terminal output of vit_part1.py: the picture size, the ViT-B/16 tensor shapes and patch arithmetic, its top-5 classes led by Egyptian cat 0.9374, the ResNet-50 feature-map shapes and its top-5 classes led by tiger cat 0.9416
The real output of vit_part1.py.
Terminal output of vit_part1_math.py: the table of attention weights for pixels versus patches, the toy convolution equivariance check with a maximum difference of 0.0, and the patch-shift table with mean cosine similarities and both models' top-1 classes for shifts of 0, 1, 4, 8, 12 and 16 pixels
The real output of vit_part1_math.py.

References

The ViT paper

  1. A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT). ICLR 2021 (OpenReview). arXiv:2010.11929, version 2 (June 2021), which is the version shown in the screenshots.
  2. Google Research. Vision Transformer code and pre-trained models. GitHub, 2020.

Papers the ViT paper cites in this part

  1. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin. Attention Is All You Need (the Transformer). NeurIPS 2017.
  2. J. Devlin, M.-W. Chang, K. Lee, K. Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (BERT). NAACL 2019.
  3. T. B. Brown et al. Language Models are Few-Shot Learners (GPT-3). NeurIPS 2020.
  4. D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, Z. Chen. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding (GShard). ICLR 2021.
  5. Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, L. D. Jackel. Backpropagation Applied to Handwritten Zip Code Recognition. Neural Computation 1(4), 1989.
  6. A. Krizhevsky, I. Sutskever, G. E. Hinton. ImageNet Classification with Deep Convolutional Neural Networks (AlexNet). NeurIPS 2012.
  7. K. He, X. Zhang, S. Ren, J. Sun. Deep Residual Learning for Image Recognition (ResNet). CVPR 2016.
  8. X. Wang, R. Girshick, A. Gupta, K. He. Non-local Neural Networks. CVPR 2018.
  9. N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, S. Zagoruyko. End-to-End Object Detection with Transformers (DETR). ECCV 2020.
  10. P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, J. Shlens. Stand-Alone Self-Attention in Vision Models. NeurIPS 2019.
  11. H. Wang, Y. Zhu, B. Green, H. Adam, A. Yuille, L.-C. Chen. Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation ("Wang et al., 2020a" in the paper). ECCV 2020.
  12. D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, L. van der Maaten. Exploring the Limits of Weakly Supervised Pretraining. ECCV 2018.
  13. Q. Xie, M.-T. Luong, E. Hovy, Q. V. Le. Self-training with Noisy Student improves ImageNet classification (Noisy Student). CVPR 2020.
  14. A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, N. Houlsby. Big Transfer (BiT): General Visual Representation Learning (BiT). ECCV 2020.
  15. J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. CVPR 2009.
  16. C. Sun, A. Shrivastava, S. Singh, A. Gupta. Revisiting Unreasonable Effectiveness of Data in Deep Learning Era (JFT-300M). ICCV 2017.

Other sources used in this part

  1. X. Zhai, J. Puigcerver, A. Kolesnikov, et al. A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark (VTAB). arXiv 2019. The benchmark behind the 77.63% number.
  2. L. Beyer, O. J. Hénaff, A. Kolesnikov, X. Zhai, A. van den Oord. Are we done with ImageNet? (ImageNet-ReaL). arXiv 2020. The cleaned-up labels behind the 90.72% number.
  3. A. Krizhevsky. The CIFAR-10 and CIFAR-100 datasets. University of Toronto, 2009.
  4. Hugging Face model cards: google/vit-base-patch16-224 and microsoft/resnet-50; the sample picture is the huggingface/cats-image dataset.
  5. Code for this part: vit_part1.py and vit_part1_math.py.