Research papers

  • ViT, explained

    The 2020 paper that cut an image into 16×16 patches, fed them to an unchanged Transformer, and beat the best convolutional networks once it was given enough data. Explained line by line in six parts, with the paper's own text, simple pictures, worked equations and real code on the released models.

  • BERT, explained

    The 2018 paper that taught one model to read text in both directions at once, then reuse that skill for almost any language task. Explained line by line in six parts, with the paper's own text, simple pictures and real code.