Canonical CV papers

Year Author/Title Link Notes
1959 Hubel and Wiesel. This paper analyses the operation of the Visual Cortex that later inspire and inform the design of convolutional neural networks. They win the Nobel Prize in 1981.
1989 Yann Le Cun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard and L. D. Jackel, Handwritten Digit Recognition with a Back-Propagation Network https://proceedings.neurips.cc/paper/1989/file/53c3bce66e43be4f209556518c2fcb54-Paper.pdf
1997 Normalized Cuts and Image Segmentation, Jianbo Shi & Jitendra Malik https://people.eecs.berkeley.edu/~malik/papers/shi-malik97.pdf
2012 Alex Krizhevsky, Ilya Sutskever and Geoffrey E. Hinton ImageNet Classification with Deep Convolutional Neural Networks https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
2013 Diederik P. Kingma and Max Welling, *Auto-Encoding Variational Bayes “VAE paper”* https://arxiv.org/abs/1312.6114
2014 Ross Girshick, Jeff Donahue, Trevor Darrell and Jitendra Malik, *Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation R-CNN paper* https://www.cv-foundation.org/openaccess/content_cvpr_2014/papers/Girshick_Rich_Feature_Hierarchies_2014_CVPR_paper.pdf
2014 Dropout: A Simple Way to Prevent Neural Networks from
Overfitting *Nitish Srivastava
Geoffrey Hinton, Alex Krizhevsky ,Ilya Sutskever,Ruslan Salakhutdinov* https://jmlr.csail.mit.edu/papers/volume15/srivastava14a/srivastava14a.pdf
2015 Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun, Deep Residual Learning for Image Recognition https://ieeexplore.ieee.org/document/7780459
2015 Oriol Vinyals, Alexander Toshev, Samy Bengio and Dumitru Erhan, Show and Tell: A Neural Image Caption Generator https://arxiv.org/abs/1411.4555
2015 Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy, Li Fei-Fei https://arxiv.org/pdf/1412.2306
2016 Layer NormalizationJimmy Lei BaJamie Ryan KirosGeoffrey E. Hinton https://arxiv.org/abs/1607.06450
2016 Grad-CAM visualisations https://openaccess.thecvf.com/content_ICCV_2017/papers/Selvaraju_Grad-CAM_Visual_Explanations_ICCV_2017_paper.pdf https://github.com/ramprs/grad-cam/
2020 SimCLR: A simple framework for contrastive learning of visual representations https://arxiv.org/abs/2002.05709
2020 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale https://arxiv.org/abs/2010.11929
2021 Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger and Ilya Sutskever, ***Learning Transferable Visual Models From Natural Language Supervision. CLIP*** https://arxiv.org/abs/2103.00020

| ‣ | | 2022 | OpenCLIP: Reproducible scaling laws for contrastive language-image learning Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, Jenia Jitsev | https://arxiv.org/abs/2212.07143 | | | 2022 | Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer. High-Resolution Image Synthesis With Latent Diffusion Models | https://arxiv.org/abs/2112.10752 | ‣

|

Recent Computer Vision papers

Year Link Description
2020 https://arxiv.org/abs/2005.12872 DETR
2020 https://arxiv.org/abs/2002.05709 SimCLR paper
2020 https://arxiv.org/abs/2006.07733 BYOL paper
2021 https://arxiv.org/abs/2103.03206 DeepMind. Different type of encoder.
2022 https://arxiv.org/abs/2204.14198 First VLM paper. Deepmind.
2023 https://arxiv.org/abs/2304.08485 Open VLM
2023 https://arxiv.org/abs/2305.17409 Pruning / training
2023 https://arxiv.org/abs/2303.15343 SigLip paper Google Deepmind. Lucas Beyer
2024 https://arxiv.org/abs/2311.06242 Florence 2 A foundation model for vision. All tasks possible. Microsoft.
2024 https://arxiv.org/abs/2407.21783 LLama 3 paper Meta
2024 https://arxiv.org/abs/2407.01449 ColPali uses Late Interaction to retrieve all information from images including text and graphs.
2025 https://arxiv.org/abs/2504.05299 Blog: https://huggingface.co/papers/2504.05299
2025 https://arxiv.org/abs/2406.11251 Multimodal - text, graphs, images
2025 https://arxiv.org/abs/2507.19468 DINO paper, Meta
2025 https://arxiv.org/abs/2508.10104 Dino V3 Meta
2025 https://www.nature.com/articles/s41586-025-09631-6 Google Deepmind Aligning machine and human visual representations across abstraction levels
2026 https://arxiv.org/abs/2608.02980
2026 https://arxiv.org/abs/2601.09661

Computer Vision free text books & courses


Year Title/author Link Description
Deep Learning- Yoshua Bengio https://www.deeplearningbook.org/contents/convnets.html CNN’s in Yoshua Bengio’s Deep Learning book
MIT Vision boek https://visionbook.mit.edu/
MIT vision course https://introtocv.github.io/materials.html
Stanford CV course https://cs231n.stanford.edu/