Year Author/Title Link Notes
1959 Hubel and Wiesel. This paper analyses the operation of the Visual Cortex that later inspire and inform the design of convolutional neural networks. They win the Nobel Prize in 1981.
1989 Yann Le Cun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard and L. D. Jackel, Handwritten Digit Recognition with a Back-Propagation Network https://proceedings.neurips.cc/paper/1989/file/53c3bce66e43be4f209556518c2fcb54-Paper.pdf
1997 Normalized Cuts and Image Segmentation, Jianbo Shi & Jitendra Malik https://people.eecs.berkeley.edu/~malik/papers/shi-malik97.pdf
2012 Alex Krizhevsky, Ilya Sutskever and Geoffrey E. Hinton ImageNet Classification with Deep Convolutional Neural Networks https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
2013 Diederik P. Kingma and Max Welling, *Auto-Encoding Variational Bayes “VAE paper”* https://arxiv.org/abs/1312.6114
2014 Ross Girshick, Jeff Donahue, Trevor Darrell and Jitendra Malik, *Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation R-CNN paper* https://www.cv-foundation.org/openaccess/content_cvpr_2014/papers/Girshick_Rich_Feature_Hierarchies_2014_CVPR_paper.pdf
2014 Dropout: A Simple Way to Prevent Neural Networks from
Overfitting *Nitish Srivastava
Geoffrey Hinton, Alex Krizhevsky ,Ilya Sutskever,Ruslan Salakhutdinov* https://jmlr.csail.mit.edu/papers/volume15/srivastava14a/srivastava14a.pdf
2015 Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun, Deep Residual Learning for Image Recognition https://ieeexplore.ieee.org/document/7780459
2015 Oriol Vinyals, Alexander Toshev, Samy Bengio and Dumitru Erhan, Show and Tell: A Neural Image Caption Generator https://arxiv.org/abs/1411.4555
2015 Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy, Li Fei-Fei https://arxiv.org/pdf/1412.2306
2016 Layer NormalizationJimmy Lei BaJamie Ryan KirosGeoffrey E. Hinton https://arxiv.org/abs/1607.06450
2016 Grad-CAM visualisations https://openaccess.thecvf.com/content_ICCV_2017/papers/Selvaraju_Grad-CAM_Visual_Explanations_ICCV_2017_paper.pdf https://github.com/ramprs/grad-cam/
2020 SimCLR: A simple framework for contrastive learning of visual representations https://arxiv.org/abs/2002.05709
2020 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale https://arxiv.org/abs/2010.11929
2021 Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger and Ilya Sutskever, ***Learning Transferable Visual Models From Natural Language Supervision. CLIP*** https://arxiv.org/abs/2103.00020

| ‣ | | 2022 | OpenCLIP: Reproducible scaling laws for contrastive language-image learning Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, Jenia Jitsev | https://arxiv.org/abs/2212.07143 | | | 2022 | Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer. High-Resolution Image Synthesis With Latent Diffusion Models | https://arxiv.org/abs/2112.10752 | ‣

| | | | | |

Recent Vision papers

Year Name Link Description
2020 DETR: End-to-End Object Detection with TransformersNicolas CarionFrancisco MassaGabriel SynnaeveNicolas UsunierAlexander KirillovSergey Zagoruyko https://arxiv.org/abs/2005.12872
2020 A Simple Framework for Contrastive Learning of Visual RepresentationsTing ChenSimon KornblithMohammad NorouziGeoffrey Hinton https://arxiv.org/abs/2002.05709 SimCLR paper
2020 https://arxiv.org/abs/2006.07733 https://arxiv.org/abs/2006.07733 BYOL paper
2021 Perceiver: General Perception with Iterative Attention
Andrew Jaegle, Felix Gimeno Andrew Brock Andrew Zisserman Oriol Vinyals 1 Joao Carreira https://arxiv.org/abs/2103.03206 DeepMind. Different type of encoder.
2022 Flamingo: a Visual Language Model for Few-Shot Learning https://arxiv.org/pdf/2204.14198 First VLM paper. Deepmind.
2023 LLaVA: Visual Instruction Tuning. Haotian LiuChunyuan LiQingyang WuYong Jae Lee https://arxiv.org/abs/2304.08485 Open VLM
2023 On the special role of class-selective neurons in early training https://arxiv.org/abs/2305.17409 Pruning / training
2023 Sigmoid Loss for Language Image Pre-TrainingXiaohua ZhaiBasil MustafaAlexander KolesnikovLucas Beyer https://arxiv.org/abs/2303.15343 SigLip paper Google Deepmind. Lucas Beyer
2024 Florence 2 https://arxiv.org/abs/2311.06242 A foundation model for vision. All tasks possible. Microsoft.
2024 LLama 3 paper https://arxiv.org/abs/2407.21783 Describes the multimodal llama 3 herd of models. Meta
2024 ColPali: Efficient Document Retrieval with Vision Language Models https://arxiv.org/abs/2407.01449 ColPali uses Late Interaction to retrieve all information from images including text and graphs. (needs GPU)
2025 SmolVLM https://arxiv.org/abs/2504.05299 Blog: https://huggingface.co/papers/2504.05299
2025 Unifying Multimodal Retrieval via Document Screenshot Embedding https://arxiv.org/abs/2406.11251 Multimodal - text, graphs, images
2025 **Back to the Features: DINO as a
Foundation for Video World Models**
Federico Baldassarre Marc Szafraniec Basile Terver Vasil Khalidov Francisco Massa, Yann LeCun Patrick Labatut Maximilian Seitzer Piotr Bojanowski https://arxiv.org/pdf/2507.19468 Meta
2025 DinoV3 - Simeoni et all https://arxiv.org/abs/2508.10104 Meta
2025 Aligning machine and human visual representations across abstraction levelsLukas MuttenthalerKlaus GreffFrieda BornBernhard SpitzerSimon KornblithMichael C. MozerKlaus-Robert MüllerThomas Unterthiner & Andrew K. Lampinen https://www.nature.com/articles/s41586-025-09631-6 Google Deepmind

Computer Vision free text books & courses


Year Title/author Link Description
Deep Learning- Yoshua Bengio https://www.deeplearningbook.org/contents/convnets.html CNN’s in Yoshua Bengio’s Deep Learning book
MIT Vision boek https://visionbook.mit.edu/
MIT vision course https://introtocv.github.io/materials.html
Stanford CV course https://cs231n.stanford.edu/