| ‣ | | 2022 | OpenCLIP: Reproducible scaling laws for contrastive language-image learning Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, Jenia Jitsev | https://arxiv.org/abs/2212.07143 | | | 2022 | Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer. High-Resolution Image Synthesis With Latent Diffusion Models | https://arxiv.org/abs/2112.10752 | ‣
|
| Year | Link | Description |
|---|---|---|
| 2020 | https://arxiv.org/abs/2005.12872 | DETR |
| 2020 | https://arxiv.org/abs/2002.05709 | SimCLR paper |
| 2020 | https://arxiv.org/abs/2006.07733 | BYOL paper |
| 2021 | https://arxiv.org/abs/2103.03206 | DeepMind. Different type of encoder. |
| 2022 | https://arxiv.org/abs/2204.14198 | First VLM paper. Deepmind. |
| 2023 | https://arxiv.org/abs/2304.08485 | Open VLM |
| 2023 | https://arxiv.org/abs/2305.17409 | Pruning / training |
| 2023 | https://arxiv.org/abs/2303.15343 | SigLip paper Google Deepmind. Lucas Beyer |
| 2024 | https://arxiv.org/abs/2311.06242 | Florence 2 A foundation model for vision. All tasks possible. Microsoft. |
| 2024 | https://arxiv.org/abs/2407.21783 | LLama 3 paper Meta |
| 2024 | https://arxiv.org/abs/2407.01449 | ColPali uses Late Interaction to retrieve all information from images including text and graphs. |
| 2025 | https://arxiv.org/abs/2504.05299 | Blog: https://huggingface.co/papers/2504.05299 |
| 2025 | https://arxiv.org/abs/2406.11251 | Multimodal - text, graphs, images |
| 2025 | https://arxiv.org/abs/2507.19468 | DINO paper, Meta |
| 2025 | https://arxiv.org/abs/2508.10104 | Dino V3 Meta |
| 2025 | https://www.nature.com/articles/s41586-025-09631-6 | Google Deepmind Aligning machine and human visual representations across abstraction levels |
| 2026 | https://arxiv.org/abs/2608.02980 | |
| 2026 | https://arxiv.org/abs/2601.09661 |
| Year | Title/author | Link | Description |
|---|---|---|---|
| Deep Learning- Yoshua Bengio | https://www.deeplearningbook.org/contents/convnets.html | CNN’s in Yoshua Bengio’s Deep Learning book | |
| MIT Vision boek | https://visionbook.mit.edu/ | ||
| MIT vision course | https://introtocv.github.io/materials.html | ||
| Stanford CV course | https://cs231n.stanford.edu/ |