- CLIP, the classic image-text alignment model that uses contrastive learning to embed images and captions into a shared space.
- LLaVA, which combines a vision encoder (CLIP ViT) with a large language model (Llama 2, Vicuna) for instruction following, dialogue, and vision-language reasoning.
In this chapter
The chapter has three parts. Representation learning covers how a network learns an embedding that exposes the factors of a scene, ending with CLIP, which learns one embedding for images and text together. Vision-language models then feed such visual embeddings into a large language model, and a fine-tuning lab adapts a small one to a task.Representation learning
Learning Representations from Pixels
Autoencoders, k-means, and contrastive learning, where the chosen augmentation decides what the embedding keeps.
CLIP
Contrastive Language-Image Pretraining, the foundation of modern vision-language models.
CLIP Zero-Shot Classification
Hands-on section: zero-shot classification as a linear classifier whose weights come from text prompts.
Vision-language models
LLaVA
Large Language and Vision Assistant, instruction-tuned multimodal dialogue.
LLaVA deep dive
A closer look at the LLaVA family: architecture, training stages, and data.
Fine-tuning lab
SmolVLM SFT for food extraction
Supervised fine-tuning of a small vision-language model to extract structured data from food images.
References
- Johnson, J., Hariharan, B., Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., et al. (2016). CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning. arXiv [cs.CV].
- Lu, J., Xiong, C., Parikh, D., Socher, R. (2016). Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning. arXiv [cs.CV].
- Vinyals, O., Toshev, A., Bengio, S., Erhan, D. (2016). Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge. arXiv [cs.CV].
- Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., et al. (2015). Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. arXiv [cs.LG].

