1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do multimodal LLMs process images? Describe the typical architecture.
30-second answerSay your answer out loud first, then reveal.

Design choices
- Image resolution and tiling: high-resolution images (documents, screenshots) are split into tiles, plus a low-resolution global view. More visual tokens give better detail at higher cost (an image can be hundreds to thousands of tokens).
- Projector type:
MLP (LLaVA-style): simple, keeps all patch tokens.
Resampler / Q-Former (Flamingo, BLIP-2): compresses to a fixed number of tokens. - Fusion: early fusion (visual tokens in the input sequence) vs cross-attention layers (Flamingo-style).
- Training stages: (a) pretrain the projector on image–caption pairs with the encoder and LLM frozen; (b) visual instruction tuning (VQA, OCR, charts, grounding); (c) preference / RL tuning.
Practical considerations
- Image token costs and latency (resize images sensibly).
- OCR-heavy tasks (invoices, forms) need high resolution; check model support.
- Hallucination about image contents ("object hallucination") needs evaluation.
- Audio and video follow similar patterns: a modality encoder + projector + LLM, or native multimodal training.
Related
This is what real progress feels like.