Content blocks: images and files in a message
Content blocks are the typed parts of a message, so one message can carry text alongside an image, a file or a PDF instead of a plain string.
Last updated: 27 Sep, 2026 · LangChain 1.4
A question with pictures in it
A message's content does not have to be one string. It can be a list of parts, each with a type, so text and pictures travel in one message. The video builds such a message in a multimodal RAG project: it retrieves text and images from a PDF and sends them together to a GPT-4 vision model, one that can read images.
Its create_multimodal_message function starts the content list with a text part holding the question. It then separates the retrieved documents into text and images: the text goes into one text part as context, and each image gets a label and an image_url part carrying the image data. The function returns the list inside a HumanMessage. Stripped down to its shape, with the image as base64 text inside a data URL:
from langchain.messages import HumanMessage
message = HumanMessage(content=[
{"type": "text", "text": "Question: what does the chart on page 1 show about revenue trends?"},
{"type": "text", "text": "[Image from page 1]:"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}},
{"type": "text", "text": "\n\nPlease answer the question based on the provided text and images."},
])The model reads the parts in order: the question, a label, the picture the label names, and a closing instruction. The video's pipeline passes the retrieved documents to this function and sends the message with llm.invoke. Asked what the chart on page one shows about revenue trends, the model answered from the picture: revenue rose steadily over three quarters, and Q1, the first blue bar, was the lowest. The shape of each part comes from OpenAI's documentation for its multimodal models, which is why the video's code uses OpenAI's image_url format. LangChain's content_blocks describe the same parts in one format every provider understands.
Building a message with parts
Pass content_blocks a list of typed parts. Each part names its type and its data.
from langchain.messages import HumanMessage
msg = HumanMessage(content_blocks=[
{"type": "text", "text": "What is in this picture?"},
{"type": "image", "url": "https://example.com/cat.png"},
])A text-and-image message end to end
The whole snippet, printing the block types the message carries.
from langchain.messages import HumanMessage
msg = HumanMessage(content_blocks=[
{"type": "text", "text": "What is in this picture?"},
{"type": "image", "url": "https://example.com/cat.png"},
])
print([b["type"] for b in msg.content_blocks])['text', 'image']
What the message holds
- The message carries two parts: a
textblock and animageblock. content_blocksgives every part atype, so a model that accepts images reads them.- A plain-string message is the one-block case: only text.
Sending a local image
A picture on your disk goes in as base64 text with its type, instead of a URL. Here shape.png is a red circle on a white background; use any PNG you have:
import base64
from langchain.messages import HumanMessage
with open("shape.png", "rb") as file:
data = base64.b64encode(file.read()).decode()
msg = HumanMessage(content_blocks=[
{"type": "text", "text": "What is in this picture?"},
{"type": "image", "base64": data, "mime_type": "image/png"},
])
print([b["type"] for b in msg.content_blocks])['text', 'image']
openai/gpt-oss-120b, the course model, reads text only. To get an answer about the picture, send the message to a vision model such as Gemini, with a free Gemini key set as GOOGLE_API_KEY (from aistudio.google.com/apikey) and pip install "langchain-google-genai==4.4.0". Add these lines to the end of the same file:
from langchain.chat_models import init_chat_model
vision = init_chat_model("google_genai:gemini-2.5-flash")
print(vision.invoke([msg]).text)This is the **flag of Japan**. It features a red circle (representing the sun) on a white background.
Gemini saw more than a red circle: it recognised the shape as the flag of Japan. A vision model interprets a picture the way a language model interprets text, which is useful, and is also why its answer about an image deserves the same checking as any other answer.
Plain string vs content blocks
| content="..." | content_blocks=[...] | |
|---|---|---|
| Carries | Text only | Text, images, files, audio |
| Use for | An ordinary turn | Asking about a picture or a document |
| Model support | Every chat model | Models that accept that media |
When a message needs media
- Asking a vision model what is in an image.
- Sending a PDF or a file for the model to read.
- Mixing an instruction with the media it refers to.
Related
- Previous: Messages: a conversation as a list
- Next: BaseChatModel: a model of your own
- Reference: LangChain docs
- Add a third block for a file with a
urland print the types. - Send a message with an image URL to the Gemini model above and read its answer.
- Add a second
textblock after the image and check the order of the types printed.
Slow is fine. Stopping is the only problem.