LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

Content blocks: images and files in a message

Content blocks are the typed parts of a message, so one message can carry text alongside an image, a file or a PDF instead of a plain string.

Last updated: 27 Sep, 2026 · LangChain 1.4

A message with text and images together · from the Step By Step Process To Build MultiModal RAG With LangChain (PDF And Images) · 40:25 to 44:49

A question with pictures in it

A message's content does not have to be one string. It can be a list of parts, each with a type, so text and pictures travel in one message. The video builds such a message in a multimodal RAG project: it retrieves text and images from a PDF and sends them together to a GPT-4 vision model, one that can read images.

Its create_multimodal_message function starts the content list with a text part holding the question. It then separates the retrieved documents into text and images: the text goes into one text part as context, and each image gets a label and an image_url part carrying the image data. The function returns the list inside a HumanMessage. Stripped down to its shape, with the image as base64 text inside a data URL:

python
from langchain.messages import HumanMessage

message = HumanMessage(content=[
    {"type": "text", "text": "Question: what does the chart on page 1 show about revenue trends?"},
    {"type": "text", "text": "[Image from page 1]:"},
    {"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}},
    {"type": "text", "text": "\n\nPlease answer the question based on the provided text and images."},
])

The model reads the parts in order: the question, a label, the picture the label names, and a closing instruction. The video's pipeline passes the retrieved documents to this function and sends the message with llm.invoke. Asked what the chart on page one shows about revenue trends, the model answered from the picture: revenue rose steadily over three quarters, and Q1, the first blue bar, was the lowest. The shape of each part comes from OpenAI's documentation for its multimodal models, which is why the video's code uses OpenAI's image_url format. LangChain's content_blocks describe the same parts in one format every provider understands.

Building a message with parts

Pass content_blocks a list of typed parts. Each part names its type and its data.

python
from langchain.messages import HumanMessage

msg = HumanMessage(content_blocks=[
    {"type": "text", "text": "What is in this picture?"},
    {"type": "image", "url": "https://example.com/cat.png"},
])

A text-and-image message end to end

The whole snippet, printing the block types the message carries.

Example
from langchain.messages import HumanMessage

msg = HumanMessage(content_blocks=[
    {"type": "text", "text": "What is in this picture?"},
    {"type": "image", "url": "https://example.com/cat.png"},
])
print([b["type"] for b in msg.content_blocks])

What the message holds

  • The message carries two parts: a text block and an image block.
  • content_blocks gives every part a type, so a model that accepts images reads them.
  • A plain-string message is the one-block case: only text.

Sending a local image

A picture on your disk goes in as base64 text with its type, instead of a URL. Here shape.png is a red circle on a white background; use any PNG you have:

Example
import base64

from langchain.messages import HumanMessage

with open("shape.png", "rb") as file:
    data = base64.b64encode(file.read()).decode()

msg = HumanMessage(content_blocks=[
    {"type": "text", "text": "What is in this picture?"},
    {"type": "image", "base64": data, "mime_type": "image/png"},
])
print([b["type"] for b in msg.content_blocks])

openai/gpt-oss-120b, the course model, reads text only. To get an answer about the picture, send the message to a vision model such as Gemini, with a free Gemini key set as GOOGLE_API_KEY (from aistudio.google.com/apikey) and pip install "langchain-google-genai==4.4.0". Add these lines to the end of the same file:

ExampleAPI key
from langchain.chat_models import init_chat_model

vision = init_chat_model("google_genai:gemini-2.5-flash")
print(vision.invoke([msg]).text)

Gemini saw more than a red circle: it recognised the shape as the flag of Japan. A vision model interprets a picture the way a language model interprets text, which is useful, and is also why its answer about an image deserves the same checking as any other answer.

Plain string vs content blocks

content="..."content_blocks=[...]
CarriesText onlyText, images, files, audio
Use forAn ordinary turnAsking about a picture or a document
Model supportEvery chat modelModels that accept that media

When a message needs media

  • Asking a vision model what is in an image.
  • Sending a PDF or a file for the model to read.
  • Mixing an instruction with the media it refers to.
Watch out. A model only reads a block type it supports. Send an image to a text-only model and the image is ignored or rejected; check the model accepts that media first.
Try it yourself
  • Add a third block for a file with a url and print the types.
  • Send a message with an image URL to the Gemini model above and read its answer.
  • Add a second text block after the image and check the order of the types printed.

Slow is fine. Stopping is the only problem.