IngestionPipeline: load, split and embed once
IngestionPipeline is a reusable set of steps that loads, splits and embeds documents, and with a document store it skips any document that has not changed.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
The last lessons ran the reader, the splitter and the embedder as separate calls. A pipeline names those steps once and remembers what it has processed, so a re-run does no work on unchanged files.
Naming the steps
transformations run in order on the documents: the splitter makes nodes, the embedding model adds a vector to each. The document store remembers which documents the pipeline has seen and what their content was:
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.storage.docstore import SimpleDocumentStore
pipeline = IngestionPipeline(
transformations=[SentenceSplitter(chunk_size=80, chunk_overlap=0), Settings.embed_model],
docstore=SimpleDocumentStore(),
)View the code here
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
Running the pipeline twice
from llama_index.core.ingestion import IngestionPipeline
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.storage.docstore import SimpleDocumentStore
pipeline = IngestionPipeline(
transformations=[SentenceSplitter(chunk_size=80, chunk_overlap=0), Settings.embed_model],
docstore=SimpleDocumentStore(),
)
first = pipeline.run(documents=documents)
print("first run:", len(first), "nodes, embedded:", first[0].embedding is not None)
second = pipeline.run(documents=documents)
print("second run:", len(second), "nodes")first run: 6 nodes, embedded: True second run: 0 nodes
Reading the two runs
- The first run returns six nodes, each with an embedding, the same split as the chunking lesson.
- The second run returns nothing: every document was unchanged, so nothing was split or embedded again.
- With a hosted embedding model that skipped work is money saved on every re-run.
Editing a file and re-running
Append a line to one file, reload with a stable id, and run again. filename_as_id=True gives each document an id based on its file name:
with open("help/delivery.md", "a") as f:
f.write("\nOrders placed on a Sunday are sent on Monday.\n")
updated = SimpleDirectoryReader("help", file_metadata=only_name, filename_as_id=True).load_data()Re-running after an edit
from llama_index.core import SimpleDirectoryReader
with open("help/delivery.md", "a") as f:
f.write("\nOrders placed on a Sunday are sent on Monday.\n")
updated = SimpleDirectoryReader("help", file_metadata=only_name, filename_as_id=True).load_data()
print("after an edit:", len(pipeline.run(documents=updated)), "nodes")after an edit: 2 nodes
Reading the edited run
- Two nodes come back, both from the delivery file, now long enough to split in two.
- Only the changed file was processed again; the docstore matched the others by id and content and skipped them.
- filename_as_id=True is what let the edited file be recognised as the same document with new content.
Without a docstore vs with a docstore
| No docstore | With a docstore | |
|---|---|---|
| Re-run cost | Re-splits and re-embeds all | Only changed documents |
| Detects a change | No | Yes, by id and content |
| Needs stable ids | No | Yes, filename_as_id=True |
| Good for | A one-off build | Documents that change over time |
When to use a pipeline with a docstore
- A help centre or wiki that is re-ingested on a schedule, where most files are unchanged.
- Any hosted embedding model, where re-embedding unchanged text costs money.
- A build you want to name once and reuse, rather than re-wiring the reader and splitter each time.
filename_as_id=True an edited file loads with a fresh random id, so the docstore treats it as a new document and re-processes everything. Stable ids are what make the skip work.Related
- Previous: Metadata: what gets embedded and what the model sees
- Next: Retrievers: the search step on its own
- Reference: Ingestion pipeline
- Run the pipeline a third time after the edit and read the node count.
- Remove
filename_as_id=Truefrom the second reader and run it. What happens, and why? - Add a second splitter step with a smaller chunk size and count the nodes.
Slow is fine. Stopping is the only problem.