LlamaIndexllama-index-core 0.14 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Persisting an index: not embedding twice

Persisting is a save step that writes a built index to disk, so a later run loads it instead of embedding every chunk again.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

Embedding thousands of chunks takes time, and money with a hosted model. Build the index once, save it, and every later run reads the saved files.

Saving and loading an index

persist writes the index to a folder; load_index_from_storage reads it back through a storage context.

python
index.storage_context.persist(persist_dir="storage")   # save once

from llama_index.core import StorageContext, load_index_from_storage
storage = StorageContext.from_defaults(persist_dir="storage")
index = load_index_from_storage(storage)                # load later

Building and saving

Load the files, split them, build the index, then persist it. The last line lists the files that were written.

python
nodes = SentenceSplitter(chunk_size=80, chunk_overlap=0).get_nodes_from_documents(documents)
index = VectorStoreIndex(nodes)
index.storage_context.persist(persist_dir="storage")
print(sorted(os.listdir("storage")))
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps

The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.

All lamps come with a two year guarantee against electrical faults.

Bulbs are not covered by the refund policy once they have been used.

The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds

You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.

Items bought in a sale can be refunded too, but the delivery charge is not returned.

To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.

Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery

Standard delivery takes 3 to 5 working days and is free on orders over 40.

Express delivery arrives the next working day if you order before 2pm. It costs 6.

We deliver to the mainland only. Parcels to islands take 2 extra working days.

If a parcel has not arrived after 10 working days, contact us and we will send a replacement.

Writing the index to a folder

The whole program builds the index and saves it. The printed list is what persist wrote.

Example
index.storage_context.persist(persist_dir="storage")
print(sorted(os.listdir("storage")))

docstore.json holds the chunks' text, index_store.json describes the index, and default__vector_store.json holds the embeddings. The graph and image stores are empty files kept for other index types.

Loading without re-embedding

A second run loads the saved index and searches it. No document is read or embedded again; only the question is embedded, with the same model.

Example
from llama_index.core import Settings
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
from llama_index.core import StorageContext, load_index_from_storage

storage = StorageContext.from_defaults(persist_dir="storage")
index = load_index_from_storage(storage)
print(len(index.docstore.docs), "chunks loaded")
print([n.metadata["file_name"] for n in index.as_retriever(similarity_top_k=1).retrieve("How do I get my money back?")])

What loading gave back

  • Six chunks loaded means the saved docstore came back whole, without re-reading the files.
  • The refund file was retrieved for the money question, so the loaded index searches exactly as the built one did.
  • Only the question was embedded on this run, which is the work that was saved.

Rebuilding every run vs persisting once

ApproachEmbeds the documentsStartup
Rebuild each runEvery runSlow, and costly with a hosted model
Persist onceOnceFast; later runs only load

When to persist an index

  • Any app that restarts, so it does not re-embed on every boot.
  • A large corpus where embedding takes minutes.
  • A hosted embedding model you pay for per call.
Watch out. The embedding model used to build the index must be set the same way when loading. Vectors from one model are meaningless to another, and nothing checks for you, so the search silently returns wrong results.
Try it yourself
  • Delete the storage folder and run the loader. What error do you get?
  • Load the index and print index.docstore.docs to see the saved chunks.
  • Change the embedding model on the loading run and read how the retrieved file changes.

Little by little, you're building something great.