LlamaIndexllama-index-core 0.14 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Integrations: swapping in real backends

An integration is a package that replaces a keyless stand-in with a real backend, such as a hosted model or a vector database, through the same LlamaIndex interface.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

This course ran on keyless stand-ins: local embeddings, an extractive answering model, an in-memory index, and a folder on disk. Each has a real counterpart you swap in with one or two lines, because they share the same interface.

What you used and the real one

What you used (stand-in)A real onePackage / classOne-line swap
ExtractiveLLM answering modelA hosted LLMllama-index-llms-groq → GroqSettings.llm = Groq(model="openai/gpt-oss-120b")
HuggingFaceEmbedding (local)Hosted embeddingsllama-index-embeddings-openai → OpenAIEmbeddingSettings.embed_model = OpenAIEmbedding()
in-memory VectorStoreIndexA vector databasellama-index-vector-stores-chroma → ChromaVectorStorepass a StorageContext with the store
local persist() folderServer-backed storageChroma, Qdrant or Postgres storessame StorageContext, a different vector store
MockFunctionCallingLLM (agents)A tool-calling modelllama-index-llms-openai → OpenAISettings.llm = OpenAI(model="gpt-4o-mini")

Swapping the vector store to Chroma

The store swap is a real drop-in you can run. Build a Chroma vector store, wrap it in a storage context, and pass it to the same from_documents call. Everything after it is unchanged.

python
import chromadb
from llama_index.core import StorageContext
from llama_index.vector_stores.chroma import ChromaVectorStore

client = chromadb.EphemeralClient()
vector_store = ChromaVectorStore(chroma_collection=client.create_collection("shop"))
storage_context = StorageContext.from_defaults(vector_store=vector_store)

The same course code on Chroma

The whole program stores the shop documents in Chroma instead of memory, then answers and reports how many vectors the store holds. The query and answer are the same as before.

Example
import chromadb
from llama_index.core import Settings, SimpleDirectoryReader, StorageContext, VectorStoreIndex
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.chroma import ChromaVectorStore

from extractive_llm import ExtractiveLLM

Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")

# the same documents, now stored in Chroma instead of in memory
client = chromadb.EphemeralClient()
vector_store = ChromaVectorStore(chroma_collection=client.create_collection("shop"))
storage_context = StorageContext.from_defaults(vector_store=vector_store)

docs = SimpleDirectoryReader("help").load_data()
index = VectorStoreIndex.from_documents(docs, storage_context=storage_context)

engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)
print(engine.query("How long until my refund money reaches my card?"))
print("stored vectors:", client.get_collection("shop").count())

Why the swap is one line

  • The index still answered the refund question, so retrieval works the same over Chroma.
  • Three vectors are stored in Chroma, which means the documents landed in the real store, not in memory.
  • Only the store changed: the reader, splitter, query engine and answering model are untouched.

Stand-in vs real backend

PieceKeyless stand-inReal backend
Answering modelExtractiveLLMA hosted LLM through a provider
Vector storein-memory VectorStoreIndexChroma, Qdrant or Postgres
Embeddingslocal HuggingFaceEmbeddingalready real, or a hosted model

When you move off the stand-ins

  • Going to production, where an index must be shared and survive restarts.
  • Needing fluent answers a real model writes, not extracted sentences.
  • A corpus too large for memory, which belongs in a vector database.

The one line that swaps in a real model

Set one default and every query engine and agent in the course uses the real model. It needs an API key and so is not run here; no other line changes.

python
from llama_index.llms.groq import Groq

Settings.llm = Groq(model="openai/gpt-oss-120b")  # set GROQ_API_KEY; real answers, no other change
Watch out. A real embedding model must be the same one at build time and query time. If you build with local embeddings and later query with a hosted model, the vectors do not match and the search returns nonsense with no error.
Try it yourself
  • Give the Chroma collection a different name and rerun; the vector count stays the same.
  • Persist the Chroma client to a path and reopen it to confirm the vectors survive.
  • Read the Groq swap line and note which single default it changes.

Slow is fine. Stopping is the only problem.