Integrations: swapping in real backends
An integration is a package that replaces a keyless stand-in with a real backend, such as a hosted model or a vector database, through the same LlamaIndex interface.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
This course ran on keyless stand-ins: local embeddings, an extractive answering model, an in-memory index, and a folder on disk. Each has a real counterpart you swap in with one or two lines, because they share the same interface.
What you used and the real one
| What you used (stand-in) | A real one | Package / class | One-line swap |
|---|---|---|---|
ExtractiveLLM answering model | A hosted LLM | llama-index-llms-groq → Groq | Settings.llm = Groq(model="openai/gpt-oss-120b") |
HuggingFaceEmbedding (local) | Hosted embeddings | llama-index-embeddings-openai → OpenAIEmbedding | Settings.embed_model = OpenAIEmbedding() |
in-memory VectorStoreIndex | A vector database | llama-index-vector-stores-chroma → ChromaVectorStore | pass a StorageContext with the store |
local persist() folder | Server-backed storage | Chroma, Qdrant or Postgres stores | same StorageContext, a different vector store |
MockFunctionCallingLLM (agents) | A tool-calling model | llama-index-llms-openai → OpenAI | Settings.llm = OpenAI(model="gpt-4o-mini") |
Swapping the vector store to Chroma
The store swap is a real drop-in you can run. Build a Chroma vector store, wrap it in a storage context, and pass it to the same from_documents call. Everything after it is unchanged.
import chromadb
from llama_index.core import StorageContext
from llama_index.vector_stores.chroma import ChromaVectorStore
client = chromadb.EphemeralClient()
vector_store = ChromaVectorStore(chroma_collection=client.create_collection("shop"))
storage_context = StorageContext.from_defaults(vector_store=vector_store)The same course code on Chroma
The whole program stores the shop documents in Chroma instead of memory, then answers and reports how many vectors the store holds. The query and answer are the same as before.
import chromadb
from llama_index.core import Settings, SimpleDirectoryReader, StorageContext, VectorStoreIndex
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.chroma import ChromaVectorStore
from extractive_llm import ExtractiveLLM
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
# the same documents, now stored in Chroma instead of in memory
client = chromadb.EphemeralClient()
vector_store = ChromaVectorStore(chroma_collection=client.create_collection("shop"))
storage_context = StorageContext.from_defaults(vector_store=vector_store)
docs = SimpleDirectoryReader("help").load_data()
index = VectorStoreIndex.from_documents(docs, storage_context=storage_context)
engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)
print(engine.query("How long until my refund money reaches my card?"))
print("stored vectors:", client.get_collection("shop").count())Why the swap is one line
- The index still answered the refund question, so retrieval works the same over Chroma.
- Three vectors are stored in Chroma, which means the documents landed in the real store, not in memory.
- Only the store changed: the reader, splitter, query engine and answering model are untouched.
Stand-in vs real backend
| Piece | Keyless stand-in | Real backend |
|---|---|---|
| Answering model | ExtractiveLLM | A hosted LLM through a provider |
| Vector store | in-memory VectorStoreIndex | Chroma, Qdrant or Postgres |
| Embeddings | local HuggingFaceEmbedding | already real, or a hosted model |
When you move off the stand-ins
- Going to production, where an index must be shared and survive restarts.
- Needing fluent answers a real model writes, not extracted sentences.
- A corpus too large for memory, which belongs in a vector database.
The one line that swaps in a real model
Set one default and every query engine and agent in the course uses the real model. It needs an API key and so is not run here; no other line changes.
from llama_index.llms.groq import Groq
Settings.llm = Groq(model="openai/gpt-oss-120b") # set GROQ_API_KEY; real answers, no other changeRelated
- Previous: Updating documents: refresh and delete
- Next: LlamaIndex project: a help-centre assistant
- See also: Persisting an index: not embedding twice
- Reference: Integrations: vector stores
- Give the Chroma collection a different name and rerun; the vector count stays the same.
- Persist the Chroma client to a path and reopen it to confirm the vectors survive.
- Read the Groq swap line and note which single default it changes.
Slow is fine. Stopping is the only problem.