Integrations: swapping in real backends
An integration is a package that replaces a keyless stand-in with a real backend, such as a hosted model or a vector database, through the same LlamaIndex interface.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
This course ran on keyless stand-ins: local embeddings, an extractive answering model, an in-memory index, and a folder on disk. Each has a real counterpart you swap in with one or two lines, because they share the same interface.
What you used and the real one
| What you used (stand-in) | A real one | Package / class | One-line swap |
|---|---|---|---|
ExtractiveLLM answering model | A hosted LLM | llama-index-llms-groq → Groq | Settings.llm = Groq(model="openai/gpt-oss-120b") |
HuggingFaceEmbedding (local) | Hosted embeddings | llama-index-embeddings-openai → OpenAIEmbedding | Settings.embed_model = OpenAIEmbedding() |
in-memory VectorStoreIndex | A vector database | llama-index-vector-stores-chroma → ChromaVectorStore | pass a StorageContext with the store |
local persist() folder | Server-backed storage | Chroma, Qdrant or Postgres stores | same StorageContext, a different vector store |
MockFunctionCallingLLM (agents) | A tool-calling model | llama-index-llms-openai → OpenAI | Settings.llm = OpenAI(model="gpt-4o-mini") |
Swapping the vector store to Chroma
The store swap is a real drop-in you can run. Build a Chroma vector store, wrap it in a storage context, and pass it to the same from_documents call. Everything after it is unchanged.
import chromadb
from llama_index.core import StorageContext
from llama_index.vector_stores.chroma import ChromaVectorStore
client = chromadb.EphemeralClient()
vector_store = ChromaVectorStore(chroma_collection=client.create_collection("shop"))
storage_context = StorageContext.from_defaults(vector_store=vector_store)View the code here
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
import re
from llama_index.core.llms import CompletionResponse, CustomLLM, LLMMetadata
from llama_index.core.llms.callbacks import llm_completion_callback
def stems(text):
"""Words longer than three letters, cut to five letters, so refund and refunds match."""
return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}
class ExtractiveLLM(CustomLLM):
"""Answers with the context sentence that shares most words with the question."""
@property
def metadata(self):
return LLMMetadata(model_name="extractive")
@llm_completion_callback()
def complete(self, prompt, formatted=False, **kwargs):
context = prompt.split("---------------------")[1]
question = prompt.split("Query:")[1].split("Answer:")[0]
asked = stems(question)
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
sentences = [s for s in sentences if s and not s.startswith("#") and ": " not in s[:20]]
best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
if len(asked & stems(best)) < 2:
return CompletionResponse(text="I could not find that in the documents.")
return CompletionResponse(text=best)
@llm_completion_callback()
def stream_complete(self, prompt, formatted=False, **kwargs):
yield self.complete(prompt)
The same course code on Chroma
The whole program stores the shop documents in Chroma instead of memory, then answers and reports how many vectors the store holds. The query and answer are the same as before.
import chromadb
from llama_index.core import Settings, SimpleDirectoryReader, StorageContext, VectorStoreIndex
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from llama_index.vector_stores.chroma import ChromaVectorStore
from extractive_llm import ExtractiveLLM
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
# the same documents, now stored in Chroma instead of in memory
client = chromadb.EphemeralClient()
vector_store = ChromaVectorStore(chroma_collection=client.create_collection("shop"))
storage_context = StorageContext.from_defaults(vector_store=vector_store)
docs = SimpleDirectoryReader("help").load_data()
index = VectorStoreIndex.from_documents(docs, storage_context=storage_context)
engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)
print(engine.query("How long until my refund money reaches my card?"))
print("stored vectors:", client.get_collection("shop").count())The money goes back to the card you paid with within 5 working days of us receiving the item. stored vectors: 3
Why the swap is one line
- The index still answered the refund question, so retrieval works the same over Chroma.
- Three vectors are stored in Chroma, which means the documents landed in the real store, not in memory.
- Only the store changed: the reader, splitter, query engine and answering model are untouched.
Stand-in vs real backend
| Piece | Keyless stand-in | Real backend |
|---|---|---|
| Answering model | ExtractiveLLM | A hosted LLM through a provider |
| Vector store | in-memory VectorStoreIndex | Chroma, Qdrant or Postgres |
| Embeddings | local HuggingFaceEmbedding | already real, or a hosted model |
When you move off the stand-ins
- Going to production, where an index must be shared and survive restarts.
- Needing fluent answers a real model writes, not extracted sentences.
- A corpus too large for memory, which belongs in a vector database.
The one line that swaps in a real model
Set one default and every query engine and agent in the course uses the real model. It needs an API key and so is not run here; no other line changes.
from llama_index.llms.groq import Groq
Settings.llm = Groq(model="openai/gpt-oss-120b") # set GROQ_API_KEY; real answers, no other changeRelated
- Previous: Updating documents: refresh and delete
- Next: LlamaIndex project: a help-centre assistant
- See also: Persisting an index: not embedding twice
- Reference: Integrations: vector stores
- Give the Chroma collection a different name and rerun; the vector count stays the same.
- Persist the Chroma client to a path and reopen it to confirm the vectors survive.
- Read the Groq swap line and note which single default it changes.
Slow is fine. Stopping is the only problem.