CrewAICrewAI 1.15 · Python 3.10 to 3.13
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
35 small wins to finish your pathNext lesson →

Knowledge sources: a reference library for agents

Knowledge is a set of reference documents an agent searches before it answers, so its reply is grounded in facts you loaded rather than in the model alone.

Last updated: 28 Sep, 2026 · CrewAI 1.15

The desk from the flow lesson answers about orders, but a shop also has fixed rules: a refund window, a shipping time, a returns address. Those belong in Knowledge: you load them once, and CrewAI finds the closest ones to each question and puts them in front of the agent. Knowledge is read before an answer; it is not the running memory of a conversation, which is a separate idea covered next.

Turning text into a vector with an embedder

Knowledge finds the closest document by comparing vectors, so it needs an embedder: something that turns text into a list of numbers. The default embedder calls a hosted service that needs a key. To keep this lesson runnable with no key, write a small one that any machine can run.

python
from crewai.rag.embeddings.providers.custom.embedding_callable import CustomEmbeddingFunction

class WordHash(CustomEmbeddingFunction):
    def name(self):
        return "word_hash"
    def __call__(self, input):
        out = []
        for text in input:
            v = [0.0] * 64
            for word in text.lower().split():
                v[int(__import__("hashlib").md5(word.strip(".,?!").encode()).hexdigest(), 16) % 64] += 1.0
            out.append(v)
        return out

WordHash puts each word into one of 64 slots picked by its hash. Two texts that share words point the same way, so their vectors are close. A hosted embedder understands meaning; this one only matches words, which is enough to see retrieval work.

Loading the shop's rules as sources

A StringKnowledgeSource wraps a piece of text. You can also load files with the CSV, PDF, JSON and text sources, but a string is the shortest way to see it work.

python
from crewai.knowledge.source.string_knowledge_source import StringKnowledgeSource

policy = StringKnowledgeSource(content="A refund is offered within 30 days of delivery.")
shipping = StringKnowledgeSource(content="Every order ships within two working days.")

Each source is one fact here. In a real shop a source is a whole policy file, which CrewAI splits into chunks before it embeds them.

Building the store and asking it

Knowledge ties the sources to the embedder and stores the vectors in ChromaDB on disk. add_sources embeds and saves; query returns the closest chunks.

python
from crewai import Knowledge
from crewai.rag.embeddings.providers.custom.custom_provider import CustomProvider

kb = Knowledge(collection_name="shop_policy", sources=[policy, shipping],
               embedder=CustomProvider(embedding_callable=WordHash))
kb.add_sources()
hits = kb.query(["When can I get a refund?"], results_limit=1)

CustomProvider is how CrewAI accepts an embedder you wrote: it takes the class, not an instance. The store keeps its files under CrewAI's local storage folder, so a second run reads them back.

Retrieving the right source for each question

Exampleknowledge_desk.py
import os, hashlib, math
os.environ["OTEL_SDK_DISABLED"] = "true"
os.environ["CREWAI_DISABLE_TELEMETRY"] = "true"
from crewai import Knowledge
from crewai.knowledge.source.string_knowledge_source import StringKnowledgeSource
from crewai.rag.embeddings.providers.custom.embedding_callable import CustomEmbeddingFunction
from crewai.rag.embeddings.providers.custom.custom_provider import CustomProvider

class WordHash(CustomEmbeddingFunction):
    def name(self):
        return "word_hash"
    def __call__(self, input):
        out = []
        for text in input:
            v = [0.0] * 64
            for word in text.lower().split():
                word = word.strip(".,?!")
                v[int(hashlib.md5(word.encode()).hexdigest(), 16) % 64] += 1.0
            length = math.sqrt(sum(x * x for x in v)) or 1.0
            out.append([x / length for x in v])
        return out

policy = StringKnowledgeSource(content="A refund is offered within 30 days of delivery.")
shipping = StringKnowledgeSource(content="Every order ships within two working days.")
kb = Knowledge(collection_name="shop_policy", sources=[policy, shipping],
               embedder=CustomProvider(embedding_callable=WordHash))
kb.reset()
kb.add_sources()
for q in ["When can I get a refund?", "How fast does an order ship?"]:
    top = kb.query([q], results_limit=1)[0]
    print(q, "->", top["content"])

Reading what came back

  • The refund question returned the refund line, not the shipping line: the two share the word "refund", and no shipping word is in the question.
  • The order question returned the shipping line: it shares "order" with that source, so the vectors line up there instead.
  • Each answer tracks its question. Change the words in a question and the match changes, which is what makes this retrieval and not a fixed reply.
  • A result is a dict with content, score, id and metadata; a higher score is a closer match.

Knowledge next to Memory

Both search by vector, but they hold different things and are written at different times.

KnowledgeMemory
What it holdsReference facts you load up frontWhat a crew saw during its runs
Written whenBefore the run, by youDuring and after runs, by the crew
Read howRetrieved into the prompt before the agent answersRecalled when a task asks for it
Stored withChromaDBLanceDB by default

Giving the sources to a crew

An agent or a crew takes knowledge_sources and an embedder of its own. Before the agent answers, CrewAI retrieves the closest chunks and adds them to the prompt, so the reply is built on the loaded rules.

python
from crewai import Crew

crew = Crew(agents=[clerk], tasks=[reply],
            knowledge_sources=[policy, shipping],
            embedder=CustomProvider(embedding_callable=WordHash))

When a reference library fits

  • Fixed rules that many tickets need: refund windows, shipping times, a returns address.
  • Product facts or documentation an agent should quote instead of guessing.
  • Anything you want the agent to read from, not invent, and that does not change per conversation.
Watch out
The default embedder calls a hosted service and needs a key. Build Knowledge with no embedder and no key, and add_sources fails the moment it tries to embed. Set an embedder you can run, or provide the key, before you load anything.
Try it yourself
  • Add a third source about the returns address and ask "Where do I send a return?".
  • Change a question to share no words with any source and read the score you get back.
  • Raise results_limit to 2 and print both hits with their scores.

Little by little, you're building something great.