LlamaIndexllama-index-core 0.14 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Workflows: a custom RAG pipeline with @step

A Workflow is a class whose steps pass events to each other, so you build a retrieve-then-answer pipeline you control step by step.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

A query engine hides retrieve and answer inside one call. A workflow splits them into named steps joined by events, so you can add a step, store something between steps, or change the order.

Declaring an event and steps

An Event is a small object one step returns and the next receives. A method marked @step becomes a step; its argument type says which event starts it.

python
from llama_index.core.workflow import Workflow, step, StartEvent, StopEvent, Event, Context

class Retrieved(Event):
    question: str
    context: str

The retrieve step

The first step takes the StartEvent, retrieves nodes, saves the source names in ctx.store, and returns a Retrieved event with the joined text.

python
    @step
    async def retrieve(self, ctx: Context, ev: StartEvent) -> Retrieved:
        nodes = self.retriever.retrieve(ev.question)
        await ctx.store.set("sources", sorted({n.metadata["file_name"] for n in nodes}))
        context = "\n".join(n.get_content() for n in nodes)
        return Retrieved(question=ev.question, context=context)

The answer step

The second step receives the Retrieved event, builds the prompt, asks the stand-in model, reads the sources back out of the context store, and returns a StopEvent, which ends the run.

python
    @step
    async def answer(self, ctx: Context, ev: Retrieved) -> StopEvent:
        prompt = f"---------------------{ev.context}---------------------\nQuery:{ev.question}\nAnswer:"
        reply = self.llm.complete(prompt).text
        sources = await ctx.store.get("sources")
        return StopEvent(result=f"{reply} (from {', '.join(sources)})")

The workflow answering two questions

The whole program. The two steps run in order, joined by the Retrieved event, and the context store carries the source names from the first step to the second.

Example
import asyncio

from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.workflow import Context, Event, StartEvent, StopEvent, Workflow, step
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

from extractive_llm import ExtractiveLLM

Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")


class Retrieved(Event):
    question: str
    context: str


class RAGWorkflow(Workflow):
    def __init__(self, index, **kwargs):
        super().__init__(**kwargs)
        self.retriever = index.as_retriever(similarity_top_k=2)
        self.llm = ExtractiveLLM()

    @step
    async def retrieve(self, ctx: Context, ev: StartEvent) -> Retrieved:
        nodes = self.retriever.retrieve(ev.question)
        await ctx.store.set("sources", sorted({n.metadata["file_name"] for n in nodes}))
        context = "\n".join(n.get_content() for n in nodes)
        return Retrieved(question=ev.question, context=context)

    @step
    async def answer(self, ctx: Context, ev: Retrieved) -> StopEvent:
        prompt = f"---------------------{ev.context}---------------------\nQuery:{ev.question}\nAnswer:"
        reply = self.llm.complete(prompt).text
        sources = await ctx.store.get("sources")
        return StopEvent(result=f"{reply} (from {', '.join(sources)})")


async def main():
    docs = SimpleDirectoryReader("help").load_data()
    index = VectorStoreIndex.from_documents(docs)
    workflow = RAGWorkflow(index, timeout=60)
    for question in ["How long until my refund money reaches my card?", "What are your opening hours?"]:
        result = await workflow.run(question=question)
        print(question)
        print("  ", result)


asyncio.run(main())

How the events moved through the run

  • The start event carried the question into retrieve, which returned a Retrieved event.
  • The Retrieved event triggered answer, because its argument type is Retrieved.
  • The context store passed the source names between steps without putting them in the event.
  • The absent question still ran both steps; the stand-in model refused because no sentence matched.

Query engine vs a workflow

ApproachStepsControl
Query engineFixed: retrieve then answerLittle; one call
WorkflowWhatever steps you writeFull: add steps, branch, store state

When a workflow is worth the extra code

  • Adding a step, such as a check or a rewrite, between retrieval and answering.
  • Branching on the question before deciding how to answer.
  • Passing state between steps that should not travel inside the events.
Watch out. A step runs only when some step returns the event type it asks for. If nothing returns your event, the step never fires and the run stalls until it times out; check the return type of the step before it.
Try it yourself
  • Add a middle step that prints the question before answering.
  • Store the number of retrieved nodes in ctx.store and print it in the answer.
  • Change similarity_top_k to 1 and see whether the refund question still finds its sentence.

Slow is fine. Stopping is the only problem.