Document Q&A API: answer questions about a PDF
The month 1 project of the AI Forward Deployed Engineer roadmap: a FastAPI service that takes a PDF once and then answers plain questions about it, each answer naming the page it came from. It runs in Docker, has tests, and lives on GitHub with a README a stranger can follow.
The problem
A small support team answers the same questions about the same handful of PDF policies all day: refund windows, delivery times, what counts as damaged. Today they open the file and skim, or paste pages into a chat window and hope. Both are slow and neither shows where the answer came from, so nobody trusts it for a customer.
They want one place to upload a policy PDF and ask a question in plain words, and get back a short answer and the page it is on, so they can check it before repeating it to a customer. It has to be a service, not a notebook: something a colleague can start with one command and call from their own tool.
Architecture
Two paths share one store. Ingest runs once per document: the upload endpoint reads the PDF into page-numbered text and splits it into passages that each remember their page. Ask runs on every question: it finds the passages most likely to answer, hands them and the question to a model, and returns a typed object with the answer and the page. Keeping the two apart is what lets a reader upload a 20-page policy once and then ask twenty questions cheaply.
The one decision worth making early is how the retriever finds passages. Plain word overlap is enough for a handful of short policies and needs nothing installed; an embedding search scales better and is month 2's subject. Either way the passages carry their page, which is what makes the citation honest.
What it draws on
The project uses all four weeks of month 1. Three have lessons on the site; the rest are upcoming courses, so use the official docs for them until they open.
- Python for AI: Python for production: Pydantic models for the typed answer, async calls, retries and pytest
- Python for AI: files and JSON: reading the upload and writing the passage store
- LLM Fundamentals: how LLMs work: the model call, the context window, and why the passages have to be short
- FastAPI, PDF text extraction and Docker: courses upcoming. Use the FastAPI tutorial, the docs of a PDF library such as pypdf, and Docker's Get started guide
What done looks like
| Requirement | Done when |
|---|---|
| FastAPI service | POST /documents accepts a PDF and POST /ask answers a question about it |
| Typed answer | An answer comes back as an object with the answer and a page number, never free text the caller has to parse |
| Honest citation | The page in the answer is the page the passage actually came from, checkable in the PDF |
| Refuses cleanly | A scanned PDF with no text, or a question the document does not cover, returns a clear message, not a made-up answer or a crash |
| Tested | pytest covers both endpoints, the empty-PDF case and the no-answer case |
| Dockerised | docker compose up starts the service with the model key read from the environment |
| On GitHub | A public repository with a README that gets a stranger from clone to first answer in under five minutes |
Where to start
Build it back to front, so something works at every step. Get one page of text out of a PDF and print it. Answer one hard-coded question from that text with a model. Then wrap the two steps in the upload and ask endpoints, add the page to the response, and handle the empty-PDF and no-answer cases. Write the tests as you go, and the Dockerfile and README last, by following your own README on a clean machine.
Every expert started right here.