Secure enterprise copilot: SSO, permissions and an audit trail
The month 5 project of the AI Forward Deployed Engineer roadmap: a copilot staff sign in to with their company account, that only retrieves documents each user is allowed to see, stops injected instructions, records every answer in an audit trail, and can run inside a private network with an open-weight model as the fallback.
The problem
Your month 2 assistant answers well. Now a bank or an insurer wants it, and their security team has questions. Who can use it? Can a junior analyst get an answer from a board paper? What happens when a document contains instructions aimed at the model? Where do prompts go, and who can read them? Can it run with no data leaving our network?
This project answers each of those questions in the system itself, not in a slide. It is the version of a RAG assistant an enterprise can actually approve.
Architecture
The user signs in through the client's identity provider, and their identity travels with the request through every step. Input rails check the question before the agent runs. Retrieval adds a filter built from the user's groups, so it can only return documents they could open themselves. Models sit behind one gateway: the client's private endpoint first, an open-weight model served with vLLM as the fallback. Every answer is written to an audit log.
Enforce permissions in retrieval, never in the prompt. If a document the user may not see reaches the model, no instruction will reliably keep it out of the answer. Store each document's access list as metadata when you index it, and filter on it in every query.
What it draws on
- Topic: SSO, OAuth and RBAC
- LlamaIndex: metadata filters, for per-user retrieval
- LlamaIndex: incremental indexing
- Topic: customer data and PII
- NeMo Guardrails: prompt-injection defences
- LangGraph: the agent and its state, carrying the user's identity
- Connectors, Guardrails AI for DLP, compliance, private deployment and vLLM: courses upcoming. Use the official docs for your identity provider, your cloud's private endpoints and vLLM's OpenAI-compatible server
What done looks like
| Requirement | Done when |
|---|---|
| SSO | Users sign in through an OIDC or SAML identity provider; there is no separate password |
| Per-user permissions | Two test users with different groups get different sources for the same question, and neither ever sees a document outside their groups |
| Injection | A document or message telling the model to ignore its instructions is caught by a rail, and the rail is named in the log |
| Audit trail | Each answer records who asked, the question, the documents retrieved, the rails that fired and the answer, with a timestamp |
| PII | Personal data is masked before it reaches logs, traces and eval sets |
| Private deployment | A written deployment plan, and ideally a working setup, for a private VPC with no public model endpoint |
| Open-weight fallback | With the main model switched off, the copilot answers through an open-weight model behind the same gateway |
Where to start
Start from your month 2 assistant. Add two fake users and an access list on every document, and make the permission test pass before anything else. Then add sign-in, the input rails and the audit log. Leave the private deployment and the open-weight fallback for last, once everything else works behind the gateway.
Every expert started right here.