Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q33IntermediateConcept

How do you handle multi-tenancy and noisy neighbours in a shared LLM platform?

30-second answerSay your answer out loud first, then reveal.

Mechanisms

ProblemSolution
One tenant's batch job saturates GPUsSeparate batch and interactive queues; quotas; schedule batch off-peak
Large prompts from one tenant blow up latencyPer-tenant max tokens; long-context pool
Provider rate limits shared across tenantsToken-bucket per tenant within the global budget; fair queuing
Premium SLAsPriority classes; reserved capacity
Data isolationTenant ID enforced in retrieval filters, cache keys, adapter routing, logs
Cost recoveryPer-tenant metering; usage-based pricing or caps
Fair queuing example. Weighted fair queuing by tenant tier. Each tenant gets a share of capacity proportional to its plan weight when the system is contended, and can burst when it's idle.

Monitoring: per-tenant latency, errors, usage vs quota, and a "top consumers" dashboard to spot abuse or unexpected load early.

Slow is fine. Stopping is the only problem.