Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q5EasyConcept

What are the main options for deploying LLMs, and how do you choose?

30-second answerSay your answer out loud first, then reveal.
OptionProsConsGood for
Model provider APIBest models, zero infra, fast iterationData leaves network (with contractual controls), rate limits, vendor dependencyMost new features, complex reasoning
Managed cloud platform (in your cloud account)Regional processing, IAM integration, consolidated billingModel selection varies by region; some lagEnterprises with cloud commitments and residency needs
Self-hosted (Kubernetes + inference server)Control, privacy, predictable cost at high utilisation, fine-tuned modelsGPU ops, capacity planning, on-callHigh-volume narrow tasks, sensitive data, custom models
Edge / on-deviceOffline, privacy, no per-token costSmall models only, device fragmentation, update logisticsMobile features, field devices

Decision inputs: eval results per option, data classification, monthly token volume vs GPU cost at realistic utilisation, latency SLOs, team capability, and lock-in tolerance. Use a gateway abstraction so you can switch or mix.

Every expert started right here.