1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are the main options for deploying LLMs, and how do you choose?
30-second answerSay your answer out loud first, then reveal.
| Option | Pros | Cons | Good for |
|---|---|---|---|
| Model provider API | Best models, zero infra, fast iteration | Data leaves network (with contractual controls), rate limits, vendor dependency | Most new features, complex reasoning |
| Managed cloud platform (in your cloud account) | Regional processing, IAM integration, consolidated billing | Model selection varies by region; some lag | Enterprises with cloud commitments and residency needs |
| Self-hosted (Kubernetes + inference server) | Control, privacy, predictable cost at high utilisation, fine-tuned models | GPU ops, capacity planning, on-call | High-volume narrow tasks, sensitive data, custom models |
| Edge / on-device | Offline, privacy, no per-token cost | Small models only, device fragmentation, update logistics | Mobile features, field devices |
Decision inputs: eval results per option, data classification, monthly token volume vs GPU cost at realistic utilisation, latency SLOs, team capability, and lock-in tolerance. Use a gateway abstraction so you can switch or mix.
Related
Every expert started right here.