An enterprise asks: should we use LLM APIs or self-host open models? How do you analyse it?
1. Capability: run the eval set on the top API models vs candidate open models (possibly fine-tuned). If only frontier APIs meet the bar, the decision is mostly made.
2. Data and compliance: regulated data, residency (e.g. data must stay in India), air-gapped environments. Check whether enterprise API offerings (regional endpoints, zero data retention, private deployments through cloud providers) satisfy requirements.
3. Cost model (illustrative method)
API monthly cost = Σ (input_tokens × in_price + output_tokens × out_price)
Self-host monthly = GPUs × hourly_rate × 730 + engineering + ops + overhead
Break-even: volume where they're equal, at REALISTIC GPU utilisation- Self-hosting is cheap per token only if GPUs stay busy. At 20% utilisation it's often more expensive.
- Include engineering salaries, on-call, redundancy (N+1), upgrades and evaluations.
4. Operations: autoscaling, model updates, security patches, monitoring. Does the team have the skills?
5. Latency and control: on-prem for low-latency or edge needs; full control over model versions; custom fine-tuning.
Typical recommendation pattern
| Workload | Choice |
|---|---|
| Complex reasoning, agentic tasks, low or medium volume | Frontier API |
| High-volume narrow tasks (classification, extraction) | Fine-tuned small open model, self-hosted |
| Highly sensitive data with strict residency | Self-hosted or private cloud deployment |
| Prototype / unknown demand | API first; revisit at scale |
Plus an abstraction layer (gateway) to switch between models without rewriting the application.
Related
Slow is fine. Stopping is the only problem.