1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What's different about operating LLMs on edge devices or on-device?
30-second answerSay your answer out loud first, then reveal.
Operational aspects
| Area | Practice |
|---|---|
| Model packaging | Quantized formats (e.g. INT4/INT8 for mobile runtimes, GGUF for llama.cpp-class runtimes); per-hardware builds |
| Distribution | Download on first use or Wi-Fi; delta updates; versioning compatible with app versions |
| Capability tiers | Detect device memory and NPU; choose model size; fall back to cloud for complex tasks |
| Evaluation | Quality on target quantization + latency, battery and thermal impact per device class |
| Telemetry | Opt-in, aggregated, privacy-preserving metrics (latency, crashes, fallback rates) |
| Safety | On-device guardrails (small classifiers) since server-side filters may be bypassed offline |
| Security | Model integrity checks; assume weights can be extracted, so don't embed secrets |
| Rollback | Remote config to disable on-device features or switch to cloud mode |
When edge makes sense: privacy-sensitive data (on-device summarisation), offline use (field workers, low connectivity areas), latency-critical interactions, and cost reduction at a very large user scale.
Related
Slow is fine. Stopping is the only problem.