Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q48HardConcept

What's different about operating LLMs on edge devices or on-device?

30-second answerSay your answer out loud first, then reveal.

Operational aspects

AreaPractice
Model packagingQuantized formats (e.g. INT4/INT8 for mobile runtimes, GGUF for llama.cpp-class runtimes); per-hardware builds
DistributionDownload on first use or Wi-Fi; delta updates; versioning compatible with app versions
Capability tiersDetect device memory and NPU; choose model size; fall back to cloud for complex tasks
EvaluationQuality on target quantization + latency, battery and thermal impact per device class
TelemetryOpt-in, aggregated, privacy-preserving metrics (latency, crashes, fallback rates)
SafetyOn-device guardrails (small classifiers) since server-side filters may be bypassed offline
SecurityModel integrity checks; assume weights can be extracted, so don't embed secrets
RollbackRemote config to disable on-device features or switch to cloud mode

When edge makes sense: privacy-sensitive data (on-device summarisation), offline use (field workers, low connectivity areas), latency-critical interactions, and cost reduction at a very large user scale.

Slow is fine. Stopping is the only problem.