1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Design an AI code completion system like GitHub Copilot's inline suggestions.
30-second answerSay your answer out loud first, then reveal.

Key design points
- Latency is everything: suggestions must appear within a few hundred ms or they're useless. Use small models (1–15B class) optimised for speed, nearest-region serving, persistent connections, speculative decoding, and short max output (a line or block).
- Fill-in-the-middle: models trained with
<prefix> ... <suffix> ... <middle>formats use code after the cursor, which is crucial for good insertions. - Context retrieval: the best gains often come from relevant snippets of other files (Jaccard or embedding similarity over recently viewed files, symbol definitions via the language server).
- Request economics: many requests per user per hour. Debouncing, caching (typing through a suggestion should reuse it) and cancellation cut load significantly.
- Post-processing: cut at syntactically sensible points (using indentation and brackets); avoid repeating existing code; filter secrets; optionally block suggestions matching public code (licence concerns).
- Evaluation: offline (exact match / execution-based on held-out repositories, FIM benchmarks); online (acceptance rate, characters retained, latency percentiles, by language).
- Privacy: enterprise settings for no retention or training; on-prem options.
Separate from chat/agent features: the multi-file "agent mode" uses larger models and different latency expectations (see the Agentic AI guide).
Related
You understood something today that you didn't yesterday.