Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q42HardSystem design

Design a self-service fine-tuning platform for internal teams (upload data → train → evaluate → deploy).

30-second answerSay your answer out loud first, then reveal.
A dataset is uploaded, validated and versioned, then a job config goes to a GPU scheduler and training workers; the result is auto-evaluated, stored in a model registry, passes an approval gate, and is served with multi-LoRA under monitoring and rollback.

Key components

  1. Data governance: consent and usage rights, PII redaction, licence checks; dataset versioning and lineage.
  2. Templates and guardrails: sensible defaults (LoRA rank, LR, epochs); prevent common mistakes (missing EOS tokens, chat template mismatch, train/test leakage).
  3. Compute management: GPU quotas per team, priority queues, preemption with checkpoint/resume, cost reporting.
  4. Experiment tracking: metrics, configs and artifacts (MLflow / W&B-style).
  5. Evaluation gate: must beat the baseline on the task eval and not regress safety or general abilities beyond thresholds.
  6. Serving efficiency: many LoRA adapters on one base model (multi-LoRA serving) instead of a GPU per fine-tuned model.
  7. Lifecycle: retraining when data drifts; deprecating adapters when the base model is upgraded (re-train on the new base).

Success metrics: time from data to deployed model, GPU utilisation, percentage of fine-tunes beating prompting baselines, incidents.

You understood something today that you didn't yesterday.