Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q6EasyConcept

What is a vector database, and how do approximate nearest-neighbour (ANN) indexes like HNSW work?

30-second answerSay your answer out loud first, then reveal.

Main ANN index families

IndexIdeaStrengthsTrade-offs
Flat (brute force)Compare against every vector100% recallSlow at scale; fine under ~100K vectors
HNSWMulti-layer graph: sparse top layers for long jumps, dense bottom layer for fine searchFast, high recall, supports insertsMemory-hungry (graph + full vectors in RAM)
IVFCluster vectors (k-means); search only the nearest nprobe clustersLower memory, scalableRecall depends on nprobe; needs training
PQ / quantizationCompress vectors into short codes4–32x less memorySome accuracy loss; often combined (IVF-PQ)
DiskANN-styleGraph index on SSDBillion-scale on one machineHigher latency than in-RAM

HNSW tuning knobs: M (neighbours per node: higher means better recall and more memory), ef_construction (build quality), ef_search (query-time accuracy vs speed).

Vector DB features beyond ANN: metadata filtering, hybrid (keyword + vector) search, CRUD and upserts, replication and sharding, multi-tenancy, backups.

Options: dedicated (Pinecone, Weaviate, Qdrant, Milvus, Chroma) or extensions of existing stores (pgvector for Postgres, Elasticsearch/OpenSearch, MongoDB Atlas, Redis).

Good answer about choosing. "If we already run Postgres and have under ~10M vectors, pgvector is often enough and avoids a new system. At larger scale, or with heavy filtering and hybrid needs, a dedicated engine makes sense."

Follow-ups to expect

  • What happens to HNSW recall when you add restrictive metadata filters? It can drop, because the graph search gets cut off from neighbours. Engines handle this with filtered-HNSW, pre-filtering, or switching to brute force for small filtered sets.

Little by little, you're building something great.