# Overview
This tutorial teaches how a vector database works by building one step-by-step in Python. The implementation uses NumPy and a small sentence-transformer model to encode text into fixed-size floating-point vectors. The guide is hands-on: add each code block to a script and run it to see incremental behavior.
# What you build
You build a minimal VectorDB implementation that:
- loads an embedding model (sentence-transformers/all-MiniLM-L6-v2 in the example),
- encodes documents into fixed-length vectors,
- stores vectors alongside metadata,
- supports retrieval by cosine similarity,
- adds input validation and metadata filtering,
- persists index data to disk,
- includes a test suite to verify behavior.
The example corpus included in the repo has 25 simulated documents with topic tags.
# Why fixed-size vectors matter
Every document, regardless of length, becomes the same vector size. In the tutorial the embeddings use dimension 384 and float32 values, so each document occupies 384 × 4 = 1,536 bytes. The provided example prints a summary: VectorDB(25 docs, dim=384, model='sentence-transformers/all-MiniLM-L6-v2') and an index size of about 37.5 KiB.
# Files and how to run
- vector_db.py: the VectorDB implementation you explore.
- corpus.py: sample documents (25 docs) and metadata used for demos.
- test.py: a test suite you can run to check the implementation.
# Performance notes and scaling
The tutorial measures timings: model load (example 1.64s) and encoding (example 0.14s for 25 documents, ~6 ms per document). It explains that brute-force cosine similarity is simple and exact but scales linearly with corpus size. When latency or corpus size increase, consider approximate nearest neighbor (ANN) indexing strategies (the article explains when brute-force becomes impractical).
# Incremental features covered
The ten steps are designed to each teach one idea: setting up the environment, building the index, encoding documents, adding metadata and filtering, validating inputs, implementing persistence, and evaluating search behavior. Example helper functions (show and header) keep output readable while you run the script as it grows.
# How to follow the tutorial
Create tutorial.py and append each step's code in sequence, running the script after each addition. The printed output at each stage helps you confirm that the code does what the explanatory text says. Running python test.py gives additional reassurance that the database behaves as implemented.
# Practical next decisions
After you understand the minimal implementation, decide whether you need: larger or custom embedding models, an ANN index for larger corpora, or more robust persistence and access controls for production. The tutorial gives enough detail to make those choices informed rather than speculative.