Towards Data Science iconTowards Data ScienceSep 23, 2026 ~9 min source read

From Words to Vectors: How TF‑IDF Converts Text into Numbers

A step‑by‑step walkthrough of tokenization, vocabulary, term frequency, and the TF‑IDF calculation using a small review dataset — presented so you can see what the numbers actually represent before moving on to embeddings and other NLP representations.

Share this story

Send the public story page.

Useful takeaways from this story.

TF‑IDF combines term frequency (how often a word appears in a document) with inverse document frequency (how common that word is across all documents) by multiplying TF and IDF.

Representing text for models starts with tokenization and building a fixed vocabulary so every document yields a numeric vector with one entry per vocabulary term.

# Overview

# Dataset and goal The worked dataset consists of four short reviews (labeled D1–D4). After tokenization each review becomes a list of words, for example:

  • D1: ["the", "food", "was", "good", "and", "fresh"]
  • D2: ["the", "food", "was", "good", "and", "tasty"]
  • D3: ["the", "food", "was", "bad", "and", "stale"]
  • D4: ["the", "food", "was", "bad", "and", "tasteless"]

The article keeps the dataset small so you can see how each step affects the numeric output.

# Step 1 — Tokenization Tokenization breaks text into individual words (tokens). TF‑IDF operates at the token level, so tokenization is the first required preprocessing step. Each document is represented as a sequence of tokens.

# Step 2 — Building the vocabulary Next you collect every unique token across all documents to create the vocabulary. The vocabulary defines the coordinates of the vector space: each token gets a fixed position in every document vector. Even if a document does not contain a token, that token must still have an entry (zero) in that document's vector.

# Step 3 — Term Frequency (TF) Term frequency measures how often a token appears inside a single document, normalized by the document length. The formula shown is:

TF(t, d) = (Number of times t appears in d) / (Total number of terms in d)

For D1, which has six tokens, each word that appears once has TF = 1/6 ≈ 0.1667. Words that do not appear in D1 have TF = 0. The TF vector for a document lists these TF values in the vocabulary order.

# Step 4 — Inverse Document Frequency and TF‑IDF TF‑IDF pairs TF with inverse document frequency (IDF), which downweights tokens that are common across the whole corpus and upweights rarer tokens. The two components are multiplied:

Multiplying TF and IDF produces a numeric vector for each document where each entry indicates how important that token is to that document relative to the collection.

# What the vectors mean and what you can do with them TF‑IDF vectors place documents as points in a vector space whose dimensions are vocabulary tokens. Practical consequences:

  • Similarity: documents with similar token importance patterns end up near each other in that space. That supports retrieval and clustering.
  • Classification: the vectors are usable as features for supervised models that learn to map patterns of token importance to labels.
  • Interpretability: TF‑IDF weights are easy to interpret term by term (higher weight = more helpful token for distinguishing a document).

This representation is a classical, sparse, interpretable starting point. The article positions TF‑IDF as a foundational step before moving to denser learned representations such as embeddings.

# Bottom line Follow the four concrete steps — tokenize, build vocabulary, compute TF, multiply by IDF — to convert text into TF‑IDF vectors. With a small example the transformations are transparent, which helps when you later compare TF‑IDF to embeddings and transformer‑based representations.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app