# What this is about
# From text to numbers: tokenization and embeddings Before any model computation, input text is split into tokens. A token can be a whole word, a subword, or a character sequence depending on the tokenizer. Tokens are then converted into embeddings: numeric vectors that represent tokens' meanings and relationships. Tokens that appear in similar contexts tend to have similar embeddings, which helps the model reason about word relationships rather than treating words as isolated symbols.
# Transformers and self‑attention: modeling context
# How LLMs are trained Training typically happens in stages.
- Pre‑training: The model ingests large corpora—books, websites, papers, code depending on the dataset—and adjusts millions or billions of parameters to predict text patterns. This stage teaches grammar, common sense associations, and broad language structure.
- Fine‑tuning: After pre‑training, models are fine‑tuned on smaller, task‑specific datasets to improve performance on particular tasks or to follow instructions more reliably.
# How responses are generated: next‑token prediction LLMs generate text by repeatedly predicting the most likely next token given the tokens already produced. The model does not retrieve a stored answer. Instead it uses learned statistical patterns to assign probabilities to candidate tokens, selects one (often using sampling or beam strategies), appends it to the sequence, and repeats. This iterative process yields fluent and context‑aware outputs but is fundamentally probabilistic.
# Why LLMs hallucinate
# Practical implications LLMs are useful for drafting text, answering questions, and assisting complex tasks across domains. Use them for ideation and first drafts, but verify factual claims, especially in high‑stakes contexts. Understand that model outputs reflect statistical prediction based on training data and alignment choices, not a direct connection to external truth sources.
# Takeaway LLMs work by converting text to tokens and embeddings, applying transformer layers with self‑attention to build context, training in stages to learn language patterns and desired behaviors, and generating responses token by token. Their probabilistic generation can produce useful outputs but also hallucinations, so human oversight is required when accuracy matters.