# Why Polars feels faster
# The core mental model: expressions + lazy execution
The simplest way to see the model is the difference between read_csv and scan_csv. read_csv reads a file into memory immediately. scan_csv reads only the header and returns a lazy frame. Every chain of transformations after scan_csv is a declarative description of intent. Nothing executes until you call collect. That delay gives the optimizer room to push filters down to the file and read only necessary columns.
When a dataset is larger than available memory, use collect(engine="streaming"). Instead of materializing the full frame, streaming collect processes the file in chunks so you can work with data that doesn't fit in memory.
# Use expressions, not imperative loops
Polars shines when you express transformations as expressions (select, filter, with_columns). Those expressions are combined and optimized before any data moves. Common verbs you'll use:
- select, filter, with_columns for column-level transformations
- group_by and agg for grouped summaries
- pivot and unpivot for reshaping
- when / then / otherwise for conditional logic
Because these are expressions, Polars can reorder and fuse operations to minimize intermediate data and I/O.
# Window-style computations with over
The over construct runs an aggregation per group but returns a value for every row. That means you can compute each region's share of its group total, or a rank within a category, using a single expression—no manual groupby-and-join required. It reads like a normal column expression and simplifies per-row group calculations.
# Important type behavior: null vs NaN
Polars uses two different states: null for missing values and NaN as a real floating-point value. They have different methods for handling and are not interchangeable. Expect confusion if you assume they behave the same.
# Joins, semi/anti joins, and namespaces
Polars supports the usual joins and also semi and anti joins, which filter rows based on another frame without widening your result. Type-specific namespaces—.str for string ops and.dt for datetime ops—expose focused operations for those types.
# Output paths and interoperability
# Practical takeaway
# Where to go next
Download the KDnuggets Polars cheat sheet for concise reference examples of scan_csv vs read_csv, over usage, the core verb set, join variants,.str/.dt operations, and output methods like sink_parquet and to_pandas.