# What Vessel 0.3 delivers
Vessel 0.3 is a major relaunch of the Ruby crawling project that had been quiet for several years. The internals were rewritten and the feature set expanded to match what a full crawling framework should provide: request scheduling, concurrent fetching, URL deduplication, retries, and a processing pipeline for extracted items.
# Why this matters for Ruby developers
# Core features in this release
- Pluggable drivers: choose between running a real Chrome browser or using plain HTTP requests depending on page complexity.
- Fields API: a structured way to declare and extract data fields before sending them through the pipeline.
- Middleware pipeline: insert custom middleware to manipulate requests or responses at defined stages.
- Proxy rotation and cookies: built-in support to rotate proxies and manage cookies during crawls.
- Retries and callbacks: configurable retry behavior and callbacks for handling events and extracted requests.
- CLI project generator: scaffolds whole projects to get you started quickly.
# How Vessel works in practice
At the highest level you subclass Vessel::Cargo, declare the domain and start URLs, and implement handler methods to parse pages and yield new requests. Vessel schedules those requests, enforces deduplication so each URL is visited only once, runs concurrent fetches to keep the crawl efficient, retries transient errors, and forwards extracted items through a processing pipeline for storage or further processing.
# When to choose a driver
Use the Chrome driver when pages rely on client-side JavaScript that must be executed. Use the HTTP driver for faster, lighter crawls when you can fetch and parse raw HTML. Pluggable drivers give you the flexibility to mix approaches across different targets.
# Typical workflow and benefits
- 1Generate a project with the CLI.
- 2Subclass Vessel::Cargo and configure domain and start URLs.
- 3Write handler methods to extract fields via the fields API and yield follow-up requests.
- 4Configure middleware, proxies, retry policies, and pipeline stages.
The benefit is practical: less boilerplate, fewer ad-hoc integrations, and a single place to control crawl behavior at scale.
# Short comparison to ad-hoc toolchains
Without a framework, Rubyists combine Nokogiri, Ferrum, Mechanize, Faraday, and custom queueing or deduplication code. That works for small tasks but becomes brittle for large crawls. Vessel bundles those concerns into one idiom and automates the crawl lifecycle.
# What to expect next
# Bottom line