Serpapi iconSerpapiSep 17, 2026 ~7 min source read

Vessel 0.3: Ruby gets a full-featured crawling framework similar to Scrapy

After four years away, Vessel returns as a rewritten crawling framework for Ruby with pluggable drivers, a fields API, middleware, proxy rotation, retries, cookies, callbacks, and a CLI for generating projects.

Vessel 0.3: Ruby finally gets its Scrapy crawling framework

Share this story

Send the public story page.

Useful takeaways from this story.

Vessel 0.3 rebuilds the project into a cohesive crawling framework that handles scheduling, concurrency, deduplication, retries, and pipelines for Ruby developers.

The release introduces pluggable drivers (real Chrome or plain HTTP), a fields API, middleware pipeline, proxy rotation, cookie support, retries, and callbacks.

# What Vessel 0.3 delivers

Vessel 0.3 is a major relaunch of the Ruby crawling project that had been quiet for several years. The internals were rewritten and the feature set expanded to match what a full crawling framework should provide: request scheduling, concurrent fetching, URL deduplication, retries, and a processing pipeline for extracted items.

# Why this matters for Ruby developers

# Core features in this release

  • Pluggable drivers: choose between running a real Chrome browser or using plain HTTP requests depending on page complexity.
  • Fields API: a structured way to declare and extract data fields before sending them through the pipeline.
  • Middleware pipeline: insert custom middleware to manipulate requests or responses at defined stages.
  • Proxy rotation and cookies: built-in support to rotate proxies and manage cookies during crawls.
  • Retries and callbacks: configurable retry behavior and callbacks for handling events and extracted requests.
  • CLI project generator: scaffolds whole projects to get you started quickly.

# How Vessel works in practice

At the highest level you subclass Vessel::Cargo, declare the domain and start URLs, and implement handler methods to parse pages and yield new requests. Vessel schedules those requests, enforces deduplication so each URL is visited only once, runs concurrent fetches to keep the crawl efficient, retries transient errors, and forwards extracted items through a processing pipeline for storage or further processing.

# When to choose a driver

Use the Chrome driver when pages rely on client-side JavaScript that must be executed. Use the HTTP driver for faster, lighter crawls when you can fetch and parse raw HTML. Pluggable drivers give you the flexibility to mix approaches across different targets.

# Typical workflow and benefits

  1. Generate a project with the CLI.
  2. Subclass Vessel::Cargo and configure domain and start URLs.
  3. Write handler methods to extract fields via the fields API and yield follow-up requests.
  4. Configure middleware, proxies, retry policies, and pipeline stages.

The benefit is practical: less boilerplate, fewer ad-hoc integrations, and a single place to control crawl behavior at scale.

# Short comparison to ad-hoc toolchains

Without a framework, Rubyists combine Nokogiri, Ferrum, Mechanize, Faraday, and custom queueing or deduplication code. That works for small tasks but becomes brittle for large crawls. Vessel bundles those concerns into one idiom and automates the crawl lifecycle.

# What to expect next

# Bottom line

More context around this story.

Shopify's structural linter for Ruby codebases
Rubyweekly iconRubyweeklySep 10, 2026

Shopify's structural linter for Ruby codebases

Originally appeared on Ruby Weekly . #​816 — September 10, 2026 Read on the Web Ruby Weekly RubyKaigi's Talks Are Finally Online RubyKaigi , the conference where Ruby's core team gathers each year, took place in Japan this April, but the talks have just made it onto YouTube with 68 videos in this playlist . Some hi

How to scrape Zillow
Serpapi iconSerpapiSep 4, 2026

How to scrape Zillow

Scrape Zillow real estate listings with SerpApi in Python, JavaScript, Ruby, or with just plain cURL. Retrieve homes for sale, rentals, and recently sold properties in structured JSON.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app