37signals icon37signalsOct 2, 2026 ~8 min source read

Warming the Puma master before it forks cut Basecamp’s post-deploy queues

Basecamp reduced multi-hundred to multi-thousand request backlogs at deploy time by routing signed-in requests through the Puma master process before forking workers, avoiding repeated cold work across 63 single-threaded workers.

Warming up the Puma master before it forks

Share this story

Send the public story page.

Useful takeaways from this story.

Cold work repeated by each worker (YJIT compilation, compiled templates, schema cache reads, inline caches) caused large deploy-time slowdowns when 63 single-threaded workers started serving traffic at once.

Running signed-in requests through the already-loaded Puma master before it forked workers warmed shared caches and reduced per-request CPU spikes and request queue lengths.

Basecamp runs Puma in cluster mode with preload_app! and 63 single-threaded workers per 48-core host. The master process loads the app once, then forks workers that share memory copy-on-write. That design reduces memory costs, but it doesn't eliminate per-worker first-use work: anything initialized on first use after fork runs independently in each worker.

What was being done repeatedly by workers

  • YJIT compilation: methods compiled only after several calls, so the master calls very little during boot and workers compile hot methods independently.
  • Compiled templates: the first render of an ERB template compiles it into a Ruby method.
  • Inline caches and memoized values used by Ruby, Rails, and the application.

In a test environment with YJIT enabled, a cold first request to a project page took 652 ms, of which 151 ms was YJIT compiling. The same request to a warm process took 28 ms. In production, CPU time per request peaked around 200 ms while traffic moved to the new container, versus roughly 30 ms once workers were warmed.

Why the architecture magnified the problem

Early mitigations that were tried and ruled out included pre-opening worker database connections in before_worker_boot. That change made no difference, indicating the issue wasn't merely connection setup. The root cause was first-use work that happens only when the code path is exercised.

They cut post-deploy request queues by running signed-in requests through the already-loaded Puma master before it forked workers. Doing this triggers YJIT compilation, template compilation, schema reads, and other memoizations in the master process. The workers inherit those warmed resources via fork and copy-on-write, so they don't repeat the costly first-use work when they start serving traffic.

On the busiest hosts this adjustment reduced the CPU spikes and long request queues seen during deploys, avoiding the need for immediate hardware additions. The change relies on the preload_app! + fork approach and leverages the master's single-instance warm-up to amortize cold-start costs across all workers.

More context around this story.

Hardly Promethean
Jardo iconJardoSep 29, 2026

Hardly Promethean

Originally appeared on Jardo.dev: Blog . Last week I published What About Rails , a dive into DHH's Rails World keynote. Smarter people than me had interesting things to say about it: You have the most powerful tool you ever had, you have become a 1000x maker, and you can't think of how to make your stack 10x better? J

What About Rails?
Jardo iconJardoSep 24, 2026

What About Rails?

Originally appeared on Jardo.dev: Blog . David Heinemeier Hansson is, for better or worse, still in charge of Ruby on Rails. I'd love to stop paying attention to him, but I build applications with Rails, so his actions affect me and my clients. Yesterday, he gave the opening keynote at Rails World 2026, where he laid o

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app