Google iconGoogleSep 29, 2026 ~7 min source read

GKE Agent Sandbox cuts agentic RL sandbox startup from minutes to seconds, boosting research throughput

Google’s GKE Agent Sandbox and accompanying SDK reduce time-to-first-command by up to 45x, shrink worst-case sandbox waits from minutes to under 10 seconds, and reduce control-plane churn by roughly 3x — addressing the main infrastructure bottlenecks for large-scale agentic reinforcement learning and evaluation.

Accelerating agentic RL and evaluation research velocity with 45x faster GKE Agent Sandbox

Share this story

Send the public story page.

Useful takeaways from this story.

Time-to-first-command improved 10x–45x (cold-starts down to 1–9 seconds), keeping expensive accelerators active.

An in-place pod recycling strategy yields ~3x fewer pod creations during rollout bursts, stabilizing the Kubernetes control plane.

SandboxWarmPool plus GKE Image Streaming make thousands of large, distinct OCI images (e.g., 4,578 R2E images) usable at scale without crippling image-pull delays.

Agentic reinforcement learning (RL) scales poorly on standard Kubernetes setups because thousands of ephemeral CPU sandboxes are required alongside GPUs that run model-inference. When each sandbox takes minutes to provision, expensive accelerators sit idle and synchronous RL steps are blocked by the slowest sandbox. Google introduces GKE Agent Sandbox, an open Kubernetes primitive plus an RL orchestration SDK and native integrations, to remove this bottleneck for large-scale agentic workloads.

Infrastructure pain points for agentic RL

Agentic RL workloads create three practical, recurring problems:

  • Accelerator idle costs: time-to-first-command (TTFC) for CPU sandboxes can be tens of seconds to minutes, which stalls GPU-based policy execution until all sandboxes are ready.
  • Massive image cardinality: agentic evaluations often require thousands of distinct, multi-gigabyte OCI images (the example cited includes 4,578 R2E images). Multiple rollouts per image multiply sandbox counts (e.g., four rollouts → 18,312 tasks), producing massive image-pull pressure and storage friction.
  • Control-plane saturation: bursty rollouts create tens of thousands of ephemeral pods, saturating the Kubernetes API server and causing pod state errors, queuing, and false node-health evictions.

GKE Agent Sandbox is a purpose-built sandbox layer that combines several primitives:

  • SandboxWarmPool: keeps pre-initialized, healthy sandboxes ready to eliminate cold-start overhead.
  • GKE Image Streaming: enables efficient use of high image cardinality workloads by streaming large container images instead of full pre-pulls.
  • Agent Sandbox RL orchestration SDK: an in-cluster driver and controller that claims warm pods and implements in-place recycling of pods across rollouts.
  • Secure runtime options (example test setup used a 10-node gVisor sandbox pool).

Google deliberately stress-tested the system with agentic benchmarks such as SWE-bench and high-cardinality datasets. The tests focused on burst rollouts, many large images, and control-plane behavior. Key measured outcomes reported:

  • The in-place recycling strategy reduced pod creation by about 3x, lowering API server churn and improving control-plane stability.

Practical implications for labs and platforms

For teams running large-scale agentic RL or evaluation:

  • Shorter sandbox startup means GPUs spend more time executing policies and less time waiting, improving utilization and reducing cost-per-experiment.
  • Lower tail latency removes a common blocking factor in synchronous batch training, allowing larger batch sizes and faster iteration cycles.
  • Support for thousands of images through image streaming reduces storage and network bottlenecks associated with SWE-bench–style workloads.
  • Fewer pod creations smooths the Kubernetes control plane under bursty, large-scale workloads.

Complementary tooling and ecosystem notes

GKE Agent Sandbox is presented alongside the Agent Sandbox RL SDK and native integrations for popular RL gyms and harnesses. Related Google announcements in the same timeframe include Agent Substrate (higher-density agent runtime) and case studies reporting substantial cost reductions for customers using these primitives.

GKE Agent Sandbox targets the sandbox-level bottleneck that slows agentic RL research. By combining warm pools, image streaming, and pod-reuse strategies, the solution reduces cold-start latency and control-plane churn, enabling more efficient, large-scale parallel rollouts without waiting minutes to begin execution.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app