Amazon iconAmazonSep 3, 2026 ~6 min source read

How to deploy a customer-operated LiteLLM gateway for OpenAI Codex on Amazon ECS and Amazon Bedrock

A concise how-to and decision guide for placing a LiteLLM gateway on Amazon ECS (Fargate) between a developer’s local OpenAI Codex agent and an OpenAI model hosted on Amazon Bedrock, so teams can centralize model routing, scoped identities, budgets, rate limits, and telemetry.

Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock

Share this story

Send the public story page.

Useful takeaways from this story.

Deploy LiteLLM on Amazon ECS with Fargate to act as a centralized gateway that routes Codex Requests to OpenAI models on Amazon Bedrock while keeping local tool execution on developer workstations.

Reference deployment components include ALB (and optional AWS WAF), Amazon RDS for PostgreSQL, Secrets Manager and KMS for keys, CloudWatch for logs/metrics, and Amazon ECR for the gateway image.

# Overview This brief explains how to place a customer-operated LiteLLM gateway on Amazon Elastic Container Service (Amazon ECS) using AWS Fargate, connect it to an OpenAI model via Amazon Bedrock, and configure OpenAI Codex to route inference through the gateway's Responses API. The approach centralizes model access controls and usage telemetry while keeping Codex's local task and tool-execution loop on developer workstations.

# Why use a LiteLLM gateway between Codex and Bedrock

Direct access to Amazon Bedrock is lower complexity when AWS-native identity, IAM policies, and CloudTrail logs already meet requirements. The trade-off for a customer-operated gateway is additional operational responsibility: you own gateway availability, database lifecycle, upgrades, incident response, and capacity planning.

# Architecture and request flow The pattern places LiteLLM between Codex and Amazon Bedrock. A single model turn follows five steps:

  • Codex sends the current task context and available tool definitions to the gateway's /v1/responses endpoint.
  • An Application Load Balancer (ALB) and optional AWS WAF apply network and web-layer controls and forward the request to LiteLLM running on Fargate.
  • LiteLLM authenticates the caller, validates the configured model and consumption policy, and uses its ECS task role to invoke the approved model on Amazon Bedrock.
  • If the model requests a tool, Codex runs it locally under its sandbox and approval policy, then sends results back through LiteLLM on the next Responses request.

# Reference deployment components

  • Amazon ECS on AWS Fargate for the gateway runtime.
  • Application Load Balancer (ALB) and optional AWS WAF for network and web-layer protections and source-IP rate limiting.
  • Amazon RDS for PostgreSQL to store LiteLLM state, usage, and budget data.
  • AWS Secrets Manager and AWS KMS for gateway and scoped-key storage.
  • Amazon CloudWatch logs, CloudWatch Container Insights, alarms, and deployment health metrics.
  • Amazon ECR for storing an immutable gateway image.

# Functional validation and features to test

# When to choose direct IAM Identity Center or a managed gateway Use direct Bedrock access when your security and auditing requirements are satisfied by AWS-native identity and IAM controls. Use a customer-operated LiteLLM gateway when you need uniform controls across teams or providers, finer-grained scoped keys and budgets, or centralized routing/fallback. A managed gateway offering such as Portkey is an alternative if you prefer offloading operational responsibility.

# Prerequisites (deployment checklist)

  • AWS account with permissions to create VPC, ECS, ELB, RDS, ECR, WAF, IAM, KMS, Secrets Manager, and CloudWatch resources.
  • Access to the selected OpenAI model on Amazon Bedrock in the deployment Region.
  • AWS CLI v2 with authenticated profile, Docker with Buildx, Codex CLI, and Python 3 for the walkthrough and scripts.

More context around this story.

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app