# Problem: data silos and the access gap Enterprise data commonly lives in specialized systems that require different expertise: SQL for databases, Athena or Trino for S3 batch data, streaming tooling for Kinesis, and distinct APIs for SaaS applications. That concentrates access knowledge with data engineers and forces business users to file tickets or rely on stale dashboards when they need ad-hoc answers.
Rather than moving or pre-integrating all data, the blog proposes letting AI agents query each system directly through a standardized protocol. The Model Context Protocol (MCP) wraps diverse data sources behind a uniform interface for tool discovery, invocation, and response handling. Agents use MCP servers to find which tool to call, invoke it, and handle results without needing to know per-source authentication, API, or query language details.
# How the reference architecture fits together The reference architecture in the post targets a streaming media use case but generalizes to mixed sources. Key components:
- Data layer: batch datasets on Amazon S3, real-time telemetry in Amazon Kinesis, and transactional customer data in a relational store such as Amazon Aurora. Batch datasets in the demo are generated with Python scripts and streaming telemetry produced by AWS Lambda.
- Governance and compute: AWS Glue Data Catalog and query engines such as Amazon Athena (and Trino in general) for accessing S3 data.
- Generative AI layer: AI agents running on Amazon Bedrock AgentCore (for example, a Strands agent) that reason over natural-language questions and decide which MCP servers to call.
- Gateway: Amazon Bedrock AgentCore Gateway aggregates multiple MCP servers behind a single endpoint and handles tool discovery, authentication, and routing.
Request flow example (user question to answer):
- 1User submits a natural-language question through a React app served via CloudFront and S3.
- 2Amazon Cognito authenticates the user and passes an identity token to the agent layer.
- 3A Strands agent running on AgentCore reasons about the question and selects MCP servers to query.
- 4AgentCore Gateway routes the request to the appropriate MCP server(s).
# Three reference access patterns
- MCP servers expose a catalog (for example, via AWS Glue Data Catalog) that centralizes discovery. Agents consult the catalog to decide which source to query. This is useful when teams want a single discovery surface for many datasets.
- Agents query MCP servers that directly wrap the source with minimal cataloging. Use this when sources are well-known to the agent or when teams prefer source-specific access and tooling.
- Combine a catalog for discovery with direct source MCP servers for execution. Discovery routes the agent to the source, and execution happens against the source MCP server to retrieve live or specialized data.
# Demo and deployability The blog's demo uses synthetic data and provides a GitHub repository with source code so readers can deploy the reference architecture. Prerequisites include an AWS account with AgentCore access, AWS CLI configured, Python 3.10+, and Docker or Finch.
# Practical implications for teams Teams can reduce one-off engineering requests by enabling agents to reach disparate systems. MCP servers and AgentCore Gateway define a consistent control plane for discovery, authentication, and routing, while agents handle orchestration and result composition. The approach preserves data in place and leverages existing query engines and catalogs.
# Bottom line