← Back to AI Tool Lab

Lab Notebook: Remote MCP Security and Service Boundaries

The problem

When I deployed the image generation service as a remote MCP endpoint, I ran into a design question that is not covered by the protocol spec: where should the root info page live? The MCP protocol runs on one path, but humans and agents need to discover connection details, parameters, and current status. Serving that information on the same host introduces real security trade-offs around service boundaries.

I modelled three architectures and evaluated each against three criteria:

  • operational visibility: can the page show live metrics?
  • attack surface: does it add routes or ports that could be abused?
  • deployment complexity: how many moving parts are needed?

Approach 1: pure static serving

The safest and fastest option is to let a reverse proxy serve a static HTML file at the root route, completely bypassing the Python application.

Advantages:

  • The static file cannot crash the MCP service.
  • It cannot access application state.
  • It cannot leak secrets.
  • It is cheap to serve.

Disadvantages:

  • The page cannot show real-time metrics such as queue depth, API key validity, or newly added tools.
  • Documentation is decoupled from code, so deployment requires pushing the HTML separately.

This is the most secure option, but it loses the operational story.

Approach 2: dual-port architecture

A common production pattern is to bind two ports: one public port for MCP protocol traffic and one internal port for diagnostics, metrics, and the info page.

Diagnostics ports are normally restricted behind firewalls or VPNs. If the info page lives on the diagnostics port, the public cannot see it. If the info route is reverse-proxied to the public port, there is a risk of misconfiguring the proxy and accidentally exposing sensitive /diagnostics endpoints.

Advantages:

  • Strong separation between public protocol and internal diagnostics.
  • Live metrics are available on the internal port.

Disadvantages:

  • Two ports to configure, monitor, and secure.
  • Proxy misconfiguration can defeat the boundary.
  • Agents on the public internet cannot see the info page without being inside the firewall.

Approach 3: embedded route

The approach I adopted for the image generation service is to register a custom GET / route directly inside the MCP application itself, on the same public port as the protocol.

@mcp.custom_route("/", ["GET"], name="root_info")
async def root_info_handler(request: Request):
    job_metrics = await client.job_manager.get_metrics()
    return HTMLResponse(content=html)

The handler injects high-level, read-only metrics into an HTML template. It does not expose mutation capability or internal diagnostics. It also keeps deployment simple because the HTML travels with the application container and needs only one reverse-proxy rule.

Advantages:

  • Immediate, actionable context for public clients.
  • Single port and single application.
  • HTML and metrics stay in sync with the deployed code.

Disadvantages:

  • Any vulnerability in the root handler could affect the MCP service.
  • The handler must be strictly read-only and carefully validated.

Comparing the three

Pure static serving wins on performance and security. Dual-port wins on separation but introduces proxy and operational risk. Embedded wins on visibility and maintenance simplicity.

My choice

System design is the art of acceptable compromise. For most remote MCP services, the embedded route provides the best balance: enough visibility to keep the service reliable, without enlarging the attack surface or complicating deployment.

I hardened the embedded route with the following rules:

  • It accepts only GET.
  • It returns only HTML; no JSON that could be used as a machine-readable API surface.
  • It reads only from the job manager's public metrics, never from internal diagnostics.
  • It does not accept query parameters, so injection vectors are limited.

If the service becomes high-traffic or high-sensitivity, I will revisit the dual-port option. For now, the embedded route gives me the operational clarity I need with a single public endpoint.