Lab Notebook: Background Replacement with an Async MCP Image Service
Why this problem matters
I have been running image-generation tasks through an MCP service for a while, and one request keeps coming up: "replace the background behind this subject." It sounds like a single model call, but in practice it is a composition problem. The model has to keep the foreground intact, generate a plausible new scene behind it, and match lighting, scale, and perspective. The slowest part is not the prompt; it is the model's internal inference and encoding loop.
This matters because MCP tools are usually called synchronously. The agent sends a request and waits. If the model takes more than a few seconds, the transport layer starts dropping connections. I needed to understand whether I could make background replacement reliable enough to hand to an autonomous agent.
Hypothesis and first attempt
My first hypothesis was simple: a synchronous generate_image tool should work if I choose a fast enough provider and keep the prompt short. I built a thin FastMCP server that exposed one tool:
@mcp.tool()
async def generate_image(prompt: str) -> str:
...The tool called the provider with requests.post, blocked until the response arrived, base64-encoded the image, and returned it. For trivial prompts like "a red apple on a white table" this finished in under five seconds. For background replacement prompts, it failed within minutes.
The failure mode was not a provider error. The FastMCP/FastAPI event loop froze because the blocking network call ran in the main thread. Health checks stopped responding, and the reverse proxy closed the idle connection with a 504 Gateway Timeout. The agent had no feedback. It could not tell whether the image was still generating, had failed, or had been silently dropped.
Switching to a job lifecycle model
I decided to stop treating image generation as a function call and start treating it as a managed job. The key idea is that the MCP tool returns immediately with a job_id, and the agent polls a separate status tool while the work runs elsewhere.
I defined a strict status machine:
Each state has a single responsible function in job_manager.py:
QUEUED: the request has been accepted and is waiting for a worker slot.STARTING: a worker thread is being allocated.GENERATING: the blocking provider call is running insideasyncio.to_thread.ENCODING: the returned image is being converted to base64 and wrapped for the MCP result format.COMPLETEDorFAILED: the result is ready to return.
The tools became:
generate_image_async(prompt, background_prompt, foreground_prompt)→ returnsjob_id.get_image_generation_status(job_id)→ returns state and progress.get_image_generation_result(job_id)→ returns the base64 image once completed.
This separates the transport concern from the compute concern. The MCP server stays responsive because the blocking call is off the event loop.
Prompt design for background replacement
Background replacement has an extra prompt-engineering challenge. The agent usually supplies a foreground description and a desired background scene. I found that the provider produced better results when I normalised both into a single composite prompt rather than passing two separate strings.
My template became:
Keep the following foreground exactly as-is and place it in front of the described background.
Foreground: {foreground}
Background: {background}
Match lighting, perspective, and scale. Preserve all edges.I also added negative prompts for common failure modes: "do not distort fingers," "do not blur text," "do not change the subject's clothing." The provider does not always respect these, but the failure rate dropped noticeably.
Progress without provider progress
The provider API does not stream progress. Once a job enters GENERATING, the agent receives no updates until the call returns. I wanted a smooth polling experience, so I built an estimated progress interpolator based on historical run times.
The method is crude but useful. I store the last N durations for each model and compute an exponential moving average. When a new job starts, I predict the total time and report progress as elapsed / predicted_total. If the actual run exceeds the prediction, I switch to a slower asymptotic curve so the bar never appears to stall at 99%.
This is not accurate in an absolute sense, but it gives the agent enough information to decide whether to keep waiting or cancel.
Concurrency and thread safety
I wrote test_concurrent_async_jobs to understand how the service behaves under load. The test fires six background replacement requests at once and polls each until completion.
The first run revealed a thread-safety bug in my progress callback. The worker thread calls progress_callback when a partial result is ready, but asyncio.run_coroutine_threadsafe requires a running event loop in the main thread. I had to capture the main loop at startup and use it as the target for all cross-thread callbacks.
self._loop = asyncio.get_running_loop()
async def progress_callback(self, job_id: str, progress: float):
await asyncio.run_coroutine_threadsafe(
self._update_progress(job_id, progress), self._loop
)Once that was in place, the test passed. The service now accepts multiple concurrent background replacement jobs, queues them, and tracks each independently.
Encoding and delivery
The final image must travel back through the MCP result channel. Providers usually return PNG or JPEG bytes, but some return URLs that expire quickly. To avoid a second download step, I encode the image as base64 before returning it.
The trade-off is message size. A 1024x1024 PNG can be a few hundred kilobytes after base64. I added an optional quality parameter so the agent can request a lower-resolution preview if bandwidth or context-window size is a concern.
Test results
I now run a layered test suite before any change to this service:
test_mcp_client_connection— proves the MCP transport layer is healthy within ten seconds.test_provider_direct— bypasses the MCP server and confirms the provider accepts the prompt and API key.test_async_job_lifecycle— walks a single job through every state.test_concurrent_async_jobs— fires multiple background replacements and verifies independent tracking.
The background replacement flow has moved from "occasionally works for simple prompts" to "reliably handles composite prompts with status visibility and concurrency." The main lesson is that the interface to a slow model must be asynchronous, even if the first version feels like a single function call.