← Back to AI Tool Lab

Lab Notebook: From Sync to Async in the Image Generation Service

Initial observation

I built a small FastMCP service for image generation with a single tool called generate_image. The tool accepted a prompt, called an external image provider, and returned a base64-encoded image. It worked well for small, fast models. When I switched to higher-quality models like flux-dev, the tool became unusable.

The provider routinely took between forty-five seconds and three minutes to return an image. During that time, the FastMCP/FastAPI event loop was blocked. Health checks stopped responding, reverse proxies dropped idle sockets, and the calling agent had no way to know whether the job was still alive.

Reproducing the failure

My first test, test_generate_image_basic_call, hung indefinitely. I could not tell whether the hang was in the MCP transport, the provider API, or the model itself. To get visibility, I split the problem into three separate tests.

  • test_mcp_client_connection used a strict 10-second timeout to confirm the MCP server was discoverable.
  • test_provider_direct called the provider's API without the MCP layer, proving that credentials, payload shape, and model availability were correct.
  • test_async_with_timeout wrapped the full tool invocation in asyncio.wait_for so a hang produced a clean TimeoutError instead of an infinite block.

Once these tests were in place, the diagnosis was clear: the provider was not broken; it was simply slower than the synchronous timeout model allowed.

Design of the job manager

I introduced an in-memory job manager. Each image generation request is assigned a UUID and follows a state machine. The states are not arbitrary labels; they map directly to the lifecycle of the underlying provider call.

The job manager runs on the main event loop and is updated from a worker thread through asyncio.run_coroutine_threadsafe. The blocking provider call itself runs inside asyncio.to_thread:

await loop.run_in_executor(
    None,
    functools.partial(self._provider_generate, job_id, prompt, model)
)

This keeps the event loop free for status requests, health checks, and new tool invocations.

MCP tool surface

I replaced the single generate_image tool with three tools:

  • generate_image_async(prompt, model, size, quality) → returns job_id.
  • get_image_generation_status(job_id) → returns state, progress, and message.
  • get_image_generation_result(job_id) → returns the base64 image if state == COMPLETED, otherwise an error.

The sequence diagram reflects a key design decision: the agent, not the server, decides how often to poll and when to fetch the result. This gives the agent control over context-window usage and patience.

Progress estimation without a streaming provider

The provider does not send progress events. I needed a way to keep the agent informed while the image was still generating. I collected historical run durations per model and used them to estimate completion.

My method:

  • On each completed job, store the duration in a per-model ring buffer.
  • When a new job starts, compute a weighted average of the last ten durations.
  • Report progress = elapsed / expected_total, clamped to 0.01 at start and 0.99 while the job is still running.
  • If the actual duration exceeds the estimate, switch to an asymptotic curve that approaches 0.99 slowly, so the progress bar never appears to freeze at 100%.

This is not an exact progress measurement, but it is enough for the agent to know the job is advancing and for a human observer to trust the system.

Concurrency test

I wrote test_concurrent_async_jobs to verify that the service could handle multiple image requests at once. The test launched six simultaneous jobs and polled each until completion.

The first run exposed a subtle ordering bug. Because all jobs were created in the same event-loop tick, their QUEUED timestamps were identical, and the queue ordering became non-deterministic. I added a monotonic created_at counter to the job object to preserve insertion order. The second run passed.

Encoding and final delivery

Once the provider returns image bytes, the job enters the ENCODING state. I support PNG and JPEG outputs. PNG is the default because it preserves quality, but for prompts that produce large photographic images I offer an optional format='jpeg' argument.

Base64 encoding is performed before the result is stored in the job manager. This means get_image_generation_result is a fast lookup rather than a repeated conversion.

The result object returned to the agent includes:

  • image: the base64-encoded image.
  • format: png or jpeg.
  • width and height: parsed from the image headers.
  • generation_time_seconds: wall-clock time from QUEUED to COMPLETED.

Conclusion

The original synchronous design failed because it treated a slow, external process as if it were a local function. The asynchronous job model turned image generation into a managed lifecycle with status, concurrency, and failure recovery. The service is now reliable enough for real agent workloads, and the test suite gives me confidence before any change.