How AI Engineers Run Local LLMs Inside Docker
For AI/ML engineers · Based on ByteMonk Docker-to-AI Infrastructure Framework
// TL;DR
AI/ML engineers use this framework to embed a locally running LLM into a Docker stack with an OpenAI-compatible API — switching from OpenAI by changing only the base URL, with no API key, no per-request cost, and no data leaving the machine. It covers Docker Model Runner (Compose provider syntax, GPU via Metal or CUDA), Ollama on the host via host.docker.internal, and connecting AI agents to external tools through Docker's 270+ MCP servers with credentials as Docker secrets. Use it for privacy-sensitive, cost-sensitive, or offline AI features.
Why run an LLM locally instead of calling a cloud API?
Because you eliminate per-request cost, API keys, and data leaving the machine. Docker Model Runner (shipped April 2025) runs models locally with llama.cpp and GPU acceleration — Metal on Apple Silicon, CUDA on Nvidia. Crucially, it exposes an OpenAI-compatible API. Any service already pointing at OpenAI switches by changing only the base URL. That makes local AI a drop-in for prototypes and privacy-sensitive workloads.
How do you add a local LLM to a Compose stack?
Two paths, depending on resources:
- Docker Model Runner — declare the model as a `provider` dependency in `compose.yaml`. Your services call it via the internal endpoint using the OpenAI-compatible API. Everything stays inside the Docker workflow.
- Ollama — run Ollama on the host and reference it from inside containers via `host.docker.internal`. Lower resource footprint, often better for development.
Either way, set at least a 45-second timeout on the first call. The first call loads the model into memory and is slow; once resident, subsequent calls are fast. Too short a timeout is the most common first-call failure.
What does an AI microservice look like in this stack?
Create an AI service folder with, say, a FastAPI app and its own Dockerfile. Register it in `compose.yaml`. A summarise endpoint validates input, constructs a prompt, calls the model over HTTP at the internal endpoint (or `host.docker.internal` for Ollama), and returns generated text. No external API key, no data leaving the machine. Because it's just another service, it joins the same user-defined network and reaches other containers by name.
How do you connect the AI agent to real tools?
Use the Model Context Protocol (MCP) — Anthropic's standard, described as USB-C for AI agents. Docker's MCP catalog offers 270+ containerised servers (GitHub, Stripe, MongoDB, Neo4j, and more). Enable one with `docker mcp server enable
For a workflow like 'summarise this issue and create a tracking ticket,' your AI service (1) calls the local LLM for a summary, (2) sends the summary and task details to the MCP Gateway, and (3) the Gateway forwards to the external tool and returns the created resource.
Why add the MCP Gateway when running multiple servers?
The MCP Gateway presents a single unified endpoint to your agent, logs every tool call, enforces access controls, and blocks suspicious requests before they reach the actual tool. Configure your AI service with the Gateway URL as an environment variable. It centralises auth, secrets, and governance instead of wiring each server individually — essential once you have more than one MCP integration.
How do you keep the AI image production-ready?
Use a multi-stage build so the runtime stage has no build tools. Switch the base to a Docker Hardened Image for 95%+ fewer CVEs. Add a non-root user and switch to it so an exploited model-serving process doesn't inherit root. Run `docker scout` on the image before shipping.
Next step: Add a local-LLM microservice to your existing stack using Model Runner's provider syntax, wire it to one MCP server through the Gateway with credentials as Docker secrets, and confirm the whole flow runs with no external API key.
// FREQUENTLY ASKED QUESTIONS
Can I switch my existing OpenAI app to a local model without rewriting it?
Yes — Docker Model Runner exposes an OpenAI-compatible API, so any app already pointing at OpenAI switches by changing only the base URL. No code rewrite, no API key, no per-request cost, and no data leaving the machine. This makes local models a genuine drop-in replacement for prototypes and privacy-sensitive workloads.
Why does my first inference call keep timing out?
The first call loads the model into memory, which takes significantly longer than later calls when the model stays resident. Set at least a 45-second timeout on the first call. This is the recommended minimum and is especially important with Ollama referenced via host.docker.internal, where cold-start loading is the usual cause of first-call failures.
Is running MCP servers in containers actually safe?
Yes — each MCP server runs in its own isolated container and cannot reach the host filesystem or sibling containers it isn't authorised to reach. Credentials are injected as Docker secrets and never appear in application code. A compromised server is contained and can't pivot to your host or other services, which is the core security benefit.