ByteMonk Docker-to-AI Infrastructure Framework

Design, build, and run a fully containerised multi-service application — including locally running AI models and MCP tool integrations — using Docker's complete 2026 toolchain, without touching cloud APIs or external credentials.

// TL;DR

The ByteMonk Docker-to-AI Infrastructure Framework is a complete method for designing, building, and running containerised multi-service applications — including locally running LLMs and MCP tool integrations — using Docker's 2026 toolchain, with no cloud APIs or external credentials required. Use it when you need to containerise a new or existing app, wire up multi-service stacks with Docker Compose, add persistent storage or inter-container networking, harden images for production, or embed a local AI model into a Docker workflow. It carries one mental model throughout: a container is a process with a private view, not a lightweight VM.

// When should you use the Docker-to-AI Infrastructure Framework?

Use this skill whenever you need to containerise a new or existing application, wire up multi-service stacks with Docker Compose, add persistent storage or inter-container networking, secure images for production, or integrate a locally running LLM or MCP tool server into a Docker-based workflow.

// What information do you need before containerising your application?

  • Application typerequired
    What kind of app are you building? e.g. Python API + Node frontend + Postgres DB, or a single service
  • Language and runtimerequired
    Primary language(s) and runtime versions (e.g. Python 3.12, Node 20)
  • Storage requirementsrequired
    Does any service need data to survive container restarts? Which services write data?
  • Service count and dependenciesrequired
    How many services? Which services depend on which? e.g. API depends on DB being healthy
  • AI requirements
    Do you need a locally running LLM? Which model (e.g. Gemma, Phi-3)? Text only or multimodal?
  • External tool integrations
    Do you need MCP servers (e.g. GitHub, Stripe, MongoDB)? Which tools need to be connected to the AI agent?
  • Target environmentrequired
    Development only, or production-bound? Determines security hardening and volume strategy.

// What core principles govern building Docker stacks the right way?

Container = Process with a Private View

A container is not a lightweight VM. It is a process running on the host kernel that uses Linux Cgroups (resource limits) and namespaces (isolated process tree, network, filesystem) to feel like a separate system. Carry this mental model through every decision: it explains startup speed, image size, and networking behaviour.

Layer Caching determines build speed

Every Dockerfile instruction creates a cached layer. If a layer has not changed, Docker skips it. Structure your Dockerfile so slow, infrequent steps (dependency installation) come before fast, frequent steps (copying source code). Reversing this order turns a 2-second build into a 3-minute build.

localhost means the container, not the host

Inside a container, localhost refers to that container's own loopback interface — not the host machine and not any other container. Services on different containers must communicate via a user-defined Docker network, using container names as hostnames resolved by Docker's internal DNS.

Named volumes survive; bind mounts accelerate dev

Named volumes are managed by Docker and outlive any individual container — the right choice for any data that must survive restarts or deletion. Bind mounts point directly at a host directory, giving instant reflection of file changes inside the container — the right choice for fast development iteration. tmpfs lives in memory and disappears on stop — the right choice for sensitive temporary data that must never touch disk.

Publish only what needs external access

Use the -p flag only for ports that genuinely need to be reachable from outside Docker. Databases should almost never have a published port. Internal services stay on the Docker network, keeping the attack surface small and the network layout clean.

Docker hardened images = 95%+ fewer CVEs

Most container vulnerabilities come from the base image, not application code. Docker Hardened Images (DHI) are minimal by design, ship with a full software bill of materials, carry cryptographic provenance, and receive critical CVE patches within 7 days. Use them in production instead of standard base images.

Docker Model Runner = local LLM, OpenAI-compatible

Docker Model Runner (shipped April 2025) lets you run large language models locally inside your Docker workflow using llama.cpp with GPU acceleration (Metal on Apple Silicon, CUDA on Nvidia). It exposes an OpenAI-compatible API, so any app already pointing at OpenAI can switch by changing only the base URL — no API key, no cost per request, no data leaving the machine.

MCP = USB-C for AI agents

Model Context Protocol (MCP), developed by Anthropic, is the standard protocol for AI agents to connect to external tools. Docker's MCP catalog (270+ containerised MCP servers) and MCP toolkit manage auth, secrets, and connections. Each MCP server runs in its own isolated container — a compromised server cannot reach the host filesystem or other containers.

// How do you build a containerised multi-service AI stack step by step?

  1. 1

    Establish the mental model before writing any commands

    Confirm you are thinking 'container = process with a private view of its environment', not 'container = lightweight VM'. Every networking, storage, and debugging decision flows from this. If something feels heavier or more confusing than it should, return here first.

  2. 2

    Verify the three-part Docker setup is working

    Run: `docker pull hello-world`, `docker run hello-world`, `docker --version`. A successful hello-world confirms the Docker client, Docker daemon, and Docker registry connection are all functional. On Linux, add your user to the Docker group to avoid sudo on every command.

  3. 3

    Author the Dockerfile with layer-cache order enforced

    Structure: FROM (use slim or Alpine variant) → WORKDIR → COPY dependency manifest only (e.g. requirements.txt or package.json) → RUN install dependencies → COPY application source → EXPOSE port → CMD. Never copy source before installing dependencies — this defeats layer caching and makes every code-change rebuild slow. Always create a .dockerignore file to exclude node_modules, .git history, and .env files from the build context.

  4. 4

    Apply multi-stage builds for production images

    Use a first stage with all build tools (compilers, bundlers, test frameworks) to produce the compiled artefact. Copy only the final output into a second minimal stage. The runtime stage has no compilers, no test frameworks, no source files — smaller image, smaller attack surface, faster pulls.

  5. 5

    Choose the correct storage strategy per service

    Named volumes for any data that must survive container deletion (databases, uploaded files). Bind mounts for development iteration (source code reflecting instantly inside a running container without image rebuilds). tmpfs for sensitive temporary data (short-lived tokens, one-time secrets) that must never touch disk. Default container writable layer = ephemeral; assume data is gone when the container is deleted unless a volume or mount is configured.

  6. 6

    Create a user-defined Docker network and use container-name DNS

    Never attempt to connect containers via localhost. Place all services that need to communicate on the same user-defined network. Docker's internal DNS resolver maps container names to their current IP automatically. The container name becomes the hostname — e.g. the API reaches the database at `db:5432`, not `localhost:5432`.

  7. 7

    Write the Docker Compose file as the single source of truth for the stack

    One service = one container. Define all services, their build context or image, environment variables (loaded from .env — never hardcoded), volume mounts, network membership, exposed ports, health checks, and depends_on with `condition: service_healthy`. Use `docker compose` (space, no hyphen) — the hyphenated form is officially removed as of 2026. Add .env to .gitignore immediately. Never put actual passwords in the YAML file.

  8. 8

    Encode health checks and startup ordering in Compose

    Without `depends_on: condition: service_healthy`, dependent services (e.g. API) will attempt to start while the dependency (e.g. Postgres) is still initialising, fail to connect, crash, and enter a restart loop. The health check must be defined on the dependency service so Compose has a signal to wait on before starting the dependent.

  9. 9

    Harden images for production using Docker Hardened Images and non-root users

    Switch base images to Docker Hardened Images (DHI) — use `dhi-ctl` CLI plugin or Gordon AI assistant to migrate existing Dockerfiles. Add `RUN adduser --disabled-password appuser` and `USER appuser` to every Dockerfile so the application process never runs as root. Run `docker scout` on every image before shipping to CI/CD; integrate Scout into the pipeline to catch vulnerabilities before production.

  10. 10

    Add a locally running LLM using Docker Model Runner or Ollama

    For Docker Model Runner: declare the model as a provider dependency in compose.yaml using the `provider` syntax; services call it via the internal endpoint with an OpenAI-compatible API by changing only the base URL. For Ollama (lower resource, better for dev): run Ollama on the host, reference it from inside containers via `host.docker.internal`; set a 45-second timeout on the first call to accommodate model load into memory — subsequent calls are fast because the model stays resident.

  11. 11

    Extend the AI layer with MCP servers via the Docker MCP catalog and toolkit

    Enable a server with `docker mcp server enable <server-name>`. Each MCP server runs in its own isolated container — it cannot reach the host filesystem or sibling containers it is not authorised to reach. Credentials are injected as Docker secrets and are never visible in application code. For multiple MCP servers, add the MCP Gateway service: it presents a single unified endpoint, logs every tool call, and enforces access controls before requests reach the actual tool.

  12. 12

    Operate and debug the running stack with standard Compose commands

    `docker compose up --build -d` to start with latest image changes. `docker compose exec <service> sh` to drop into a shell inside a running container for inspection. `docker compose logs -f <service>` for live log tailing. `docker compose down` stops and removes containers — named volumes survive. `docker compose down -v` also deletes volumes (development clean-slate only; never run on production data without explicit intent — this is the command that ends careers).

// What does this framework look like applied to real applications?

A three-tier web app: REST API backend, JavaScript frontend, relational database

Create three services in compose.yaml: DB (official Postgres image, named volume for data persistence, health check on pg_isready), API (build from ./backend Dockerfile, depends_on DB with condition: service_healthy, non-root user, internal network only), Frontend (build from ./frontend Dockerfile using npm ci not npm install, publishes UI port externally). Secrets loaded from .env file kept out of version control. Bring up with `docker compose up --build`. Verify data survives `docker compose down` and `docker compose up` cycle. Use `docker compose down -v` only for clean-slate resets in development.

Adding an AI summarisation microservice to an existing Compose stack

Add an AI service folder with a FastAPI app and its own Dockerfile. The service calls a locally running LLM (via Docker Model Runner endpoint internally, or Ollama on host via host.docker.internal) with a 45-second timeout on the first call. Register the service in compose.yaml. The summarise endpoint validates input, constructs a prompt, calls the model over HTTP, and returns the generated text. No external API key, no data leaving the machine.

Connecting the AI service to an external tool (e.g. issue tracker) via MCP

Enable the relevant MCP server from the Docker MCP catalog. Add MCP Gateway as a service in compose.yaml on an internal port. Configure the AI service with the MCP Gateway URL as an environment variable. Create a new endpoint in the AI service that: (1) calls the local LLM for a summary, (2) sends the summary plus task details to the MCP Gateway, (3) the Gateway forwards to the external tool and returns the created resource details. Credentials stored as Docker secrets, never in application code.

// What mistakes should you avoid when using Docker and Docker Compose?

  • Thinking of containers as lightweight VMs — this mismodel makes networking, storage, and debugging consistently confusing. A container is a process with a private view of its environment.
  • Copying source code before installing dependencies in the Dockerfile, which destroys layer caching and turns fast builds into slow ones on every code change.
  • Omitting a .dockerignore file, causing Docker to send the entire project directory (including node_modules, .git history, and .env secrets) as the build context.
  • Trying to connect containers via localhost — inside a container localhost means that container only; use user-defined networks and container-name DNS instead.
  • Hardcoding secrets or passwords in docker-compose.yaml, which gets committed to version control; always load from a .env file and add .env to .gitignore immediately.
  • Starting dependent services without `depends_on: condition: service_healthy`, causing the API to crash-loop while the database is still initialising.
  • Running `docker compose down -v` without understanding it deletes named volumes — this destroys persistent data and in production is catastrophic.
  • Using the legacy `docker-compose` command (with hyphen) — officially removed as of 2026; the correct form is `docker compose` with a space.
  • Using `npm install` inside Dockerfiles instead of `npm ci` — npm install can silently modify the lockfile; npm ci installs exactly what is locked and nothing else.
  • Running application processes as root inside containers — if the process is exploited, the attacker inherits full system privileges; always add a non-root user and switch to it.
  • Publishing database ports externally with -p — databases should almost never be reachable outside the Docker network.
  • Setting too short a timeout on the first LLM call — the first call loads the model into memory and takes significantly longer; 45 seconds is the recommended minimum.

// What are the key Docker and AI infrastructure terms you should know?

Container = Process with a Private View
The core mental model: a container is not a VM but a host process that uses Linux Cgroups (resource limits) and namespaces (isolated process tree, network, filesystem view) to appear like a separate system.
Cgroups
Linux kernel feature that controls how much CPU, memory, and IO a container process is allowed to consume.
Namespaces
Linux kernel feature that gives each container its own isolated process tree, network interface, and filesystem view, making it appear as a separate environment.
Docker daemon
The background process that does all the actual work: building images, starting and stopping containers, managing networks and volumes. The Docker client sends instructions to the daemon.
Docker Registry
The storage and distribution system for images. Docker Hub is the default public registry; organisations also run private registries.
Layer caching
Docker caches the result of each Dockerfile instruction as a layer. If a layer has not changed since the last build, Docker skips rebuilding it, making builds fast. Ordering instructions from slow-changing to fast-changing maximises cache hits.
Named volume
Docker-managed persistent storage identified by a name. Data survives container deletion and is completely separate from the container lifecycle — the right choice for any data that must persist.
Bind mount
A direct mapping from a host machine directory into a container. Changes on the host are immediately visible inside the container — essential for fast development iteration, not used in production.
tmpfs
In-memory storage that exists only while the container is running. Used for sensitive temporary data (e.g. short-lived tokens) that must never be written to disk.
User-defined network
A Docker network explicitly created for a set of containers. Containers on the same user-defined network can reach each other by container name via Docker's internal DNS resolver.
Docker Compose
A tool that reads a YAML blueprint (compose.yaml) describing all services, their configuration, networks, and volumes, and brings the entire multi-service application up or down with a single command.
depends_on with condition: service_healthy
A Compose directive that makes a service wait until the specified dependency's health check passes before starting, preventing crash-loop restarts during initialisation.
Multi-stage build
A Dockerfile pattern using multiple FROM stages: a build stage with all compilers and tools produces the artefact; a minimal runtime stage receives only the compiled output, resulting in a smaller, more secure image.
Docker Hardened Images (DHI)
Docker's security-focused base images, available free under open source licence since late 2025. Minimal by design, ship with a software bill of materials, cryptographic provenance, and receive critical CVE patches within 7 days. Reported to have 95%+ fewer CVEs than standard base images.
Docker Model Runner
A Docker Desktop feature (shipped April 2025) for running large language models locally using llama.cpp, with GPU acceleration via Metal (Apple Silicon) or CUDA (Nvidia). Exposes an OpenAI-compatible API accessible from inside containers.
MCP (Model Context Protocol)
Anthropic's standard protocol for AI agents to connect to external tools. Described as 'USB-C for AI agents' — one standard protocol, hundreds of compatible tool servers.
Docker MCP catalog
A catalog of 270+ containerised MCP servers (GitHub, Grafana, Stripe, MongoDB, Neo4j, and others) integrated into Docker Hub, each running in its own isolated container with credentials injected as Docker secrets.
MCP Gateway
An additional Docker service for teams running multiple MCP servers. Presents a single unified endpoint to the AI agent, logs every tool call, enforces access controls, and blocks suspicious requests before they reach the actual tool.
Gordon
Docker's AI assistant, integrated into Docker Desktop. By early 2026 it has persistent memory across sessions, context-aware terminal hints on command failure, and the ability to scan and migrate Dockerfiles to Docker Hardened Images.
docker compose down -v
The command that stops and removes containers AND deletes named volumes. Correct for development clean-slate resets; catastrophic if run against production data.
npm ci
The npm command for reproducible installs in Dockerfiles and CI pipelines. Installs exactly what is locked in package-lock.json and nothing else, unlike npm install which can silently update the lockfile.
host.docker.internal
A special DNS name resolvable from inside a Docker container that points to the host machine's IP address. Used when a containerised service needs to reach a process running on the host (e.g. Ollama).

// FREQUENTLY ASKED QUESTIONS

What is the ByteMonk Docker-to-AI Infrastructure Framework?

It's a complete method for building containerised multi-service applications with Docker's 2026 toolchain, including running LLMs locally and connecting MCP tool servers — all without cloud APIs or external credentials. It covers Dockerfiles, Docker Compose, storage, networking, production hardening, Docker Model Runner for local AI, and Docker's MCP catalog, unified under one mental model: a container is a process with a private view.

What is a Docker container actually?

A Docker container is a process running on the host kernel that uses Linux Cgroups for resource limits and namespaces for an isolated process tree, network, and filesystem — making it feel like a separate system. It is not a lightweight VM. This distinction explains why containers start fast, stay small, and network the way they do. Every storage and networking decision flows from this model.

How do I run an LLM locally inside Docker?

Use Docker Model Runner (shipped April 2025), which runs models locally with llama.cpp and GPU acceleration — Metal on Apple Silicon, CUDA on Nvidia. Declare the model as a provider dependency in compose.yaml and call it via the internal endpoint. It exposes an OpenAI-compatible API, so any app pointing at OpenAI switches by changing only the base URL — no API key, no per-request cost, no data leaving your machine.

How do I make containers talk to each other in Docker?

Put all services that need to communicate on the same user-defined Docker network, then reference them by container name — Docker's internal DNS resolves the name to the current IP. For example, an API reaches its database at db:5432, not localhost:5432. Never use localhost between containers: inside a container, localhost means that container's own loopback interface, not the host or any sibling container.

How does Docker Model Runner compare to using the OpenAI API?

Docker Model Runner runs models locally with an OpenAI-compatible API, so you switch by changing only the base URL — but you pay no per-request cost, need no API key, and no data leaves your machine. OpenAI's API offers larger frontier models and no local GPU requirement, but incurs cost, requires credentials, and sends data externally. Model Runner is ideal for privacy-sensitive, cost-sensitive, or offline workloads.

When should I use a named volume versus a bind mount?

Use a named volume for any data that must survive container deletion — databases, uploaded files — because Docker manages it independently of the container lifecycle. Use a bind mount for development iteration, since it maps a host directory directly into the container so file changes reflect instantly without rebuilding. Use tmpfs for sensitive temporary data like short-lived tokens that must never touch disk.

What results can I expect after applying this framework?

You get a fully containerised multi-service stack that starts with a single command, persists data correctly across restarts, networks cleanly by container name, and runs local AI without external credentials. Production images built with Docker Hardened Images report 95%+ fewer CVEs. Builds stay fast thanks to correct layer-cache ordering, and dependent services no longer crash-loop because health checks gate startup order.

What is MCP and why does it matter in Docker?

MCP (Model Context Protocol), developed by Anthropic, is the standard protocol for AI agents to connect to external tools — described as 'USB-C for AI agents.' Docker's MCP catalog offers 270+ containerised MCP servers (GitHub, Stripe, MongoDB, and more), each running in its own isolated container with credentials injected as Docker secrets. A compromised server cannot reach the host filesystem or sibling containers.

How do I stop my API from crashing while the database starts up?

Add depends_on with condition: service_healthy to your dependent service in Compose, and define a health check on the dependency. Without this, the API tries to connect while Postgres is still initialising, fails, crashes, and enters a restart loop. The health check (for example pg_isready on Postgres) gives Compose a signal to wait on before starting the dependent service.

Why is my Docker build so slow every time I change code?

You are almost certainly copying source code before installing dependencies in your Dockerfile, which destroys layer caching. Structure the Dockerfile so slow, infrequent steps (dependency installation) come before fast, frequent steps (copying source). Copy only the dependency manifest first, run the install, then copy source. Reversing this order turns a 2-second build into a 3-minute build on every code change.

// GET THIS SKILL — FREE

Use this skill in your AI

Every skill on SkillForge is free. Drop your email and copy this skill straight into Claude, ChatGPT, or any LLM.

We'll email you when new skills drop. Unsubscribe anytime.