Frequently Asked Questions About ByteMonk Docker-to-AI Infrastructure Framework
23 answers covering everything from basics to advanced usage.
// Basics
What is the difference between a container and a virtual machine?
A container is a single process running on the host's kernel, isolated by Cgroups (resource limits) and namespaces (private process tree, network, filesystem). A VM runs a full guest operating system with its own kernel on top of a hypervisor. That's why containers start in milliseconds, use far less memory, and ship as small images — but also why they share the host kernel. Model containers as processes, never as lightweight VMs.
What does the Docker daemon do?
The Docker daemon is the background process that does the actual work — building images, starting and stopping containers, and managing networks and volumes. The Docker client is what you type commands into; it sends instructions to the daemon. Running docker run hello-world successfully confirms the client, daemon, and registry connection are all functional.
What is layer caching and how do I use it well?
Layer caching is Docker storing the result of each Dockerfile instruction as a layer; if a layer hasn't changed, Docker skips rebuilding it. To use it well, order instructions from slow-changing to fast-changing: copy your dependency manifest and install dependencies first, then copy source code last. This way a code edit only re-runs the cheap final layers, keeping builds fast.
What is a .dockerignore file and why do I need one?
A .dockerignore file excludes files from the build context Docker sends to the daemon. Without it, Docker ships your entire project directory — including node_modules, .git history, and .env secrets — bloating builds and risking leaked credentials. Always create one and exclude node_modules, .git, and .env files at minimum.
// How To
How do I verify my Docker installation is working?
Run docker pull hello-world, then docker run hello-world, then docker --version. A successful hello-world run confirms the Docker client, the Docker daemon, and the registry connection are all functional together. On Linux, add your user to the Docker group so you don't need sudo on every command.
How do I write a Dockerfile that builds fast and stays small?
Structure it as FROM (slim or Alpine variant) → WORKDIR → COPY dependency manifest only → RUN install dependencies → COPY application source → EXPOSE → CMD. Add a .dockerignore, use npm ci not npm install, and apply multi-stage builds for production so compilers and test frameworks never reach the runtime image. This maximises cache hits and minimises attack surface.
How do I write a Docker Compose file for a multi-service app?
Follow one service = one container. Define each service's build context or image, environment variables loaded from .env, volume mounts, network membership, exposed ports, health checks, and depends_on with condition: service_healthy. Add .env to .gitignore immediately and never hardcode passwords in the YAML. Use docker compose (space, no hyphen) — the hyphenated form was removed in 2026.
How do I connect an AI service to an external tool like GitHub or Stripe?
Enable the relevant MCP server from the Docker MCP catalog with docker mcp server enable <server-name>. For multiple servers, add the MCP Gateway as a Compose service — it presents a single unified endpoint, logs every tool call, and enforces access controls. Configure your AI service with the Gateway URL as an environment variable. Credentials are injected as Docker secrets, never in application code.
How do I harden my Docker images for production?
Switch base images to Docker Hardened Images (DHI) using the dhi-ctl CLI plugin or Gordon AI assistant. Add a non-root user with adduser --disabled-password appuser and switch to it with USER appuser so the app never runs as root. Run docker scout on every image before shipping and integrate Scout into your CI/CD pipeline to catch vulnerabilities before production.
// Troubleshooting
My containers can't reach each other on localhost — what's wrong?
Inside a container, localhost refers only to that container's own loopback interface, not the host or any sibling container. Place all services that need to communicate on the same user-defined Docker network, then use container names as hostnames — Docker's internal DNS resolves them. An API reaches Postgres at db:5432, not localhost:5432.
Why does my data disappear when I restart or delete a container?
The default container writable layer is ephemeral — it's gone when the container is deleted. Any data that must survive needs a named volume, which Docker manages independently of the container lifecycle. Attach a named volume to services that write persistent data, like databases and uploaded files. Note that docker compose down -v deletes named volumes too.
Why does my first LLM call time out but later calls work fine?
The first call loads the model into memory, which takes significantly longer than subsequent calls, when the model stays resident. Set at least a 45-second timeout on the first call to accommodate model load. This is especially relevant with Ollama on the host referenced via host.docker.internal, where cold-start loading is the usual cause of first-call timeouts.
I accidentally ran docker compose down -v — what happened?
That command stopped and removed your containers AND deleted all named volumes, destroying any persistent data those volumes held. It's correct only for a development clean-slate reset. On production data it is catastrophic and irreversible without backups — this is the command that ends careers. Use plain docker compose down to stop containers while keeping named volumes intact.
Why does my API crash-loop when the stack starts?
Your dependent service is starting before its dependency is ready. Without depends_on: condition: service_healthy, the API tries to connect while Postgres is still initialising, fails, crashes, and restarts repeatedly. Define a health check on the dependency (for example pg_isready) and add the service_healthy condition to the dependent service so Compose waits for a healthy signal before starting it.
// Comparisons
How does Docker Compose compare to running docker run manually for each service?
Docker Compose is a single YAML blueprint that brings an entire multi-service stack up or down with one command, encoding networks, volumes, health checks, and startup order declaratively. Running docker run per service means manually managing networks, volumes, environment variables, and startup ordering by hand every time — error-prone and unrepeatable. Compose is the single source of truth for the stack.
How do Docker Hardened Images compare to standard base images?
Docker Hardened Images are minimal by design, ship with a full software bill of materials and cryptographic provenance, and receive critical CVE patches within 7 days — reported to have 95%+ fewer CVEs than standard base images. Standard images carry more packages and thus more vulnerabilities, since most container CVEs come from the base image, not application code. Use DHI in production.
How does Docker Model Runner compare to Ollama for local AI?
Docker Model Runner integrates directly into your Compose workflow via the provider syntax and exposes an OpenAI-compatible internal endpoint. Ollama runs on the host and is referenced from containers via host.docker.internal, using lower resources — often better for development. Model Runner keeps everything inside the Docker workflow; Ollama is a lighter-weight host process. Both keep data local with no external API key.
Why should I use npm ci instead of npm install in a Dockerfile?
npm ci installs exactly what's locked in package-lock.json and nothing else, giving reproducible builds. npm install can silently modify the lockfile, meaning your image may not match what you tested locally. In Dockerfiles and CI pipelines, reproducibility matters, so npm ci is the correct choice for deterministic, trustworthy builds.
// Advanced
When should I publish a port with the -p flag?
Use -p only for ports that genuinely need to be reachable from outside Docker, such as a frontend UI. Databases should almost never have a published port — leave them on the internal Docker network reachable by container name. Publishing only what needs external access keeps the attack surface small and the network layout clean.
How do multi-stage builds improve production images?
A multi-stage build uses a first stage containing all build tools — compilers, bundlers, test frameworks — to produce the compiled artefact, then copies only that final output into a second minimal runtime stage. The runtime stage has no compilers, no test frameworks, and no source files, producing a smaller image with a smaller attack surface and faster pulls.
What is the MCP Gateway and when do I need it?
The MCP Gateway is an additional Docker service for teams running multiple MCP servers. It presents a single unified endpoint to the AI agent, logs every tool call, enforces access controls, and blocks suspicious requests before they reach the actual tool. Add it when you're connecting more than one MCP server so you get centralised logging, auth, and governance instead of wiring each server individually.
How do isolated MCP containers protect my system?
Each MCP server runs in its own isolated container and cannot reach the host filesystem or sibling containers it isn't authorised to reach. Credentials are injected as Docker secrets and never appear in application code. This means a compromised MCP server is contained — it can't pivot to your host or other services — which is the core security advantage of running MCP servers as containers.
What is host.docker.internal used for?
host.docker.internal is a special DNS name resolvable from inside a container that points to the host machine's IP. Use it when a containerised service needs to reach a process running directly on the host — for example, an app container calling Ollama running on the host. Combine it with a 45-second first-call timeout to handle model load.