Docker Compose Production Readiness Scorecard
Score a Docker Compose stack against ten observable production controls, keep the evidence beside each answer, and export or embed the result.
Scoring principle: This scorecard gives 10 points for each evidenced Yes, 5 for Partial, and 0 for No across ten controls; a score is useful only when every answer points to observable configuration or a tested procedure.
Published and method-checked 2026-10-01 · Runs locally in your browser · No input data is submitted
1.Images are pinned to deliberate versions
Evidence to keep: Record tags or digests and the process that approves upgrades.
2.Critical services have meaningful health checks
Evidence to keep: Show healthcheck definitions and what dependency or behavior each one proves.
3.Restart behavior matches the failure model
Evidence to keep: Show restart policies and explain which failures require operator intervention.
4.Memory and CPU boundaries are documented
Evidence to keep: Show limits or reservations plus load-test or observed usage evidence.
5.Secrets are kept outside committed Compose files
Evidence to keep: Show the secret source and confirm rendered configs and repositories do not expose values.
6.Containers use least privilege
Evidence to keep: Show user, capabilities, read-only filesystem, and privileged-mode decisions.
7.Persistent data and backup coverage are mapped
Evidence to keep: List every stateful mount and the backup and restore procedure that covers it.
8.Published ports and network boundaries are intentional
Evidence to keep: List public ports, internal networks, firewall rules, and reverse-proxy boundaries.
9.Logs, metrics, and failure alerts reach an owner
Evidence to keep: Show where signals go and a recent test that reached the responsible person.
10.Upgrade and rollback have been rehearsed
Evidence to keep: Record the last test, database migration handling, and the exact rollback decision point.
Yes = 10 points · Partial = 5 · No = 0. Evidence quality is not scored automatically.
Use the result with a real self-hosted plan
Method: score evidence, not optimism
A Compose file is ready for production only when its controls and operating practices are visible, owned, and usable during failure. A successful docker compose up proves that declared services can start in one environment. It does not prove safe upgrades, useful monitoring, bounded resource use, recoverable data, or a workable response when something breaks.
This scorecard uses ten controls. For each one, choose:
- Yes — 10 points: the control exists and is supported by observable configuration or a tested procedure.
- Partial — 5 points: the intent exists, but coverage, clarity, ownership, or verification is incomplete.
- No — 0 points: the control is absent, unknown, or contradicted by the current setup.
The maximum is 100 points. Evidence quality is not a separate scored category, and a polished note cannot turn a missing control into a pass. Notes explain the selection and point the next reviewer toward proof.
Docker’s production guidance for Compose treats production as a deliberate configuration of the application, including ports, environment settings, restart behavior, and logging. Use the effective production configuration for review, including overrides and deployment-time values, rather than relying only on a development file.
What earns full credit
Full credit requires evidence another operator can observe, run, or verify. “Backups are handled” is not enough without a documented job, destination, retention rule, or restore procedure. “The service restarts” is not enough if nobody has confirmed what happens after a process or host failure.
Useful evidence includes:
- An image reference chosen for repeatability, with an update and rollback process.
- A Compose service definition showing health checks, restart behavior, resource limits, user, secrets, volumes, ports, and networks.
- A dated command result, deployment record, alert, backup report, or restore result.
- A runbook naming the owner, prerequisites, commands, decisions, and rollback path.
- A rehearsal record showing that an upgrade, restore, or rollback was performed.
Configuration and operational proof are different. A healthcheck entry shows that a check was declared; it does not prove that the check reflects real readiness. A backup command shows intent; it does not prove that data can be restored. A rollback script shows a path; a rehearsal shows whether the path works under pressure.
The scorecard rewards both because production reliability depends on the handoff between the file and the people operating it.
The ten controls
A production-ready Compose stack addresses all ten controls below, with evidence attached to each answer.
-
Deliberate image versions — Use image references chosen for repeatability and maintenance instead of accidental moving targets. Record how images are reviewed, updated, and returned to a known-good version. The goal is visible, reversible change, not freezing software permanently.
-
Meaningful health checks — Test a useful service signal with timing and retries suited to startup behavior. Compose can define or override an image health check; the health check reference documents its fields and test forms. A process-only check may miss a broken dependency or unusable application.
-
Restart behavior — Define what happens after an exited process and after a host restart. Record the restart policy and its limits. Automatic restart is not recovery: an endlessly failing service still needs an owner and a response path.
-
CPU and memory boundaries — Set resource boundaries that match workload behavior and host capacity. Docker notes that containers have no resource constraints by default and explains how memory and CPU limits affect runtime behavior in its resource constraint guidance. Record how values were chosen and what happens when a limit is reached.
-
Secrets outside committed Compose files — Keep passwords, API keys, certificates, and similar values out of plain text in Compose files and image build contexts. Compose secrets are granted to services explicitly and mounted under
/run/secrets/<secret_name>; see Docker’s secrets guidance. Also document supply, rotation, and removal from old deployments. -
Least privilege — Give each service only the privileges, capabilities, filesystem access, and host integration it needs. Review the configured user, privileged mode, capability changes, read-only paths, device access, and Docker socket mounts. The OWASP Docker Security Cheat Sheet provides a useful review checklist.
-
Persistent data and backup mapping — Identify every stateful path, its storage location, ownership model, backup method, and restore procedure. Separate disposable container state from data that must survive. A named volume is storage, not automatically a backup plan.
-
Intentional ports and networks — Expose only traffic that must reach the host or an external network, and define permitted internal communication. Review host bindings, firewall assumptions, network separation, and public database or administration ports. A reverse proxy such as Caddy may form part of this boundary; see the Caddy with Docker Compose guide.
-
Logs, metrics, and alerts reaching an owner — Send runtime signals somewhere useful and name the person or team responsible for response. Confirm log retention, service output, core health signals, escalation, and dependencies on separate monitoring systems. “We can look at logs” is weaker than an owned path and response procedure.
-
Rehearsed upgrade and rollback — Document how images are pulled, containers recreated, data changes handled, and the previous version restored. Rehearse the path with representative data or a safe staging copy, then record failures, changes, and rollback authority.
Read the rating as a decision aid
The rating summarizes documented controls; it does not certify the application, host, image supply chain, or operator readiness. A high score can still conceal one serious weakness, especially when procedures have not been tested. Treat the weakest important control as a reason to pause even when the total looks comfortable.
| Score | Rating | Next action |
|---|---|---|
| 85–100 | Ready with evidence | Confirm owners, dates, and change triggers; keep evidence current. |
| 65–84 | Close, gaps remain | Fix the lowest-scoring controls before calling the stack production-ready. |
| 40–64 | High-risk gaps | Reduce exposure, protect data, and close operational gaps before launch. |
| 0–39 | Not ready | Establish basic control, ownership, and recovery practices first. |
These bands are the disclosed OpenAlt rubric, not a guarantee or industry-wide threshold. The score is most useful alongside the notes behind it and after each material change.
A practical review workflow
The reliable way to use the scorecard is as a short review cycle, not a one-time form.
-
Review the effective deployment configuration. Include production overrides, environment substitution, profiles, and external dependencies. The configuration actually used for deployment is the evidence.
-
Mark conservatively. Choose Yes only when the control is visible and current. Choose Partial when coverage, ownership, or recent verification is incomplete.
-
Test meaningful failure paths. Confirm behavior when a process exits, a dependency is unavailable, a service reaches a resource boundary, or a deployment must be reversed. Record results rather than relying on memory.
-
Trace data from write to restore. Map volumes and bind mounts, identify backup output, and restore into an isolated environment. Confirm that the service is usable, not merely that files were copied.
-
Review access and exposure. Remove unnecessary host ports, challenge broad network access, and check for root-level integrations. If Portainer is part of the operating model, use the Portainer with Docker Compose guide and include its access path in the review.
-
Export and assign follow-up. Keep the dated JSON with the Compose revision, review date, owners, and open actions. Re-score after upgrades, host changes, new public endpoints, data migrations, or monitoring and backup changes.
For resource planning, the Docker memory calculator can support discussion, but final values should remain tied to observed workload behavior and host capacity.
Notes, JSON, and sharing
Evidence notes make the result auditable. Write short entries such as “compose.production.yaml, reviewed on 2026-10-01” or “restore rehearsal completed; owner: platform team.” Never paste passwords, private keys, tokens, or sensitive customer data into notes. The note should help the next reviewer find proof without becoming a new secret.
The widget keeps a note per control, calculates the score in the browser, and exports dated JSON. All processing is local in the browser, so the exported file becomes a portable review record. Store it where the team can control access and retention, and pair it with the Compose revision or change record that explains what was reviewed.
The generated HTML link is a normal dofollow link containing the score and rating. Present it as a snapshot of the disclosed OpenAlt rubric, not as a security certification or uptime promise. Do not imply that OpenAlt reviewed the infrastructure, and update or remove the link when its evidence is no longer current.
For broader self-hosting decisions, continue with the self-hosted guide.
Scorecard FAQ
Is a high score proof that a Compose stack is secure?
No. It shows coverage of ten visible controls, not the absence of vulnerabilities. Keep the evidence and add threat modeling, dependency review, and service-specific hardening.
Why does Partial receive five points?
Partial distinguishes an implemented but incomplete or untested control from both a verified control and a missing one. The evidence note should explain the remaining gap.
Should every container have the same limits and health checks?
No. Controls should match each workload. The score asks whether the decision is deliberate, documented, and testable, not whether every service uses identical values.
Can I embed the result in documentation?
Yes. The generated HTML is a normal dofollow link to this scorecard and includes the score and rating. Re-run the assessment when the stack changes.