9.3 KiB
Playbook: Live Service Maintenance Review
Use this playbook on a recurring basis (e.g. monthly) to validate that an already-live, externally-accessible homemade service is still healthy, secure, backed up, and current — and to apply any updates it needs.
This is not diff-driven. Apps degrade without code changes (new CVEs, expired certs, stopped backups, config drift, full disks), so every section runs every time.
When invoked, read the project directory in the current working directory. Some checks need the Docker host or the public URL — if the AI can't reach them directly, it gives the exact command to run and asks the user to paste the output.
How to Use
Tell the AI: "Use the service-update playbook to review this service: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"
The AI will:
- Read the project files and the most recent prior maintenance report (if one exists)
- Work through each section below
- For each item — report PASS, FAIL, or WARN with specific findings
- At the end, give an overall health rating and save a report
Do not proceed to the next section until the current one is resolved or explicitly deferred.
Section 1: Baseline & Drift
Goal: Confirm what's actually running matches what's in git.
- Working tree is clean and all commits are pushed to Gitea
- Running image tag/commit matches the latest git tag (or
main) — no unknown version in production - Live Docker Compose stack in Portainer/Dockhand matches the repo's
docker-compose.yml(no edits made only in the UI) - Live Nginx/NPM proxy host config matches what's documented in the repo
- Every variable in
.env.exampleis set in the live.env, and the live.envhas no stale/unused variables - Running image tags/digests are recorded in the report (this is the rollback point if anything below needs changing)
AI Action: Run git status, git log -1 --oneline, git describe --tags --abbrev=0. Ask the user for docker ps --format "{{.Names}}\t{{.Image}}\t{{.Status}}" and docker inspect --format "{{.Image}}" <container> from the host if not reachable.
Section 2: Runtime Health
Goal: The service is up and not quietly struggling.
- All containers are running and report
healthy(not justUp) - No container has a climbing restart count (restart loop)
- Health endpoint responds via the public URL (through NPM), not just localhost
- Logs from the past review period show no recurring errors/exceptions or 5xx spikes
- CPU/memory usage is within resource limits with reasonable headroom
- Host disk and Docker volumes have adequate free space (flag WARN under 20% free)
- Log rotation (
max-size/max-file) is working — log files aren't growing unbounded - Reclaimable Docker disk space (dangling images, build cache) isn't excessive
AI Action: Commands to run on the host:
docker ps -aanddocker inspect --format "{{.RestartCount}}" <container>docker stats --no-streamdocker logs --since 720h <container> 2>&1 | grep -iE "error|exception|critical|traceback"df -handdocker system dfcurl -s -o /dev/null -w "%{http_code}" https://<public-url>/health
Section 3: Functional Check
Goal: The app actually works for users, not just that the container is running.
- Primary user flows work end-to-end (smoke test the main routes/functions)
- Login works; session expiry behaves as expected
- Scheduled jobs/background tasks have run recently (check last-run timestamps or logs)
- External integrations (email, third-party APIs, webhooks) still respond
- Ntfy still delivers — send a test notification and confirm it arrives
- Database connectivity and pool are healthy (no connection errors/timeouts in logs)
AI Action: Build a short smoke-test list from the app's routers/routes and ask the user to confirm each one.
Section 4: Backups & Recovery
Goal: If the host died today, the service could be restored.
- Most recent database backup is within the expected schedule (not days/weeks stale)
- Backup files are non-zero size and not truncated
- Backups are stored off the Docker host (separate disk or remote target)
- Upload/file volumes are backed up, not just the database
- Live
.envand any UI-only config are backed up somewhere retrievable (they aren't in git) - Retention is working — old backups are pruned, not filling the disk
- A test restore has been performed within the last 6 months
AI Action: If no backup is found or the latest backup is stale, flag as CRITICAL.
Section 5: Security Posture
Goal: Nothing about the service's exposure has weakened since go-live.
5a. TLS & Headers
- TLS certificate is valid with more than 14 days to expiry, and NPM auto-renewal is enabled
- HTTP still redirects to HTTPS
- Security headers from go-live are still present (
X-Frame-Options,X-Content-Type-Options,Referrer-Policy,Content-Security-Policy,Strict-Transport-Security) - Server version header still suppressed
5b. Exposure
- Only 80/443 are exposed publicly — no app or database ports published to the host/internet
- NPM access lists / basic auth still in place for any internal-only routes
- No new containers or ports added to the stack that weren't reviewed
5c. Accounts & Secrets
- Admin and user accounts reviewed — no stale, unknown, or unexpected admin accounts
- Secrets (DB passwords, API keys, secret keys) rotated per policy, or age noted if never rotated
- No unusual failed-login spikes, rate-limit triggers, or unknown admin logins in logs/Ntfy history
5d. Container Hardening
- Containers still run as non-root, no
privileged: true, no unintended volume mounts (e.g.docker.sock) - Resource limits (
mem_limit,cpus) still set
AI Action: Run:
curl -sI https://<public-url>— verify headersecho | openssl s_client -connect <host>:443 -servername <host> 2>/dev/null | openssl x509 -noout -enddate— cert expirydocker ps --format "{{.Names}}\t{{.Ports}}"— published ports
Section 6: Vulnerabilities & Currency
Goal: Find what needs updating — new CVEs appear in code that hasn't changed.
- Running images scanned for HIGH/CRITICAL CVEs (includes base image OS packages)
- Application dependencies scanned for known CVEs
- Outdated dependencies identified — note which are security fixes vs. feature bumps
- Base images have newer patch versions available
- Runtimes/services are not end-of-life or near EOL (Python, MySQL, Node, Nginx, etc. — check https://endoflife.date)
- No dependencies that have become abandoned/unmaintained
- Dependencies and base images are still pinned (no
*, no:latest)
AI Action: Run whichever applies:
trivy image --severity HIGH,CRITICAL <image:tag>for each running imagetrivy fs . --severity HIGH,CRITICALpip-auditandpip list --outdated(Python)npm auditandnpm outdated(Node)
For a full image-level audit (SBOM, secrets, misconfig, grype), use the docker-security-audit playbook. Report new findings only if they aren't already accepted/deferred in the prior maintenance report.
Section 7: Applying Updates
Goal: Apply needed updates safely. Skip this section if Section 6 found nothing to update.
- Fresh backup taken and confirmed (Section 4) immediately before updating
- Current image tags/digests recorded as the rollback point (Section 1)
- Release notes reviewed for any major-version bump — breaking changes identified
- Versions bumped to specific pinned versions (not
:latestor unpinned) - Any DB schema/data change has a tested rollback path and was tested against a copy of real data
- New image builds and starts locally, and passes its health check, before deploying
- Deployed under a specific version tag; downtime (if any) is expected/scheduled
- Sections 2 and 3 re-run after deploy — all PASS
- New version tagged in Gitea and
Release-Notes/v{major}.{minor}.mdupdated - Old images kept until the update is confirmed stable, then cleaned up
AI Action: If any step fails after deploy, stop and walk the user through rollback to the recorded image tags (and backup restore if a migration ran).
Section 8: Documentation
- README still accurate for setup, usage, and environment variables
- Rollback / "what to do if this breaks" runbook still accurate for the current setup
- Any deferred items from this review are recorded in the report with a reason
Section 9: Maintenance Summary
After all sections are complete:
- List all unresolved findings grouped by severity: CRITICAL / HIGH / MEDIUM / LOW
- CRITICAL or HIGH unresolved = service needs action now (e.g. no working backup, expired/expiring cert, exploitable CVE, exposed DB port)
- MEDIUM/LOW unresolved = user decides whether to defer with documented acceptance
- Provide a final summary:
- Total checks: X
- Passed: X
- Failed (critical): X
- Failed (non-critical): X
- Deferred: X
- Updates applied: X
- Overall: HEALTHY / HEALTHY WITH ACTIONS / NEEDS ATTENTION
- Save the summary to
reports/maintenance-<YYYY-MM-DD>.mdin the project directory so the next review can see what was deferred and what was running