Files
AI/playbooks/service-update.md
T

9.3 KiB

Playbook: Live Service Maintenance Review

Use this playbook on a recurring basis (e.g. monthly) to validate that an already-live, externally-accessible homemade service is still healthy, secure, backed up, and current — and to apply any updates it needs.

This is not diff-driven. Apps degrade without code changes (new CVEs, expired certs, stopped backups, config drift, full disks), so every section runs every time.

When invoked, read the project directory in the current working directory. Some checks need the Docker host or the public URL — if the AI can't reach them directly, it gives the exact command to run and asks the user to paste the output.


How to Use

Tell the AI: "Use the service-update playbook to review this service: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"

The AI will:

  1. Read the project files and the most recent prior maintenance report (if one exists)
  2. Work through each section below
  3. For each item — report PASS, FAIL, or WARN with specific findings
  4. At the end, give an overall health rating and save a report

Do not proceed to the next section until the current one is resolved or explicitly deferred.


Section 1: Baseline & Drift

Goal: Confirm what's actually running matches what's in git.

  • Working tree is clean and all commits are pushed to Gitea
  • Running image tag/commit matches the latest git tag (or main) — no unknown version in production
  • Live Docker Compose stack in Portainer/Dockhand matches the repo's docker-compose.yml (no edits made only in the UI)
  • Live Nginx/NPM proxy host config matches what's documented in the repo
  • Every variable in .env.example is set in the live .env, and the live .env has no stale/unused variables
  • Running image tags/digests are recorded in the report (this is the rollback point if anything below needs changing)

AI Action: Run git status, git log -1 --oneline, git describe --tags --abbrev=0. Ask the user for docker ps --format "{{.Names}}\t{{.Image}}\t{{.Status}}" and docker inspect --format "{{.Image}}" <container> from the host if not reachable.


Section 2: Runtime Health

Goal: The service is up and not quietly struggling.

  • All containers are running and report healthy (not just Up)
  • No container has a climbing restart count (restart loop)
  • Health endpoint responds via the public URL (through NPM), not just localhost
  • Logs from the past review period show no recurring errors/exceptions or 5xx spikes
  • CPU/memory usage is within resource limits with reasonable headroom
  • Host disk and Docker volumes have adequate free space (flag WARN under 20% free)
  • Log rotation (max-size/max-file) is working — log files aren't growing unbounded
  • Reclaimable Docker disk space (dangling images, build cache) isn't excessive

AI Action: Commands to run on the host:

  • docker ps -a and docker inspect --format "{{.RestartCount}}" <container>
  • docker stats --no-stream
  • docker logs --since 720h <container> 2>&1 | grep -iE "error|exception|critical|traceback"
  • df -h and docker system df
  • curl -s -o /dev/null -w "%{http_code}" https://<public-url>/health

Section 3: Functional Check

Goal: The app actually works for users, not just that the container is running.

  • Primary user flows work end-to-end (smoke test the main routes/functions)
  • Login works; session expiry behaves as expected
  • Scheduled jobs/background tasks have run recently (check last-run timestamps or logs)
  • External integrations (email, third-party APIs, webhooks) still respond
  • Ntfy still delivers — send a test notification and confirm it arrives
  • Database connectivity and pool are healthy (no connection errors/timeouts in logs)

AI Action: Build a short smoke-test list from the app's routers/routes and ask the user to confirm each one.


Section 4: Backups & Recovery

Goal: If the host died today, the service could be restored.

  • Most recent database backup is within the expected schedule (not days/weeks stale)
  • Backup files are non-zero size and not truncated
  • Backups are stored off the Docker host (separate disk or remote target)
  • Upload/file volumes are backed up, not just the database
  • Live .env and any UI-only config are backed up somewhere retrievable (they aren't in git)
  • Retention is working — old backups are pruned, not filling the disk
  • A test restore has been performed within the last 6 months

AI Action: If no backup is found or the latest backup is stale, flag as CRITICAL.


Section 5: Security Posture

Goal: Nothing about the service's exposure has weakened since go-live.

5a. TLS & Headers

  • TLS certificate is valid with more than 14 days to expiry, and NPM auto-renewal is enabled
  • HTTP still redirects to HTTPS
  • Security headers from go-live are still present (X-Frame-Options, X-Content-Type-Options, Referrer-Policy, Content-Security-Policy, Strict-Transport-Security)
  • Server version header still suppressed

5b. Exposure

  • Only 80/443 are exposed publicly — no app or database ports published to the host/internet
  • NPM access lists / basic auth still in place for any internal-only routes
  • No new containers or ports added to the stack that weren't reviewed

5c. Accounts & Secrets

  • Admin and user accounts reviewed — no stale, unknown, or unexpected admin accounts
  • Secrets (DB passwords, API keys, secret keys) rotated per policy, or age noted if never rotated
  • No unusual failed-login spikes, rate-limit triggers, or unknown admin logins in logs/Ntfy history

5d. Container Hardening

  • Containers still run as non-root, no privileged: true, no unintended volume mounts (e.g. docker.sock)
  • Resource limits (mem_limit, cpus) still set

AI Action: Run:

  • curl -sI https://<public-url> — verify headers
  • echo | openssl s_client -connect <host>:443 -servername <host> 2>/dev/null | openssl x509 -noout -enddate — cert expiry
  • docker ps --format "{{.Names}}\t{{.Ports}}" — published ports

Section 6: Vulnerabilities & Currency

Goal: Find what needs updating — new CVEs appear in code that hasn't changed.

  • Running images scanned for HIGH/CRITICAL CVEs (includes base image OS packages)
  • Application dependencies scanned for known CVEs
  • Outdated dependencies identified — note which are security fixes vs. feature bumps
  • Base images have newer patch versions available
  • Runtimes/services are not end-of-life or near EOL (Python, MySQL, Node, Nginx, etc. — check https://endoflife.date)
  • No dependencies that have become abandoned/unmaintained
  • Dependencies and base images are still pinned (no *, no :latest)

AI Action: Run whichever applies:

  • trivy image --severity HIGH,CRITICAL <image:tag> for each running image
  • trivy fs . --severity HIGH,CRITICAL
  • pip-audit and pip list --outdated (Python)
  • npm audit and npm outdated (Node)

For a full image-level audit (SBOM, secrets, misconfig, grype), use the docker-security-audit playbook. Report new findings only if they aren't already accepted/deferred in the prior maintenance report.


Section 7: Applying Updates

Goal: Apply needed updates safely. Skip this section if Section 6 found nothing to update.

  • Fresh backup taken and confirmed (Section 4) immediately before updating
  • Current image tags/digests recorded as the rollback point (Section 1)
  • Release notes reviewed for any major-version bump — breaking changes identified
  • Versions bumped to specific pinned versions (not :latest or unpinned)
  • Any DB schema/data change has a tested rollback path and was tested against a copy of real data
  • New image builds and starts locally, and passes its health check, before deploying
  • Deployed under a specific version tag; downtime (if any) is expected/scheduled
  • Sections 2 and 3 re-run after deploy — all PASS
  • New version tagged in Gitea and Release-Notes/v{major}.{minor}.md updated
  • Old images kept until the update is confirmed stable, then cleaned up

AI Action: If any step fails after deploy, stop and walk the user through rollback to the recorded image tags (and backup restore if a migration ran).


Section 8: Documentation

  • README still accurate for setup, usage, and environment variables
  • Rollback / "what to do if this breaks" runbook still accurate for the current setup
  • Any deferred items from this review are recorded in the report with a reason

Section 9: Maintenance Summary

After all sections are complete:

  • List all unresolved findings grouped by severity: CRITICAL / HIGH / MEDIUM / LOW
  • CRITICAL or HIGH unresolved = service needs action now (e.g. no working backup, expired/expiring cert, exploitable CVE, exposed DB port)
  • MEDIUM/LOW unresolved = user decides whether to defer with documented acceptance
  • Provide a final summary:
    • Total checks: X
    • Passed: X
    • Failed (critical): X
    • Failed (non-critical): X
    • Deferred: X
    • Updates applied: X
    • Overall: HEALTHY / HEALTHY WITH ACTIONS / NEEDS ATTENTION
  • Save the summary to reports/maintenance-<YYYY-MM-DD>.md in the project directory so the next review can see what was deferred and what was running