189 lines
9.3 KiB
Markdown
189 lines
9.3 KiB
Markdown
# Playbook: Live Service Maintenance Review
|
|
|
|
Use this playbook on a recurring basis (e.g. monthly) to validate that an already-live, externally-accessible homemade service is still healthy, secure, backed up, and current — and to apply any updates it needs.
|
|
|
|
This is not diff-driven. Apps degrade without code changes (new CVEs, expired certs, stopped backups, config drift, full disks), so every section runs every time.
|
|
|
|
When invoked, read the project directory in the current working directory. Some checks need the Docker host or the public URL — if the AI can't reach them directly, it gives the exact command to run and asks the user to paste the output.
|
|
|
|
---
|
|
|
|
## How to Use
|
|
|
|
Tell the AI: _"Use the service-update playbook to review this service: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_
|
|
|
|
The AI will:
|
|
1. Read the project files and the most recent prior report in `reports/` (`maintenance-*.md`, or `golive-*.md` if this is the first maintenance review)
|
|
2. Work through each section below
|
|
3. For each item — report PASS, FAIL, or WARN with specific findings
|
|
4. At the end, give an overall health rating and save a report
|
|
|
|
Do not proceed to the next section until the current one is resolved or explicitly deferred.
|
|
|
|
---
|
|
|
|
## Section 1: Baseline & Drift
|
|
|
|
Goal: Confirm what's actually running matches what's in git.
|
|
|
|
- [ ] Working tree is clean and all commits are pushed to Gitea
|
|
- [ ] Running image tag/commit matches the latest git tag (or `main`) — no unknown version in production
|
|
- [ ] Live Docker Compose stack in Portainer/Dockhand matches the repo's `docker-compose.yml` (no edits made only in the UI)
|
|
- [ ] Live Nginx/NPM proxy host config matches what's documented in the repo
|
|
- [ ] Every variable in `.env.example` is set in the live `.env`, and the live `.env` has no stale/unused variables
|
|
- [ ] Running image tags/digests are recorded in the report (this is the rollback point if anything below needs changing)
|
|
|
|
**AI Action:** Run `git status`, `git log -1 --oneline`, `git describe --tags --abbrev=0`. Ask the user for `docker ps --format "{{.Names}}\t{{.Image}}\t{{.Status}}"` and `docker inspect --format "{{.Image}}" <container>` from the host if not reachable.
|
|
|
|
---
|
|
|
|
## Section 2: Runtime Health
|
|
|
|
Goal: The service is up and not quietly struggling.
|
|
|
|
- [ ] All containers are running and report `healthy` (not just `Up`)
|
|
- [ ] No container has a climbing restart count (restart loop)
|
|
- [ ] Health endpoint responds via the public URL (through NPM), not just localhost
|
|
- [ ] Logs from the past review period show no recurring errors/exceptions or 5xx spikes
|
|
- [ ] CPU/memory usage is within resource limits with reasonable headroom
|
|
- [ ] Host disk and Docker volumes have adequate free space (flag WARN under 20% free)
|
|
- [ ] Log rotation (`max-size`/`max-file`) is working — log files aren't growing unbounded
|
|
- [ ] Reclaimable Docker disk space (dangling images, build cache) isn't excessive
|
|
|
|
**AI Action:** Commands to run on the host:
|
|
- `docker ps -a` and `docker inspect --format "{{.RestartCount}}" <container>`
|
|
- `docker stats --no-stream`
|
|
- `docker logs --since 720h <container> 2>&1 | grep -iE "error|exception|critical|traceback"`
|
|
- `df -h` and `docker system df`
|
|
- `curl -s -o /dev/null -w "%{http_code}" https://<public-url>/health`
|
|
|
|
---
|
|
|
|
## Section 3: Functional Check
|
|
|
|
Goal: The app actually works for users, not just that the container is running.
|
|
|
|
- [ ] Primary user flows work end-to-end (smoke test the main routes/functions)
|
|
- [ ] Login works; session expiry behaves as expected
|
|
- [ ] Scheduled jobs/background tasks have run recently (check last-run timestamps or logs)
|
|
- [ ] External integrations (email, third-party APIs, webhooks) still respond
|
|
- [ ] Ntfy still delivers — send a test notification and confirm it arrives
|
|
- [ ] Database connectivity and pool are healthy (no connection errors/timeouts in logs)
|
|
|
|
**AI Action:** Build a short smoke-test list from the app's routers/routes and ask the user to confirm each one.
|
|
|
|
---
|
|
|
|
## Section 4: Backups & Recovery
|
|
|
|
Goal: If the host died today, the service could be restored.
|
|
|
|
- [ ] Most recent database backup is within the expected schedule (not days/weeks stale)
|
|
- [ ] Backup files are non-zero size and not truncated
|
|
- [ ] Backups are stored off the Docker host (separate disk or remote target)
|
|
- [ ] Upload/file volumes are backed up, not just the database
|
|
- [ ] Live `.env` and any UI-only config are backed up somewhere retrievable (they aren't in git)
|
|
- [ ] Retention is working — old backups are pruned, not filling the disk
|
|
- [ ] A test restore has been performed within the last 6 months
|
|
|
|
**AI Action:** If no backup is found or the latest backup is stale, flag as CRITICAL.
|
|
|
|
---
|
|
|
|
## Section 5: Security Posture
|
|
|
|
Goal: Nothing about the service's exposure has weakened since go-live.
|
|
|
|
### 5a. TLS & Headers
|
|
- [ ] TLS certificate is valid with more than 14 days to expiry, and NPM auto-renewal is enabled
|
|
- [ ] HTTP still redirects to HTTPS
|
|
- [ ] Security headers from go-live are still present (`X-Frame-Options`, `X-Content-Type-Options`, `Referrer-Policy`, `Content-Security-Policy`, `Strict-Transport-Security`)
|
|
- [ ] Server version header still suppressed
|
|
|
|
### 5b. Exposure
|
|
- [ ] Only 80/443 are exposed publicly — no app or database ports published to the host/internet
|
|
- [ ] NPM access lists / basic auth still in place for any internal-only routes
|
|
- [ ] No new containers or ports added to the stack that weren't reviewed
|
|
|
|
### 5c. Accounts & Secrets
|
|
- [ ] Admin and user accounts reviewed — no stale, unknown, or unexpected admin accounts
|
|
- [ ] Secrets (DB passwords, API keys, secret keys) rotated per policy, or age noted if never rotated
|
|
- [ ] No unusual failed-login spikes, rate-limit triggers, or unknown admin logins in logs/Ntfy history
|
|
|
|
### 5d. Container Hardening
|
|
- [ ] Containers still run as non-root, no `privileged: true`, no unintended volume mounts (e.g. `docker.sock`)
|
|
- [ ] Resource limits (`mem_limit`, `cpus`) still set
|
|
|
|
**AI Action:** Run:
|
|
- `curl -sI https://<public-url>` — verify headers
|
|
- `echo | openssl s_client -connect <host>:443 -servername <host> 2>/dev/null | openssl x509 -noout -enddate` — cert expiry
|
|
- `docker ps --format "{{.Names}}\t{{.Ports}}"` — published ports
|
|
|
|
---
|
|
|
|
## Section 6: Vulnerabilities & Currency
|
|
|
|
Goal: Find what needs updating — new CVEs appear in code that hasn't changed.
|
|
|
|
- [ ] Running images scanned for HIGH/CRITICAL CVEs (includes base image OS packages)
|
|
- [ ] Application dependencies scanned for known CVEs
|
|
- [ ] Outdated dependencies identified — note which are security fixes vs. feature bumps
|
|
- [ ] Base images have newer patch versions available
|
|
- [ ] Runtimes/services are not end-of-life or near EOL (Python, MySQL, Node, Nginx, etc. — check https://endoflife.date)
|
|
- [ ] No dependencies that have become abandoned/unmaintained
|
|
- [ ] Dependencies and base images are still pinned (no `*`, no `:latest`)
|
|
|
|
**AI Action:** Run whichever applies:
|
|
- `trivy image --severity HIGH,CRITICAL <image:tag>` for each running image
|
|
- `trivy fs . --severity HIGH,CRITICAL`
|
|
- `pip-audit` and `pip list --outdated` (Python)
|
|
- `npm audit` and `npm outdated` (Node)
|
|
|
|
For a full image-level audit (SBOM, secrets, misconfig, grype), use the `docker-security-audit` playbook. Report new findings only if they aren't already accepted/deferred in the prior maintenance report.
|
|
|
|
---
|
|
|
|
## Section 7: Applying Updates
|
|
|
|
Goal: Apply needed updates safely. Skip this section if Section 6 found nothing to update.
|
|
|
|
- [ ] Fresh backup taken and confirmed (Section 4) immediately before updating
|
|
- [ ] Current image tags/digests recorded as the rollback point (Section 1)
|
|
- [ ] Release notes reviewed for any major-version bump — breaking changes identified
|
|
- [ ] Versions bumped to specific pinned versions (not `:latest` or unpinned)
|
|
- [ ] Any DB schema/data change has a tested rollback path and was tested against a copy of real data
|
|
- [ ] New image builds and starts locally, and passes its health check, before deploying
|
|
- [ ] Deployed under a specific version tag; downtime (if any) is expected/scheduled
|
|
- [ ] Sections 2 and 3 re-run after deploy — all PASS
|
|
- [ ] New version tagged in Gitea and `Release-Notes/v{major}.{minor}.md` updated
|
|
- [ ] Old images kept until the update is confirmed stable, then cleaned up
|
|
|
|
**AI Action:** If any step fails after deploy, stop and walk the user through rollback to the recorded image tags (and backup restore if a migration ran).
|
|
|
|
---
|
|
|
|
## Section 8: Documentation
|
|
|
|
- [ ] README still accurate for setup, usage, and environment variables
|
|
- [ ] Rollback / "what to do if this breaks" runbook still accurate for the current setup
|
|
- [ ] Any deferred items from this review are recorded in the report with a reason
|
|
|
|
---
|
|
|
|
## Section 9: Maintenance Summary
|
|
|
|
After all sections are complete:
|
|
|
|
- List all unresolved findings grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW**
|
|
- **CRITICAL or HIGH unresolved** = service needs action now (e.g. no working backup, expired/expiring cert, exploitable CVE, exposed DB port)
|
|
- **MEDIUM/LOW unresolved** = user decides whether to defer with documented acceptance
|
|
- Provide a final summary:
|
|
- Total checks: X
|
|
- Passed: X
|
|
- Failed (critical): X
|
|
- Failed (non-critical): X
|
|
- Deferred: X
|
|
- Updates applied: X
|
|
- **Overall: HEALTHY / HEALTHY WITH ACTIONS / NEEDS ATTENTION**
|
|
- Save the summary to `reports/maintenance-<YYYY-MM-DD>.md` in the project directory so the next review can see what was deferred and what was running
|