updated and corrected to better fit the goal
This commit is contained in:
1 parent
c84940ad81
commit
2f51362c96
2 files changed
+129
-89
No files matched your search
+127
-88
@@ -1,144 +1,181 @@
|
||||
# Playbook: Service Update Review
|
||||
# Playbook: Live Service Maintenance Review
|
||||
|
||||
Use this playbook before deploying an update (code change, dependency bump, config change) to an already-running, externally-accessible homemade service.
|
||||
Use this playbook on a recurring basis (e.g. monthly) to validate that an already-live, externally-accessible homemade service is still healthy, secure, backed up, and current — and to apply any updates it needs.
|
||||
|
||||
When invoked, read the project directory in the current working directory. Diff the current working tree/branch against the last git tag if one exists, or against `origin/main` if not. If neither gives a sensible comparison point, ask the user what to diff against before continuing.
|
||||
This is not diff-driven. Apps degrade without code changes (new CVEs, expired certs, stopped backups, config drift, full disks), so every section runs every time.
|
||||
|
||||
When invoked, read the project directory in the current working directory. Some checks need the Docker host or the public URL — if the AI can't reach them directly, it gives the exact command to run and asks the user to paste the output.
|
||||
|
||||
---
|
||||
|
||||
## How to Use
|
||||
|
||||
Tell the AI: _"Use the service-update playbook to review this update: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_
|
||||
Tell the AI: _"Use the service-update playbook to review this service: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_
|
||||
|
||||
The AI will:
|
||||
1. Determine what changed since the last deployed version (git diff/log)
|
||||
1. Read the project files and the most recent prior maintenance report (if one exists)
|
||||
2. Work through each section below
|
||||
3. For each item — report PASS, FAIL, or WARN with specific findings
|
||||
4. At the end, give a go/no-go recommendation for the update
|
||||
4. At the end, give an overall health rating and save a report
|
||||
|
||||
Do not proceed to the next section until the current one is resolved or explicitly deferred.
|
||||
|
||||
---
|
||||
|
||||
## Section 1: Change Summary
|
||||
## Section 1: Baseline & Drift
|
||||
|
||||
Goal: Know exactly what's changing before evaluating risk.
|
||||
Goal: Confirm what's actually running matches what's in git.
|
||||
|
||||
- [ ] List every file changed, added, or removed since the comparison point
|
||||
- [ ] Summarize the intent of the change (bug fix, feature, dependency bump, refactor)
|
||||
- [ ] Are there any TODO/FIXME/HACK comments left in the changed files?
|
||||
- [ ] Does the change include a database schema change (migration file present)?
|
||||
- [ ] Does the change add any new environment variables, config keys, or ports?
|
||||
- [ ] Does the change add any new external dependency (package, API, service)?
|
||||
- [ ] Working tree is clean and all commits are pushed to Gitea
|
||||
- [ ] Running image tag/commit matches the latest git tag (or `main`) — no unknown version in production
|
||||
- [ ] Live Docker Compose stack in Portainer/Dockhand matches the repo's `docker-compose.yml` (no edits made only in the UI)
|
||||
- [ ] Live Nginx/NPM proxy host config matches what's documented in the repo
|
||||
- [ ] Every variable in `.env.example` is set in the live `.env`, and the live `.env` has no stale/unused variables
|
||||
- [ ] Running image tags/digests are recorded in the report (this is the rollback point if anything below needs changing)
|
||||
|
||||
**AI Action:** Run `git log <comparison>..HEAD --oneline` and `git diff <comparison>..HEAD --stat`. Use this to scope the rest of the review — sections below only need to focus on what actually changed.
|
||||
**AI Action:** Run `git status`, `git log -1 --oneline`, `git describe --tags --abbrev=0`. Ask the user for `docker ps --format "{{.Names}}\t{{.Image}}\t{{.Status}}"` and `docker inspect --format "{{.Image}}" <container>` from the host if not reachable.
|
||||
|
||||
---
|
||||
|
||||
## Section 2: Backup Before Update
|
||||
## Section 2: Runtime Health
|
||||
|
||||
Goal: Make sure the update is reversible before it starts.
|
||||
Goal: The service is up and not quietly struggling.
|
||||
|
||||
- [ ] Current database/data volume has a fresh backup taken before applying the update
|
||||
- [ ] Backup is confirmed to exist and is non-zero size (don't just assume the job ran)
|
||||
- [ ] The currently-running image tag / commit hash is noted somewhere retrievable, so you know exactly what to roll back to
|
||||
- [ ] The currently-deployed Docker Compose / Nginx config is saved or committed, not just live in NPM/Portainer
|
||||
- [ ] All containers are running and report `healthy` (not just `Up`)
|
||||
- [ ] No container has a climbing restart count (restart loop)
|
||||
- [ ] Health endpoint responds via the public URL (through NPM), not just localhost
|
||||
- [ ] Logs from the past review period show no recurring errors/exceptions or 5xx spikes
|
||||
- [ ] CPU/memory usage is within resource limits with reasonable headroom
|
||||
- [ ] Host disk and Docker volumes have adequate free space (flag WARN under 20% free)
|
||||
- [ ] Log rotation (`max-size`/`max-file`) is working — log files aren't growing unbounded
|
||||
- [ ] Reclaimable Docker disk space (dangling images, build cache) isn't excessive
|
||||
|
||||
**AI Action:** If no backup step is found for this service, flag as FAIL and stop — do not proceed to deployment sections until a backup exists.
|
||||
**AI Action:** Commands to run on the host:
|
||||
- `docker ps -a` and `docker inspect --format "{{.RestartCount}}" <container>`
|
||||
- `docker stats --no-stream`
|
||||
- `docker logs --since 720h <container> 2>&1 | grep -iE "error|exception|critical|traceback"`
|
||||
- `df -h` and `docker system df`
|
||||
- `curl -s -o /dev/null -w "%{http_code}" https://<public-url>/health`
|
||||
|
||||
---
|
||||
|
||||
## Section 3: Database Migrations
|
||||
## Section 3: Functional Check
|
||||
|
||||
Goal: Don't lose or corrupt data on update. Skip this section if the change has no schema/data changes.
|
||||
Goal: The app actually works for users, not just that the container is running.
|
||||
|
||||
- [ ] Migration scripts are idempotent or safe to re-run if interrupted
|
||||
- [ ] Migration has a documented rollback (down migration, or manual revert steps if the framework doesn't support down migrations)
|
||||
- [ ] Migration has been tested against a copy of real data, not just an empty dev database
|
||||
- [ ] No destructive change (dropped column/table, changed column type with data loss) without a confirmed, tested backup
|
||||
- [ ] Migration runs automatically on deploy, or the manual step to run it is documented
|
||||
- [ ] Primary user flows work end-to-end (smoke test the main routes/functions)
|
||||
- [ ] Login works; session expiry behaves as expected
|
||||
- [ ] Scheduled jobs/background tasks have run recently (check last-run timestamps or logs)
|
||||
- [ ] External integrations (email, third-party APIs, webhooks) still respond
|
||||
- [ ] Ntfy still delivers — send a test notification and confirm it arrives
|
||||
- [ ] Database connectivity and pool are healthy (no connection errors/timeouts in logs)
|
||||
|
||||
**AI Action:** Build a short smoke-test list from the app's routers/routes and ask the user to confirm each one.
|
||||
|
||||
---
|
||||
|
||||
## Section 4: Dependency & Supply Chain Delta
|
||||
## Section 4: Backups & Recovery
|
||||
|
||||
Goal: A version bump shouldn't introduce a new vulnerability or a moving target.
|
||||
Goal: If the host died today, the service could be restored.
|
||||
|
||||
- [ ] New or updated dependencies reviewed for known CVEs
|
||||
- [ ] Dependency versions are still pinned to specific versions (not moved to `*` or unpinned)
|
||||
- [ ] No new dependency pulled from an unverified/unofficial source
|
||||
- [ ] If the Dockerfile base image changed, it's still from an official/verified source and still pinned (not moved to `:latest`)
|
||||
- [ ] Most recent database backup is within the expected schedule (not days/weeks stale)
|
||||
- [ ] Backup files are non-zero size and not truncated
|
||||
- [ ] Backups are stored off the Docker host (separate disk or remote target)
|
||||
- [ ] Upload/file volumes are backed up, not just the database
|
||||
- [ ] Live `.env` and any UI-only config are backed up somewhere retrievable (they aren't in git)
|
||||
- [ ] Retention is working — old backups are pruned, not filling the disk
|
||||
- [ ] A test restore has been performed within the last 6 months
|
||||
|
||||
**AI Action:** If no backup is found or the latest backup is stale, flag as CRITICAL.
|
||||
|
||||
---
|
||||
|
||||
## Section 5: Security Posture
|
||||
|
||||
Goal: Nothing about the service's exposure has weakened since go-live.
|
||||
|
||||
### 5a. TLS & Headers
|
||||
- [ ] TLS certificate is valid with more than 14 days to expiry, and NPM auto-renewal is enabled
|
||||
- [ ] HTTP still redirects to HTTPS
|
||||
- [ ] Security headers from go-live are still present (`X-Frame-Options`, `X-Content-Type-Options`, `Referrer-Policy`, `Content-Security-Policy`, `Strict-Transport-Security`)
|
||||
- [ ] Server version header still suppressed
|
||||
|
||||
### 5b. Exposure
|
||||
- [ ] Only 80/443 are exposed publicly — no app or database ports published to the host/internet
|
||||
- [ ] NPM access lists / basic auth still in place for any internal-only routes
|
||||
- [ ] No new containers or ports added to the stack that weren't reviewed
|
||||
|
||||
### 5c. Accounts & Secrets
|
||||
- [ ] Admin and user accounts reviewed — no stale, unknown, or unexpected admin accounts
|
||||
- [ ] Secrets (DB passwords, API keys, secret keys) rotated per policy, or age noted if never rotated
|
||||
- [ ] No unusual failed-login spikes, rate-limit triggers, or unknown admin logins in logs/Ntfy history
|
||||
|
||||
### 5d. Container Hardening
|
||||
- [ ] Containers still run as non-root, no `privileged: true`, no unintended volume mounts (e.g. `docker.sock`)
|
||||
- [ ] Resource limits (`mem_limit`, `cpus`) still set
|
||||
|
||||
**AI Action:** Run:
|
||||
- `curl -sI https://<public-url>` — verify headers
|
||||
- `echo | openssl s_client -connect <host>:443 -servername <host> 2>/dev/null | openssl x509 -noout -enddate` — cert expiry
|
||||
- `docker ps --format "{{.Names}}\t{{.Ports}}"` — published ports
|
||||
|
||||
---
|
||||
|
||||
## Section 6: Vulnerabilities & Currency
|
||||
|
||||
Goal: Find what needs updating — new CVEs appear in code that hasn't changed.
|
||||
|
||||
- [ ] Running images scanned for HIGH/CRITICAL CVEs (includes base image OS packages)
|
||||
- [ ] Application dependencies scanned for known CVEs
|
||||
- [ ] Outdated dependencies identified — note which are security fixes vs. feature bumps
|
||||
- [ ] Base images have newer patch versions available
|
||||
- [ ] Runtimes/services are not end-of-life or near EOL (Python, MySQL, Node, Nginx, etc. — check https://endoflife.date)
|
||||
- [ ] No dependencies that have become abandoned/unmaintained
|
||||
- [ ] Dependencies and base images are still pinned (no `*`, no `:latest`)
|
||||
|
||||
**AI Action:** Run whichever applies:
|
||||
- `trivy image --severity HIGH,CRITICAL <image:tag>` for each running image
|
||||
- `trivy fs . --severity HIGH,CRITICAL`
|
||||
- `pip-audit` (Python)
|
||||
- `npm audit` (Node)
|
||||
- `pip-audit` and `pip list --outdated` (Python)
|
||||
- `npm audit` and `npm outdated` (Node)
|
||||
|
||||
Report new findings only if they weren't already accepted/deferred in a prior review.
|
||||
For a full image-level audit (SBOM, secrets, misconfig, grype), use the `docker-security-audit` playbook. Report new findings only if they aren't already accepted/deferred in the prior maintenance report.
|
||||
|
||||
---
|
||||
|
||||
## Section 5: Security Delta
|
||||
## Section 7: Applying Updates
|
||||
|
||||
Goal: Re-check only what changed — this is not a full re-audit, but changed code gets the same scrutiny as go-live.
|
||||
Goal: Apply needed updates safely. Skip this section if Section 6 found nothing to update.
|
||||
|
||||
- [ ] New code introduces no hardcoded secrets, passwords, or tokens
|
||||
- [ ] New or changed endpoints/routes require authentication consistent with the rest of the app
|
||||
- [ ] New user input is validated server-side and uses parameterized queries / ORM (no string-built SQL)
|
||||
- [ ] No debug/dev tooling left enabled (e.g. `debug=True`, exposed `/docs`, verbose stack traces)
|
||||
- [ ] Any new file upload, redirect, or deserialization logic is checked for path traversal, open redirect, and unsafe object instantiation
|
||||
- [ ] Any new HTML output is escaped (no new XSS surface)
|
||||
- [ ] Fresh backup taken and confirmed (Section 4) immediately before updating
|
||||
- [ ] Current image tags/digests recorded as the rollback point (Section 1)
|
||||
- [ ] Release notes reviewed for any major-version bump — breaking changes identified
|
||||
- [ ] Versions bumped to specific pinned versions (not `:latest` or unpinned)
|
||||
- [ ] Any DB schema/data change has a tested rollback path and was tested against a copy of real data
|
||||
- [ ] New image builds and starts locally, and passes its health check, before deploying
|
||||
- [ ] Deployed under a specific version tag; downtime (if any) is expected/scheduled
|
||||
- [ ] Sections 2 and 3 re-run after deploy — all PASS
|
||||
- [ ] New version tagged in Gitea and `Release-Notes/v{major}.{minor}.md` updated
|
||||
- [ ] Old images kept until the update is confirmed stable, then cleaned up
|
||||
|
||||
**AI Action:** If any step fails after deploy, stop and walk the user through rollback to the recorded image tags (and backup restore if a migration ran).
|
||||
|
||||
---
|
||||
|
||||
## Section 6: Config & Environment Changes
|
||||
## Section 8: Documentation
|
||||
|
||||
Goal: Config drift is how services quietly become insecure between go-live and now.
|
||||
|
||||
- [ ] Any new required environment variable is added to `.env.example` with a placeholder value
|
||||
- [ ] No new secrets added directly to Docker Compose or Nginx config files
|
||||
- [ ] Nginx config changes (if any) don't remove or weaken existing security headers
|
||||
- [ ] Container `user:`, `privileged:`, volume mounts, and resource limits are unchanged or still compliant with go-live standards
|
||||
- [ ] README still accurate for setup, usage, and environment variables
|
||||
- [ ] Rollback / "what to do if this breaks" runbook still accurate for the current setup
|
||||
- [ ] Any deferred items from this review are recorded in the report with a reason
|
||||
|
||||
---
|
||||
|
||||
## Section 7: Deployment Process
|
||||
|
||||
Goal: Get the update live without unplanned downtime or a stuck rollback.
|
||||
|
||||
- [ ] Update is deployed under a specific version tag, not floating `:latest`
|
||||
- [ ] Deployment approach causes minimal/no downtime, or downtime is scheduled and expected
|
||||
- [ ] Rollback procedure (previous image tag, compose file, backup restore) is known before starting, not figured out after something breaks
|
||||
- [ ] Health check endpoint is confirmed reachable immediately after deploy
|
||||
|
||||
---
|
||||
|
||||
## Section 8: Post-Update Verification
|
||||
|
||||
Goal: Confirm the update actually works in production, not just that it deployed.
|
||||
|
||||
- [ ] Primary routes/functions respond correctly after deploy (smoke test the main flows)
|
||||
- [ ] Logs show no new errors/exceptions in the minutes following deploy
|
||||
- [ ] Ntfy/notification integration (if implemented) still fires correctly
|
||||
- [ ] No unexpected spike in CPU/memory/disk after deploy
|
||||
- [ ] Old container/image cleaned up once the new one is confirmed stable (if not needed for rollback anymore)
|
||||
|
||||
---
|
||||
|
||||
## Section 9: Documentation
|
||||
|
||||
- [ ] CHANGELOG or version number updated to reflect this update
|
||||
- [ ] README updated if setup, usage, or environment variables changed
|
||||
- [ ] Any breaking change is called out explicitly, not buried in a generic commit message
|
||||
|
||||
---
|
||||
|
||||
## Section 10: Update Decision
|
||||
## Section 9: Maintenance Summary
|
||||
|
||||
After all sections are complete:
|
||||
|
||||
- List all unresolved FINDs grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW**
|
||||
- **CRITICAL or HIGH unresolved = NO GO.** Do not deploy until fixed — this includes a missing backup or missing rollback path.
|
||||
- List all unresolved findings grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW**
|
||||
- **CRITICAL or HIGH unresolved** = service needs action now (e.g. no working backup, expired/expiring cert, exploitable CVE, exposed DB port)
|
||||
- **MEDIUM/LOW unresolved** = user decides whether to defer with documented acceptance
|
||||
- Provide a final summary:
|
||||
- Total checks: X
|
||||
@@ -146,4 +183,6 @@ After all sections are complete:
|
||||
- Failed (critical): X
|
||||
- Failed (non-critical): X
|
||||
- Deferred: X
|
||||
- **Recommendation: GO / NO GO / GO WITH CONDITIONS**
|
||||
- Updates applied: X
|
||||
- **Overall: HEALTHY / HEALTHY WITH ACTIONS / NEEDS ATTENTION**
|
||||
- Save the summary to `reports/maintenance-<YYYY-MM-DD>.md` in the project directory so the next review can see what was deferred and what was running
|
||||
Reference in new issue
Block a user