updated and corrected to better fit the goal

This commit is contained in:
derekc committed 2026-09-23 23:37:36 -07:00
1 parent c84940ad81
commit 2f51362c96
2 files changed
+129 -89

No files matched your search

+2 -1
View File
@@ -10,7 +10,8 @@ A personal repository of markdown files that define how AI systems should work w
AI/
├── CLAUDE.md # Main entry point — loaded automatically by Claude Code
├── playbooks/
│ └── service-golive.md # Pre-go-live review checklist for services exposed via NPM
│ ├── service-golive.md # Pre-go-live review checklist for services exposed via NPM
│ └── service-update.md # Recurring maintenance review for already-live services
└── preferences/
├── articles.md # Blog post writing guide for chns.tech (Hugo/Markdown)
├── communication.md # How I like responses structured and how to interact
+127 -88
View File
@@ -1,144 +1,181 @@
# Playbook: Service Update Review
# Playbook: Live Service Maintenance Review
Use this playbook before deploying an update (code change, dependency bump, config change) to an already-running, externally-accessible homemade service.
Use this playbook on a recurring basis (e.g. monthly) to validate that an already-live, externally-accessible homemade service is still healthy, secure, backed up, and current — and to apply any updates it needs.
When invoked, read the project directory in the current working directory. Diff the current working tree/branch against the last git tag if one exists, or against `origin/main` if not. If neither gives a sensible comparison point, ask the user what to diff against before continuing.
This is not diff-driven. Apps degrade without code changes (new CVEs, expired certs, stopped backups, config drift, full disks), so every section runs every time.
When invoked, read the project directory in the current working directory. Some checks need the Docker host or the public URL — if the AI can't reach them directly, it gives the exact command to run and asks the user to paste the output.
---
## How to Use
Tell the AI: _"Use the service-update playbook to review this update: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_
Tell the AI: _"Use the service-update playbook to review this service: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_
The AI will:
1. Determine what changed since the last deployed version (git diff/log)
1. Read the project files and the most recent prior maintenance report (if one exists)
2. Work through each section below
3. For each item — report PASS, FAIL, or WARN with specific findings
4. At the end, give a go/no-go recommendation for the update
4. At the end, give an overall health rating and save a report
Do not proceed to the next section until the current one is resolved or explicitly deferred.
---
## Section 1: Change Summary
## Section 1: Baseline & Drift
Goal: Know exactly what's changing before evaluating risk.
Goal: Confirm what's actually running matches what's in git.
- [ ] List every file changed, added, or removed since the comparison point
- [ ] Summarize the intent of the change (bug fix, feature, dependency bump, refactor)
- [ ] Are there any TODO/FIXME/HACK comments left in the changed files?
- [ ] Does the change include a database schema change (migration file present)?
- [ ] Does the change add any new environment variables, config keys, or ports?
- [ ] Does the change add any new external dependency (package, API, service)?
- [ ] Working tree is clean and all commits are pushed to Gitea
- [ ] Running image tag/commit matches the latest git tag (or `main`) — no unknown version in production
- [ ] Live Docker Compose stack in Portainer/Dockhand matches the repo's `docker-compose.yml` (no edits made only in the UI)
- [ ] Live Nginx/NPM proxy host config matches what's documented in the repo
- [ ] Every variable in `.env.example` is set in the live `.env`, and the live `.env` has no stale/unused variables
- [ ] Running image tags/digests are recorded in the report (this is the rollback point if anything below needs changing)
**AI Action:** Run `git log <comparison>..HEAD --oneline` and `git diff <comparison>..HEAD --stat`. Use this to scope the rest of the review — sections below only need to focus on what actually changed.
**AI Action:** Run `git status`, `git log -1 --oneline`, `git describe --tags --abbrev=0`. Ask the user for `docker ps --format "{{.Names}}\t{{.Image}}\t{{.Status}}"` and `docker inspect --format "{{.Image}}" <container>` from the host if not reachable.
---
## Section 2: Backup Before Update
## Section 2: Runtime Health
Goal: Make sure the update is reversible before it starts.
Goal: The service is up and not quietly struggling.
- [ ] Current database/data volume has a fresh backup taken before applying the update
- [ ] Backup is confirmed to exist and is non-zero size (don't just assume the job ran)
- [ ] The currently-running image tag / commit hash is noted somewhere retrievable, so you know exactly what to roll back to
- [ ] The currently-deployed Docker Compose / Nginx config is saved or committed, not just live in NPM/Portainer
- [ ] All containers are running and report `healthy` (not just `Up`)
- [ ] No container has a climbing restart count (restart loop)
- [ ] Health endpoint responds via the public URL (through NPM), not just localhost
- [ ] Logs from the past review period show no recurring errors/exceptions or 5xx spikes
- [ ] CPU/memory usage is within resource limits with reasonable headroom
- [ ] Host disk and Docker volumes have adequate free space (flag WARN under 20% free)
- [ ] Log rotation (`max-size`/`max-file`) is working — log files aren't growing unbounded
- [ ] Reclaimable Docker disk space (dangling images, build cache) isn't excessive
**AI Action:** If no backup step is found for this service, flag as FAIL and stop — do not proceed to deployment sections until a backup exists.
**AI Action:** Commands to run on the host:
- `docker ps -a` and `docker inspect --format "{{.RestartCount}}" <container>`
- `docker stats --no-stream`
- `docker logs --since 720h <container> 2>&1 | grep -iE "error|exception|critical|traceback"`
- `df -h` and `docker system df`
- `curl -s -o /dev/null -w "%{http_code}" https://<public-url>/health`
---
## Section 3: Database Migrations
## Section 3: Functional Check
Goal: Don't lose or corrupt data on update. Skip this section if the change has no schema/data changes.
Goal: The app actually works for users, not just that the container is running.
- [ ] Migration scripts are idempotent or safe to re-run if interrupted
- [ ] Migration has a documented rollback (down migration, or manual revert steps if the framework doesn't support down migrations)
- [ ] Migration has been tested against a copy of real data, not just an empty dev database
- [ ] No destructive change (dropped column/table, changed column type with data loss) without a confirmed, tested backup
- [ ] Migration runs automatically on deploy, or the manual step to run it is documented
- [ ] Primary user flows work end-to-end (smoke test the main routes/functions)
- [ ] Login works; session expiry behaves as expected
- [ ] Scheduled jobs/background tasks have run recently (check last-run timestamps or logs)
- [ ] External integrations (email, third-party APIs, webhooks) still respond
- [ ] Ntfy still delivers — send a test notification and confirm it arrives
- [ ] Database connectivity and pool are healthy (no connection errors/timeouts in logs)
**AI Action:** Build a short smoke-test list from the app's routers/routes and ask the user to confirm each one.
---
## Section 4: Dependency & Supply Chain Delta
## Section 4: Backups & Recovery
Goal: A version bump shouldn't introduce a new vulnerability or a moving target.
Goal: If the host died today, the service could be restored.
- [ ] New or updated dependencies reviewed for known CVEs
- [ ] Dependency versions are still pinned to specific versions (not moved to `*` or unpinned)
- [ ] No new dependency pulled from an unverified/unofficial source
- [ ] If the Dockerfile base image changed, it's still from an official/verified source and still pinned (not moved to `:latest`)
- [ ] Most recent database backup is within the expected schedule (not days/weeks stale)
- [ ] Backup files are non-zero size and not truncated
- [ ] Backups are stored off the Docker host (separate disk or remote target)
- [ ] Upload/file volumes are backed up, not just the database
- [ ] Live `.env` and any UI-only config are backed up somewhere retrievable (they aren't in git)
- [ ] Retention is working — old backups are pruned, not filling the disk
- [ ] A test restore has been performed within the last 6 months
**AI Action:** If no backup is found or the latest backup is stale, flag as CRITICAL.
---
## Section 5: Security Posture
Goal: Nothing about the service's exposure has weakened since go-live.
### 5a. TLS & Headers
- [ ] TLS certificate is valid with more than 14 days to expiry, and NPM auto-renewal is enabled
- [ ] HTTP still redirects to HTTPS
- [ ] Security headers from go-live are still present (`X-Frame-Options`, `X-Content-Type-Options`, `Referrer-Policy`, `Content-Security-Policy`, `Strict-Transport-Security`)
- [ ] Server version header still suppressed
### 5b. Exposure
- [ ] Only 80/443 are exposed publicly — no app or database ports published to the host/internet
- [ ] NPM access lists / basic auth still in place for any internal-only routes
- [ ] No new containers or ports added to the stack that weren't reviewed
### 5c. Accounts & Secrets
- [ ] Admin and user accounts reviewed — no stale, unknown, or unexpected admin accounts
- [ ] Secrets (DB passwords, API keys, secret keys) rotated per policy, or age noted if never rotated
- [ ] No unusual failed-login spikes, rate-limit triggers, or unknown admin logins in logs/Ntfy history
### 5d. Container Hardening
- [ ] Containers still run as non-root, no `privileged: true`, no unintended volume mounts (e.g. `docker.sock`)
- [ ] Resource limits (`mem_limit`, `cpus`) still set
**AI Action:** Run:
- `curl -sI https://<public-url>` — verify headers
- `echo | openssl s_client -connect <host>:443 -servername <host> 2>/dev/null | openssl x509 -noout -enddate` — cert expiry
- `docker ps --format "{{.Names}}\t{{.Ports}}"` — published ports
---
## Section 6: Vulnerabilities & Currency
Goal: Find what needs updating — new CVEs appear in code that hasn't changed.
- [ ] Running images scanned for HIGH/CRITICAL CVEs (includes base image OS packages)
- [ ] Application dependencies scanned for known CVEs
- [ ] Outdated dependencies identified — note which are security fixes vs. feature bumps
- [ ] Base images have newer patch versions available
- [ ] Runtimes/services are not end-of-life or near EOL (Python, MySQL, Node, Nginx, etc. — check https://endoflife.date)
- [ ] No dependencies that have become abandoned/unmaintained
- [ ] Dependencies and base images are still pinned (no `*`, no `:latest`)
**AI Action:** Run whichever applies:
- `trivy image --severity HIGH,CRITICAL <image:tag>` for each running image
- `trivy fs . --severity HIGH,CRITICAL`
- `pip-audit` (Python)
- `npm audit` (Node)
- `pip-audit` and `pip list --outdated` (Python)
- `npm audit` and `npm outdated` (Node)
Report new findings only if they weren't already accepted/deferred in a prior review.
For a full image-level audit (SBOM, secrets, misconfig, grype), use the `docker-security-audit` playbook. Report new findings only if they aren't already accepted/deferred in the prior maintenance report.
---
## Section 5: Security Delta
## Section 7: Applying Updates
Goal: Re-check only what changed — this is not a full re-audit, but changed code gets the same scrutiny as go-live.
Goal: Apply needed updates safely. Skip this section if Section 6 found nothing to update.
- [ ] New code introduces no hardcoded secrets, passwords, or tokens
- [ ] New or changed endpoints/routes require authentication consistent with the rest of the app
- [ ] New user input is validated server-side and uses parameterized queries / ORM (no string-built SQL)
- [ ] No debug/dev tooling left enabled (e.g. `debug=True`, exposed `/docs`, verbose stack traces)
- [ ] Any new file upload, redirect, or deserialization logic is checked for path traversal, open redirect, and unsafe object instantiation
- [ ] Any new HTML output is escaped (no new XSS surface)
- [ ] Fresh backup taken and confirmed (Section 4) immediately before updating
- [ ] Current image tags/digests recorded as the rollback point (Section 1)
- [ ] Release notes reviewed for any major-version bump — breaking changes identified
- [ ] Versions bumped to specific pinned versions (not `:latest` or unpinned)
- [ ] Any DB schema/data change has a tested rollback path and was tested against a copy of real data
- [ ] New image builds and starts locally, and passes its health check, before deploying
- [ ] Deployed under a specific version tag; downtime (if any) is expected/scheduled
- [ ] Sections 2 and 3 re-run after deploy — all PASS
- [ ] New version tagged in Gitea and `Release-Notes/v{major}.{minor}.md` updated
- [ ] Old images kept until the update is confirmed stable, then cleaned up
**AI Action:** If any step fails after deploy, stop and walk the user through rollback to the recorded image tags (and backup restore if a migration ran).
---
## Section 6: Config & Environment Changes
## Section 8: Documentation
Goal: Config drift is how services quietly become insecure between go-live and now.
- [ ] Any new required environment variable is added to `.env.example` with a placeholder value
- [ ] No new secrets added directly to Docker Compose or Nginx config files
- [ ] Nginx config changes (if any) don't remove or weaken existing security headers
- [ ] Container `user:`, `privileged:`, volume mounts, and resource limits are unchanged or still compliant with go-live standards
- [ ] README still accurate for setup, usage, and environment variables
- [ ] Rollback / "what to do if this breaks" runbook still accurate for the current setup
- [ ] Any deferred items from this review are recorded in the report with a reason
---
## Section 7: Deployment Process
Goal: Get the update live without unplanned downtime or a stuck rollback.
- [ ] Update is deployed under a specific version tag, not floating `:latest`
- [ ] Deployment approach causes minimal/no downtime, or downtime is scheduled and expected
- [ ] Rollback procedure (previous image tag, compose file, backup restore) is known before starting, not figured out after something breaks
- [ ] Health check endpoint is confirmed reachable immediately after deploy
---
## Section 8: Post-Update Verification
Goal: Confirm the update actually works in production, not just that it deployed.
- [ ] Primary routes/functions respond correctly after deploy (smoke test the main flows)
- [ ] Logs show no new errors/exceptions in the minutes following deploy
- [ ] Ntfy/notification integration (if implemented) still fires correctly
- [ ] No unexpected spike in CPU/memory/disk after deploy
- [ ] Old container/image cleaned up once the new one is confirmed stable (if not needed for rollback anymore)
---
## Section 9: Documentation
- [ ] CHANGELOG or version number updated to reflect this update
- [ ] README updated if setup, usage, or environment variables changed
- [ ] Any breaking change is called out explicitly, not buried in a generic commit message
---
## Section 10: Update Decision
## Section 9: Maintenance Summary
After all sections are complete:
- List all unresolved FINDs grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW**
- **CRITICAL or HIGH unresolved = NO GO.** Do not deploy until fixed — this includes a missing backup or missing rollback path.
- List all unresolved findings grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW**
- **CRITICAL or HIGH unresolved** = service needs action now (e.g. no working backup, expired/expiring cert, exploitable CVE, exposed DB port)
- **MEDIUM/LOW unresolved** = user decides whether to defer with documented acceptance
- Provide a final summary:
- Total checks: X
@@ -146,4 +183,6 @@ After all sections are complete:
- Failed (critical): X
- Failed (non-critical): X
- Deferred: X
- **Recommendation: GO / NO GO / GO WITH CONDITIONS**
- Updates applied: X
- **Overall: HEALTHY / HEALTHY WITH ACTIONS / NEEDS ATTENTION**
- Save the summary to `reports/maintenance-<YYYY-MM-DD>.md` in the project directory so the next review can see what was deferred and what was running