From 2f51362c96162d723106ecf5aa9f2b74641b4396 Mon Sep 17 00:00:00 2001 From: Derek Cooper Date: Wed, 23 Sep 2026 23:37:36 -0700 Subject: [PATCH] updated and corrected to better fit the goal --- README.md | 3 +- playbooks/service-update.md | 215 +++++++++++++++++++++--------------- 2 files changed, 129 insertions(+), 89 deletions(-) diff --git a/README.md b/README.md index aa82895..90629d7 100644 --- a/README.md +++ b/README.md @@ -10,7 +10,8 @@ A personal repository of markdown files that define how AI systems should work w AI/ ├── CLAUDE.md # Main entry point — loaded automatically by Claude Code ├── playbooks/ -│ └── service-golive.md # Pre-go-live review checklist for services exposed via NPM +│ ├── service-golive.md # Pre-go-live review checklist for services exposed via NPM +│ └── service-update.md # Recurring maintenance review for already-live services └── preferences/ ├── articles.md # Blog post writing guide for chns.tech (Hugo/Markdown) ├── communication.md # How I like responses structured and how to interact diff --git a/playbooks/service-update.md b/playbooks/service-update.md index 4d1d424..d54a4a8 100644 --- a/playbooks/service-update.md +++ b/playbooks/service-update.md @@ -1,144 +1,181 @@ -# Playbook: Service Update Review +# Playbook: Live Service Maintenance Review -Use this playbook before deploying an update (code change, dependency bump, config change) to an already-running, externally-accessible homemade service. +Use this playbook on a recurring basis (e.g. monthly) to validate that an already-live, externally-accessible homemade service is still healthy, secure, backed up, and current — and to apply any updates it needs. -When invoked, read the project directory in the current working directory. Diff the current working tree/branch against the last git tag if one exists, or against `origin/main` if not. If neither gives a sensible comparison point, ask the user what to diff against before continuing. +This is not diff-driven. Apps degrade without code changes (new CVEs, expired certs, stopped backups, config drift, full disks), so every section runs every time. + +When invoked, read the project directory in the current working directory. Some checks need the Docker host or the public URL — if the AI can't reach them directly, it gives the exact command to run and asks the user to paste the output. --- ## How to Use -Tell the AI: _"Use the service-update playbook to review this update: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_ +Tell the AI: _"Use the service-update playbook to review this service: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_ The AI will: -1. Determine what changed since the last deployed version (git diff/log) +1. Read the project files and the most recent prior maintenance report (if one exists) 2. Work through each section below 3. For each item — report PASS, FAIL, or WARN with specific findings -4. At the end, give a go/no-go recommendation for the update +4. At the end, give an overall health rating and save a report Do not proceed to the next section until the current one is resolved or explicitly deferred. --- -## Section 1: Change Summary +## Section 1: Baseline & Drift -Goal: Know exactly what's changing before evaluating risk. +Goal: Confirm what's actually running matches what's in git. -- [ ] List every file changed, added, or removed since the comparison point -- [ ] Summarize the intent of the change (bug fix, feature, dependency bump, refactor) -- [ ] Are there any TODO/FIXME/HACK comments left in the changed files? -- [ ] Does the change include a database schema change (migration file present)? -- [ ] Does the change add any new environment variables, config keys, or ports? -- [ ] Does the change add any new external dependency (package, API, service)? +- [ ] Working tree is clean and all commits are pushed to Gitea +- [ ] Running image tag/commit matches the latest git tag (or `main`) — no unknown version in production +- [ ] Live Docker Compose stack in Portainer/Dockhand matches the repo's `docker-compose.yml` (no edits made only in the UI) +- [ ] Live Nginx/NPM proxy host config matches what's documented in the repo +- [ ] Every variable in `.env.example` is set in the live `.env`, and the live `.env` has no stale/unused variables +- [ ] Running image tags/digests are recorded in the report (this is the rollback point if anything below needs changing) -**AI Action:** Run `git log ..HEAD --oneline` and `git diff ..HEAD --stat`. Use this to scope the rest of the review — sections below only need to focus on what actually changed. +**AI Action:** Run `git status`, `git log -1 --oneline`, `git describe --tags --abbrev=0`. Ask the user for `docker ps --format "{{.Names}}\t{{.Image}}\t{{.Status}}"` and `docker inspect --format "{{.Image}}" ` from the host if not reachable. --- -## Section 2: Backup Before Update +## Section 2: Runtime Health -Goal: Make sure the update is reversible before it starts. +Goal: The service is up and not quietly struggling. -- [ ] Current database/data volume has a fresh backup taken before applying the update -- [ ] Backup is confirmed to exist and is non-zero size (don't just assume the job ran) -- [ ] The currently-running image tag / commit hash is noted somewhere retrievable, so you know exactly what to roll back to -- [ ] The currently-deployed Docker Compose / Nginx config is saved or committed, not just live in NPM/Portainer +- [ ] All containers are running and report `healthy` (not just `Up`) +- [ ] No container has a climbing restart count (restart loop) +- [ ] Health endpoint responds via the public URL (through NPM), not just localhost +- [ ] Logs from the past review period show no recurring errors/exceptions or 5xx spikes +- [ ] CPU/memory usage is within resource limits with reasonable headroom +- [ ] Host disk and Docker volumes have adequate free space (flag WARN under 20% free) +- [ ] Log rotation (`max-size`/`max-file`) is working — log files aren't growing unbounded +- [ ] Reclaimable Docker disk space (dangling images, build cache) isn't excessive -**AI Action:** If no backup step is found for this service, flag as FAIL and stop — do not proceed to deployment sections until a backup exists. +**AI Action:** Commands to run on the host: +- `docker ps -a` and `docker inspect --format "{{.RestartCount}}" ` +- `docker stats --no-stream` +- `docker logs --since 720h 2>&1 | grep -iE "error|exception|critical|traceback"` +- `df -h` and `docker system df` +- `curl -s -o /dev/null -w "%{http_code}" https:///health` --- -## Section 3: Database Migrations +## Section 3: Functional Check -Goal: Don't lose or corrupt data on update. Skip this section if the change has no schema/data changes. +Goal: The app actually works for users, not just that the container is running. -- [ ] Migration scripts are idempotent or safe to re-run if interrupted -- [ ] Migration has a documented rollback (down migration, or manual revert steps if the framework doesn't support down migrations) -- [ ] Migration has been tested against a copy of real data, not just an empty dev database -- [ ] No destructive change (dropped column/table, changed column type with data loss) without a confirmed, tested backup -- [ ] Migration runs automatically on deploy, or the manual step to run it is documented +- [ ] Primary user flows work end-to-end (smoke test the main routes/functions) +- [ ] Login works; session expiry behaves as expected +- [ ] Scheduled jobs/background tasks have run recently (check last-run timestamps or logs) +- [ ] External integrations (email, third-party APIs, webhooks) still respond +- [ ] Ntfy still delivers — send a test notification and confirm it arrives +- [ ] Database connectivity and pool are healthy (no connection errors/timeouts in logs) + +**AI Action:** Build a short smoke-test list from the app's routers/routes and ask the user to confirm each one. --- -## Section 4: Dependency & Supply Chain Delta +## Section 4: Backups & Recovery -Goal: A version bump shouldn't introduce a new vulnerability or a moving target. +Goal: If the host died today, the service could be restored. -- [ ] New or updated dependencies reviewed for known CVEs -- [ ] Dependency versions are still pinned to specific versions (not moved to `*` or unpinned) -- [ ] No new dependency pulled from an unverified/unofficial source -- [ ] If the Dockerfile base image changed, it's still from an official/verified source and still pinned (not moved to `:latest`) +- [ ] Most recent database backup is within the expected schedule (not days/weeks stale) +- [ ] Backup files are non-zero size and not truncated +- [ ] Backups are stored off the Docker host (separate disk or remote target) +- [ ] Upload/file volumes are backed up, not just the database +- [ ] Live `.env` and any UI-only config are backed up somewhere retrievable (they aren't in git) +- [ ] Retention is working — old backups are pruned, not filling the disk +- [ ] A test restore has been performed within the last 6 months + +**AI Action:** If no backup is found or the latest backup is stale, flag as CRITICAL. + +--- + +## Section 5: Security Posture + +Goal: Nothing about the service's exposure has weakened since go-live. + +### 5a. TLS & Headers +- [ ] TLS certificate is valid with more than 14 days to expiry, and NPM auto-renewal is enabled +- [ ] HTTP still redirects to HTTPS +- [ ] Security headers from go-live are still present (`X-Frame-Options`, `X-Content-Type-Options`, `Referrer-Policy`, `Content-Security-Policy`, `Strict-Transport-Security`) +- [ ] Server version header still suppressed + +### 5b. Exposure +- [ ] Only 80/443 are exposed publicly — no app or database ports published to the host/internet +- [ ] NPM access lists / basic auth still in place for any internal-only routes +- [ ] No new containers or ports added to the stack that weren't reviewed + +### 5c. Accounts & Secrets +- [ ] Admin and user accounts reviewed — no stale, unknown, or unexpected admin accounts +- [ ] Secrets (DB passwords, API keys, secret keys) rotated per policy, or age noted if never rotated +- [ ] No unusual failed-login spikes, rate-limit triggers, or unknown admin logins in logs/Ntfy history + +### 5d. Container Hardening +- [ ] Containers still run as non-root, no `privileged: true`, no unintended volume mounts (e.g. `docker.sock`) +- [ ] Resource limits (`mem_limit`, `cpus`) still set + +**AI Action:** Run: +- `curl -sI https://` — verify headers +- `echo | openssl s_client -connect :443 -servername 2>/dev/null | openssl x509 -noout -enddate` — cert expiry +- `docker ps --format "{{.Names}}\t{{.Ports}}"` — published ports + +--- + +## Section 6: Vulnerabilities & Currency + +Goal: Find what needs updating — new CVEs appear in code that hasn't changed. + +- [ ] Running images scanned for HIGH/CRITICAL CVEs (includes base image OS packages) +- [ ] Application dependencies scanned for known CVEs +- [ ] Outdated dependencies identified — note which are security fixes vs. feature bumps +- [ ] Base images have newer patch versions available +- [ ] Runtimes/services are not end-of-life or near EOL (Python, MySQL, Node, Nginx, etc. — check https://endoflife.date) +- [ ] No dependencies that have become abandoned/unmaintained +- [ ] Dependencies and base images are still pinned (no `*`, no `:latest`) **AI Action:** Run whichever applies: +- `trivy image --severity HIGH,CRITICAL ` for each running image - `trivy fs . --severity HIGH,CRITICAL` -- `pip-audit` (Python) -- `npm audit` (Node) +- `pip-audit` and `pip list --outdated` (Python) +- `npm audit` and `npm outdated` (Node) -Report new findings only if they weren't already accepted/deferred in a prior review. +For a full image-level audit (SBOM, secrets, misconfig, grype), use the `docker-security-audit` playbook. Report new findings only if they aren't already accepted/deferred in the prior maintenance report. --- -## Section 5: Security Delta +## Section 7: Applying Updates -Goal: Re-check only what changed — this is not a full re-audit, but changed code gets the same scrutiny as go-live. +Goal: Apply needed updates safely. Skip this section if Section 6 found nothing to update. -- [ ] New code introduces no hardcoded secrets, passwords, or tokens -- [ ] New or changed endpoints/routes require authentication consistent with the rest of the app -- [ ] New user input is validated server-side and uses parameterized queries / ORM (no string-built SQL) -- [ ] No debug/dev tooling left enabled (e.g. `debug=True`, exposed `/docs`, verbose stack traces) -- [ ] Any new file upload, redirect, or deserialization logic is checked for path traversal, open redirect, and unsafe object instantiation -- [ ] Any new HTML output is escaped (no new XSS surface) +- [ ] Fresh backup taken and confirmed (Section 4) immediately before updating +- [ ] Current image tags/digests recorded as the rollback point (Section 1) +- [ ] Release notes reviewed for any major-version bump — breaking changes identified +- [ ] Versions bumped to specific pinned versions (not `:latest` or unpinned) +- [ ] Any DB schema/data change has a tested rollback path and was tested against a copy of real data +- [ ] New image builds and starts locally, and passes its health check, before deploying +- [ ] Deployed under a specific version tag; downtime (if any) is expected/scheduled +- [ ] Sections 2 and 3 re-run after deploy — all PASS +- [ ] New version tagged in Gitea and `Release-Notes/v{major}.{minor}.md` updated +- [ ] Old images kept until the update is confirmed stable, then cleaned up + +**AI Action:** If any step fails after deploy, stop and walk the user through rollback to the recorded image tags (and backup restore if a migration ran). --- -## Section 6: Config & Environment Changes +## Section 8: Documentation -Goal: Config drift is how services quietly become insecure between go-live and now. - -- [ ] Any new required environment variable is added to `.env.example` with a placeholder value -- [ ] No new secrets added directly to Docker Compose or Nginx config files -- [ ] Nginx config changes (if any) don't remove or weaken existing security headers -- [ ] Container `user:`, `privileged:`, volume mounts, and resource limits are unchanged or still compliant with go-live standards +- [ ] README still accurate for setup, usage, and environment variables +- [ ] Rollback / "what to do if this breaks" runbook still accurate for the current setup +- [ ] Any deferred items from this review are recorded in the report with a reason --- -## Section 7: Deployment Process - -Goal: Get the update live without unplanned downtime or a stuck rollback. - -- [ ] Update is deployed under a specific version tag, not floating `:latest` -- [ ] Deployment approach causes minimal/no downtime, or downtime is scheduled and expected -- [ ] Rollback procedure (previous image tag, compose file, backup restore) is known before starting, not figured out after something breaks -- [ ] Health check endpoint is confirmed reachable immediately after deploy - ---- - -## Section 8: Post-Update Verification - -Goal: Confirm the update actually works in production, not just that it deployed. - -- [ ] Primary routes/functions respond correctly after deploy (smoke test the main flows) -- [ ] Logs show no new errors/exceptions in the minutes following deploy -- [ ] Ntfy/notification integration (if implemented) still fires correctly -- [ ] No unexpected spike in CPU/memory/disk after deploy -- [ ] Old container/image cleaned up once the new one is confirmed stable (if not needed for rollback anymore) - ---- - -## Section 9: Documentation - -- [ ] CHANGELOG or version number updated to reflect this update -- [ ] README updated if setup, usage, or environment variables changed -- [ ] Any breaking change is called out explicitly, not buried in a generic commit message - ---- - -## Section 10: Update Decision +## Section 9: Maintenance Summary After all sections are complete: -- List all unresolved FINDs grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW** -- **CRITICAL or HIGH unresolved = NO GO.** Do not deploy until fixed — this includes a missing backup or missing rollback path. +- List all unresolved findings grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW** +- **CRITICAL or HIGH unresolved** = service needs action now (e.g. no working backup, expired/expiring cert, exploitable CVE, exposed DB port) - **MEDIUM/LOW unresolved** = user decides whether to defer with documented acceptance - Provide a final summary: - Total checks: X @@ -146,4 +183,6 @@ After all sections are complete: - Failed (critical): X - Failed (non-critical): X - Deferred: X - - **Recommendation: GO / NO GO / GO WITH CONDITIONS** + - Updates applied: X + - **Overall: HEALTHY / HEALTHY WITH ACTIONS / NEEDS ATTENTION** +- Save the summary to `reports/maintenance-.md` in the project directory so the next review can see what was deferred and what was running