diff --git a/playbooks/service-update.md b/playbooks/service-update.md new file mode 100644 index 0000000..4d1d424 --- /dev/null +++ b/playbooks/service-update.md @@ -0,0 +1,149 @@ +# Playbook: Service Update Review + +Use this playbook before deploying an update (code change, dependency bump, config change) to an already-running, externally-accessible homemade service. + +When invoked, read the project directory in the current working directory. Diff the current working tree/branch against the last git tag if one exists, or against `origin/main` if not. If neither gives a sensible comparison point, ask the user what to diff against before continuing. + +--- + +## How to Use + +Tell the AI: _"Use the service-update playbook to review this update: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_ + +The AI will: +1. Determine what changed since the last deployed version (git diff/log) +2. Work through each section below +3. For each item — report PASS, FAIL, or WARN with specific findings +4. At the end, give a go/no-go recommendation for the update + +Do not proceed to the next section until the current one is resolved or explicitly deferred. + +--- + +## Section 1: Change Summary + +Goal: Know exactly what's changing before evaluating risk. + +- [ ] List every file changed, added, or removed since the comparison point +- [ ] Summarize the intent of the change (bug fix, feature, dependency bump, refactor) +- [ ] Are there any TODO/FIXME/HACK comments left in the changed files? +- [ ] Does the change include a database schema change (migration file present)? +- [ ] Does the change add any new environment variables, config keys, or ports? +- [ ] Does the change add any new external dependency (package, API, service)? + +**AI Action:** Run `git log ..HEAD --oneline` and `git diff ..HEAD --stat`. Use this to scope the rest of the review — sections below only need to focus on what actually changed. + +--- + +## Section 2: Backup Before Update + +Goal: Make sure the update is reversible before it starts. + +- [ ] Current database/data volume has a fresh backup taken before applying the update +- [ ] Backup is confirmed to exist and is non-zero size (don't just assume the job ran) +- [ ] The currently-running image tag / commit hash is noted somewhere retrievable, so you know exactly what to roll back to +- [ ] The currently-deployed Docker Compose / Nginx config is saved or committed, not just live in NPM/Portainer + +**AI Action:** If no backup step is found for this service, flag as FAIL and stop — do not proceed to deployment sections until a backup exists. + +--- + +## Section 3: Database Migrations + +Goal: Don't lose or corrupt data on update. Skip this section if the change has no schema/data changes. + +- [ ] Migration scripts are idempotent or safe to re-run if interrupted +- [ ] Migration has a documented rollback (down migration, or manual revert steps if the framework doesn't support down migrations) +- [ ] Migration has been tested against a copy of real data, not just an empty dev database +- [ ] No destructive change (dropped column/table, changed column type with data loss) without a confirmed, tested backup +- [ ] Migration runs automatically on deploy, or the manual step to run it is documented + +--- + +## Section 4: Dependency & Supply Chain Delta + +Goal: A version bump shouldn't introduce a new vulnerability or a moving target. + +- [ ] New or updated dependencies reviewed for known CVEs +- [ ] Dependency versions are still pinned to specific versions (not moved to `*` or unpinned) +- [ ] No new dependency pulled from an unverified/unofficial source +- [ ] If the Dockerfile base image changed, it's still from an official/verified source and still pinned (not moved to `:latest`) + +**AI Action:** Run whichever applies: +- `trivy fs . --severity HIGH,CRITICAL` +- `pip-audit` (Python) +- `npm audit` (Node) + +Report new findings only if they weren't already accepted/deferred in a prior review. + +--- + +## Section 5: Security Delta + +Goal: Re-check only what changed — this is not a full re-audit, but changed code gets the same scrutiny as go-live. + +- [ ] New code introduces no hardcoded secrets, passwords, or tokens +- [ ] New or changed endpoints/routes require authentication consistent with the rest of the app +- [ ] New user input is validated server-side and uses parameterized queries / ORM (no string-built SQL) +- [ ] No debug/dev tooling left enabled (e.g. `debug=True`, exposed `/docs`, verbose stack traces) +- [ ] Any new file upload, redirect, or deserialization logic is checked for path traversal, open redirect, and unsafe object instantiation +- [ ] Any new HTML output is escaped (no new XSS surface) + +--- + +## Section 6: Config & Environment Changes + +Goal: Config drift is how services quietly become insecure between go-live and now. + +- [ ] Any new required environment variable is added to `.env.example` with a placeholder value +- [ ] No new secrets added directly to Docker Compose or Nginx config files +- [ ] Nginx config changes (if any) don't remove or weaken existing security headers +- [ ] Container `user:`, `privileged:`, volume mounts, and resource limits are unchanged or still compliant with go-live standards + +--- + +## Section 7: Deployment Process + +Goal: Get the update live without unplanned downtime or a stuck rollback. + +- [ ] Update is deployed under a specific version tag, not floating `:latest` +- [ ] Deployment approach causes minimal/no downtime, or downtime is scheduled and expected +- [ ] Rollback procedure (previous image tag, compose file, backup restore) is known before starting, not figured out after something breaks +- [ ] Health check endpoint is confirmed reachable immediately after deploy + +--- + +## Section 8: Post-Update Verification + +Goal: Confirm the update actually works in production, not just that it deployed. + +- [ ] Primary routes/functions respond correctly after deploy (smoke test the main flows) +- [ ] Logs show no new errors/exceptions in the minutes following deploy +- [ ] Ntfy/notification integration (if implemented) still fires correctly +- [ ] No unexpected spike in CPU/memory/disk after deploy +- [ ] Old container/image cleaned up once the new one is confirmed stable (if not needed for rollback anymore) + +--- + +## Section 9: Documentation + +- [ ] CHANGELOG or version number updated to reflect this update +- [ ] README updated if setup, usage, or environment variables changed +- [ ] Any breaking change is called out explicitly, not buried in a generic commit message + +--- + +## Section 10: Update Decision + +After all sections are complete: + +- List all unresolved FINDs grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW** +- **CRITICAL or HIGH unresolved = NO GO.** Do not deploy until fixed — this includes a missing backup or missing rollback path. +- **MEDIUM/LOW unresolved** = user decides whether to defer with documented acceptance +- Provide a final summary: + - Total checks: X + - Passed: X + - Failed (critical): X + - Failed (non-critical): X + - Deferred: X + - **Recommendation: GO / NO GO / GO WITH CONDITIONS**