diff --git a/playbooks/service-golive.md b/playbooks/service-golive.md index 31da3c9..0426ff2 100644 --- a/playbooks/service-golive.md +++ b/playbooks/service-golive.md @@ -13,10 +13,25 @@ The AI will: 1. Read the project files in the current directory 2. Work through each section below 3. For each item — report PASS, FAIL, or WARN with specific findings -4. At the end, give a go/no-go recommendation +4. At the end, give a go/no-go recommendation and save a report Do not proceed to the next section until the current one is resolved or explicitly deferred. +**Severity:** Items tagged `[CRITICAL]`, `[HIGH]`, `[MEDIUM]`, or `[LOW]` use that severity when they FAIL. For untagged items, the AI assigns severity based on how exploitable or disruptive the gap is and states its reasoning. + +--- + +## Section 0: Pre-flight + +Goal: Make sure the review tools are available before relying on them. + +- [ ] `trivy --version` +- [ ] `bandit --version` (Python projects) +- [ ] `gitleaks version` +- [ ] `curl` and `openssl` available + +**AI Action:** If a tool is missing, tell the user and ask whether to install it or skip the checks that depend on it. Skipped checks are reported as WARN, not PASS. + --- ## Section 1: Feature & Improvement Review @@ -24,6 +39,8 @@ Do not proceed to the next section until the current one is resolved or explicit Goal: Catch missing functionality before users find it. - [ ] Does the service have a health check endpoint (e.g. `/health` or `/ping`)? +- [ ] Does the app container have its own Compose `healthcheck` wired to that endpoint (not just the database)? +- [ ] Is `restart: unless-stopped` (or `always`) set on every service? - [ ] Are all intended routes/endpoints implemented and reachable? - [ ] Is there a meaningful error response for bad input (not raw stack traces)? - [ ] Are there any obvious UX gaps or incomplete flows in the UI (if applicable)? @@ -33,12 +50,15 @@ Goal: Catch missing functionality before users find it. ### 1a. Ntfy Admin Notifications -Goal: Ensure the super admin is alerted to significant events without having to monitor logs manually. +Goal: Ensure the super admin is alerted to significant events without having to monitor logs manually. Ntfy is required for all web apps (see `preferences/python-web.md`). -- [ ] Is Ntfy (or equivalent push notification system) integrated into the application? +- [ ] `[HIGH]` Is Ntfy integrated into the application? - [ ] Are admin-relevant events triggering Ntfy notifications? +- [ ] `[HIGH]` Is the Ntfy topic access-protected (token or user auth on the self-hosted instance) so others can't subscribe to admin alerts? +- [ ] Are Ntfy calls wrapped in `try/except` so a failed notification can't crash the app? +- [ ] Ntfy URL, topic, and token come from `.env` — not hardcoded -**If Ntfy is NOT implemented**, flag as WARN and recommend the following events for notification coverage based on what the app does: +**If Ntfy is NOT implemented**, flag as FAIL and recommend the following events for notification coverage based on what the app does: | Event | Severity | Why it matters | |---|---|---| @@ -55,7 +75,18 @@ Goal: Ensure the super admin is alerted to significant events without having to | Large data export initiated | Medium | Data exfiltration risk indicator | | Config or environment change detected | High | Unplanned changes should be visible | -**AI Action:** Search the codebase for Ntfy integration (look for `ntfy`, `ntfy.sh`, or HTTP POST calls to a notification endpoint). If none found, list the above recommended events as WARN items and ask the user whether to implement before go-live or defer. +**AI Action:** Search the codebase for Ntfy integration (look for `ntfy`, `ntfy.sh`, or HTTP POST calls to a notification endpoint). If none found, list the above recommended events as FAIL items and ask the user whether to implement before go-live. + +### 1b. Uptime Kuma Monitoring + +Goal: Get alerted when the service is down — Ntfy runs inside the app, so it can't report its own outage. + +- [ ] `[HIGH]` An Uptime Kuma monitor exists for this service, pointed at the public health endpoint URL (through NPM), not the internal container address +- [ ] The monitor has a notification configured (e.g. Ntfy) so downtime actually alerts someone +- [ ] Monitor uses an HTTP(s) status or keyword check, not just a ping/TCP check +- [ ] Certificate expiry notification is enabled on the monitor + +**AI Action:** Ask the user to confirm the monitor is set up — the AI can't see Uptime Kuma directly. --- @@ -86,95 +117,135 @@ Goal: Ensure the service won't collapse under real load. Goal: Do not put a vulnerable service on the internet. Be thorough. ### 3a. Secrets & Credentials -- [ ] No hardcoded passwords, tokens, API keys, or secrets in any source file -- [ ] `.env` file is in `.gitignore` and not committed +- [ ] `[CRITICAL]` No hardcoded passwords, tokens, API keys, or secrets in any source file +- [ ] `[CRITICAL]` `.env` file is in `.gitignore` and not committed +- [ ] `[CRITICAL]` No secrets anywhere in git history (a previously committed `.env` is still exposed even if deleted later — rotate any secret found) +- [ ] `[HIGH]` `.dockerignore` exists and excludes `.env`, `.git`, and other local-only files so they aren't baked into the image - [ ] `.env.example` exists with placeholder values only -- [ ] No secrets in Docker Compose files (use `env_file` or environment variable references, not literal values) -- [ ] No secrets in Nginx config files +- [ ] `[HIGH]` No secrets in Docker Compose files (use `env_file` or environment variable references, not literal values) +- [ ] `[HIGH]` No secrets in Nginx config files ### 3b. Authentication & Authorization -- [ ] All non-public endpoints require authentication +- [ ] `[CRITICAL]` All non-public endpoints require authentication - [ ] Authentication tokens/sessions have an expiry -- [ ] Password hashing uses bcrypt, argon2, or scrypt — not MD5/SHA1 -- [ ] There is no default admin password that ships with the service +- [ ] `[HIGH]` Password hashing uses bcrypt, argon2, or scrypt — not MD5/SHA1 +- [ ] `[CRITICAL]` There is no default admin password that ships with the service - [ ] Role/permission checks exist if the app has multiple access levels -- [ ] Failed login attempts are rate-limited or account-locked after N failures +- [ ] `[HIGH]` Failed login attempts are rate-limited or account-locked after N failures +- [ ] Login and password-reset responses don't reveal whether an account exists (same message and similar timing for valid/invalid users) +- [ ] Logout and password change invalidate existing sessions/tokens ### 3c. Input Validation & Injection - [ ] All user input is validated server-side (not just client-side) -- [ ] SQL queries use parameterized statements or ORM — no string concatenation -- [ ] File upload paths are sanitized — no path traversal possible -- [ ] HTML output is escaped to prevent XSS (or a framework handles this automatically) +- [ ] `[CRITICAL]` SQL queries use parameterized statements or ORM — no string concatenation +- [ ] `[HIGH]` File upload paths are sanitized — no path traversal possible +- [ ] `[HIGH]` HTML output is escaped to prevent XSS (or a framework handles this automatically) - [ ] Redirects only go to allowed/relative URLs — no open redirect -- [ ] JSON deserialization does not allow arbitrary object instantiation +- [ ] `[HIGH]` JSON deserialization does not allow arbitrary object instantiation ### 3d. HTTP & Nginx Security Headers Verify the Nginx config for the proxy host includes: - [ ] `X-Frame-Options: DENY` or `SAMEORIGIN` - [ ] `X-Content-Type-Options: nosniff` -- [ ] `X-XSS-Protection: 1; mode=block` +- [ ] `X-XSS-Protection: 0` or omitted — the legacy `1; mode=block` value is deprecated and can introduce vulnerabilities; CSP replaces it - [ ] `Referrer-Policy: strict-origin-when-cross-origin` - [ ] `Content-Security-Policy` header defined (even if broad to start) -- [ ] `Strict-Transport-Security` (HSTS) with `max-age` >= 31536000 +- [ ] `Permissions-Policy` restricts browser features the app doesn't use (e.g. `camera=(), microphone=(), geolocation=()`) +- [ ] `[MEDIUM]` `Strict-Transport-Security` (HSTS) with `max-age` >= 31536000 - [ ] Server version header suppressed (`server_tokens off`) - [ ] Unnecessary HTTP methods disabled (e.g. TRACE, DELETE if not used) ### 3e. TLS / HTTPS -- [ ] TLS certificate is valid and not self-signed for production -- [ ] HTTP traffic redirects to HTTPS (not served in parallel) +- [ ] `[CRITICAL]` TLS certificate is valid and not self-signed for production +- [ ] `[HIGH]` HTTP traffic redirects to HTTPS (not served in parallel) - [ ] TLS 1.0 and 1.1 disabled — only TLS 1.2+ allowed - [ ] Weak cipher suites disabled - [ ] Certificate expiry is monitored (NPM auto-renews, but verify it's configured) ### 3f. Docker & Container Security -- [ ] Containers do not run as root (check `user:` in Compose or Dockerfile `USER` instruction) -- [ ] No container has `privileged: true` unless there is a documented reason -- [ ] No unnecessary host volume mounts (especially `/var/run/docker.sock` unless intentional) +- [ ] `[HIGH]` Containers do not run as root (check `user:` in Compose or Dockerfile `USER` instruction) +- [ ] `[HIGH]` No container has `privileged: true` unless there is a documented reason +- [ ] `[HIGH]` No unnecessary host volume mounts (especially `/var/run/docker.sock` unless intentional) - [ ] Container images are not using `latest` tag in production -- [ ] Docker socket is not exposed to the external network +- [ ] `[CRITICAL]` Docker socket is not exposed to the external network - [ ] Resource limits (`mem_limit`, `cpus`) are set on containers -**AI Action:** Run the following tools if available: -- `bandit -r . -ll` — Python static security analysis -- `trivy fs . --severity HIGH,CRITICAL` — dependency and filesystem CVE scan -- `docker scout cves ` — container image vulnerability scan - -Report all FAIL/WARN findings. Do not proceed to go-live recommendation until critical issues are resolved. - ### 3g. Network & Exposure -- [ ] Only port 80/443 are exposed publicly — no app ports (e.g. 8000, 3000) directly open to internet +- [ ] `[CRITICAL]` Only port 80/443 are exposed publicly — no app ports (e.g. 8000, 3000) directly open to internet - [ ] NPM proxy host has access list or basic auth if the service is internal-only - [ ] Rate limiting is configured in Nginx or the app for API endpoints -- [ ] The service does not expose an admin panel (e.g. `/admin`, `/dashboard`) without additional auth -- [ ] Database ports (3306, 5432, 6379) are NOT exposed beyond the Docker network +- [ ] `[HIGH]` The service does not expose an admin panel (e.g. `/admin`, `/dashboard`) without additional auth +- [ ] `[HIGH]` API docs (`/docs`, `/redoc`, `/openapi.json` in FastAPI) are disabled or behind auth in production +- [ ] `[HIGH]` Debug mode is off (`debug=False`, no auto-reload, no verbose tracebacks returned to clients) +- [ ] `[CRITICAL]` Database ports (3306, 5432, 6379) are NOT exposed beyond the Docker network - [ ] SSH is not running inside any container -### 3h. Dependency & Supply Chain +### 3h. Database Access +- [ ] `[HIGH]` App connects with a dedicated MySQL user limited to its own database — not `root` +- [ ] App DB user has only the privileges it needs (no `GRANT`, `SUPER`, or global privileges) +- [ ] MySQL root password is separate from the app user's password and stored only in `.env` + +### 3i. Dependency & Supply Chain - [ ] Dependencies are pinned to specific versions (not `*` or `latest`) - [ ] Known CVEs in dependencies? (run `trivy fs .` or `pip-audit` / `npm audit`) - [ ] No abandoned or unmaintained packages with known issues - [ ] Docker base images are from official/verified sources -### 3i. Request & Session Hardening -- [ ] Session/auth cookies set `Secure`, `HttpOnly`, and `SameSite=Strict` or `Lax` -- [ ] CSRF protection is in place for state-changing (non-GET) endpoints, or the API is stateless/token-authenticated (not cookie-authenticated) and doesn't need it -- [ ] CORS policy restricts `Access-Control-Allow-Origin` to known origins — no wildcard `*` combined with `Access-Control-Allow-Credentials: true` +### 3j. Request & Session Hardening +- [ ] `[HIGH]` Session/auth cookies set `Secure`, `HttpOnly`, and `SameSite=Strict` or `Lax` +- [ ] `[HIGH]` CSRF protection is in place for state-changing (non-GET) endpoints, or the API is stateless/token-authenticated (not cookie-authenticated) and doesn't need it +- [ ] `[HIGH]` CORS policy restricts `Access-Control-Allow-Origin` to known origins — no wildcard `*` combined with `Access-Control-Allow-Credentials: true` - [ ] File uploads enforce a max size limit and restrict allowed file types/extensions -### 3j. Backups & Recovery -- [ ] Database/data volumes have an automated backup process (cron, script, or handled by the DB image) +### 3k. Backups & Recovery +- [ ] `[HIGH]` Database/data volumes have an automated backup process (cron, script, or handled by the DB image) - [ ] Backups are stored outside the container (separate volume, external disk, or remote target) so they survive the container/host being wiped - [ ] A restore from backup has been tested at least once, not just assumed to work - [ ] Backup files containing sensitive data are encrypted or access-restricted +- [ ] Live `.env` and any config that only exists in Portainer/Dockhand or NPM are backed up somewhere retrievable (they aren't in git) -### 3k. Logging & Observability +### 3l. Logging & Observability - [ ] Logs are rotated or size-capped (Docker `logging.driver`/`max-size`, or app-level rotation) so they can't fill disk over time - [ ] Logs persist outside the container (volume-mounted or shipped elsewhere) so they survive a container restart/crash - [ ] The app trusts `X-Forwarded-For`/`X-Real-IP` from the proxy correctly, so rate limiting and audit logs record the real client IP, not NPM's +- [ ] `[HIGH]` Logs never contain passwords, tokens, session IDs, API keys, or full request bodies with sensitive fields + +**AI Action:** Run the following tools if available: +- `bandit -r . -ll` — Python static security analysis +- `trivy fs . --severity HIGH,CRITICAL` — dependency and filesystem CVE scan +- `gitleaks git .` — secret scan across full git history (older gitleaks versions: `gitleaks detect --source .`) +- `docker scout cves ` or `trivy image ` — container image vulnerability scan + +Report all FAIL/WARN findings. Do not proceed to Section 4 until CRITICAL and HIGH issues are resolved. --- -## Section 4: Go-Live Decision +## Section 4: External Verification + +Goal: Confirm the live service matches what the config says. Sections 1–3 check files — this checks what the internet actually sees. + +Only start this section once Sections 1–3 have no unresolved CRITICAL/HIGH findings. Enable the NPM proxy host, then run these checks immediately. If any CRITICAL/HIGH check fails, disable the proxy host until it's fixed. + +- [ ] `[HIGH]` Security headers from 3d are present in the live response +- [ ] `[HIGH]` HTTP → HTTPS redirect works on the public URL +- [ ] `[MEDIUM]` SSL Labs test (https://www.ssllabs.com/ssltest/) scores A or better +- [ ] securityheaders.com (https://securityheaders.com) scores A or better, or gaps are documented +- [ ] `[CRITICAL]` External port scan shows only 80/443 open — run from outside the LAN (e.g. phone hotspot) or an online port scanner, not from inside the network +- [ ] `[HIGH]` `/docs`, `/redoc`, `/openapi.json`, and any admin paths return 401/403/404 from the public URL +- [ ] Error pages don't leak stack traces or server/framework versions (request a nonexistent path and a malformed request) +- [ ] Uptime Kuma monitor shows the service as up + +**AI Action:** Run: +- `curl -sI https://` — verify headers +- `curl -sI http://` — verify redirect +- `curl -s -o /dev/null -w "%{http_code}" https:///docs` (repeat for `/redoc`, `/openapi.json`, admin paths) +- `echo | openssl s_client -connect :443 -servername 2>/dev/null | openssl x509 -noout -dates -issuer` + +Ask the user to run the SSL Labs, securityheaders.com, and external port scan checks and paste the results. + +--- + +## Section 5: Go-Live Decision - [ ] Is a rollback or "what to do if this breaks/is compromised" step documented (README or runbook) for this service? @@ -190,3 +261,15 @@ After all sections are complete: - Failed (non-critical): X - Deferred: X - **Recommendation: GO / NO GO / GO WITH CONDITIONS** + +--- + +## Section 6: Baseline & Handoff + +Goal: Record what went live so the `service-update` maintenance playbook has something to compare against. Only complete after a GO or GO WITH CONDITIONS decision. + +- [ ] Go-live version tagged in Gitea (e.g. `v1.0.0`) and pushed +- [ ] `Release-Notes/v1.0.md` created +- [ ] Running image tags/digests recorded in the report +- [ ] Deferred items recorded in the report with the reason for deferring +- [ ] Report saved to `reports/golive-.md` in the project directory — the first maintenance review uses it as its prior report diff --git a/playbooks/service-update.md b/playbooks/service-update.md index d54a4a8..98b8160 100644 --- a/playbooks/service-update.md +++ b/playbooks/service-update.md @@ -13,7 +13,7 @@ When invoked, read the project directory in the current working directory. Some Tell the AI: _"Use the service-update playbook to review this service: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-update.md"_ The AI will: -1. Read the project files and the most recent prior maintenance report (if one exists) +1. Read the project files and the most recent prior report in `reports/` (`maintenance-*.md`, or `golive-*.md` if this is the first maintenance review) 2. Work through each section below 3. For each item — report PASS, FAIL, or WARN with specific findings 4. At the end, give an overall health rating and save a report