276 lines
16 KiB
Markdown
276 lines
16 KiB
Markdown
# Playbook: Service Go-Live Review
|
||
|
||
Use this playbook before exposing any service to external access through Nginx Proxy Manager (NPM).
|
||
When invoked, read the project directory in the current working directory and work through each section as an interactive checklist.
|
||
|
||
---
|
||
|
||
## How to Use
|
||
|
||
Tell the AI: _"Use the service-golive playbook to review this project: https://git.chns.tech/CHNS/AI/raw/branch/main/playbooks/service-golive.md"_
|
||
|
||
The AI will:
|
||
1. Read the project files in the current directory
|
||
2. Work through each section below
|
||
3. For each item — report PASS, FAIL, or WARN with specific findings
|
||
4. At the end, give a go/no-go recommendation and save a report
|
||
|
||
Do not proceed to the next section until the current one is resolved or explicitly deferred.
|
||
|
||
**Severity:** Items tagged `[CRITICAL]`, `[HIGH]`, `[MEDIUM]`, or `[LOW]` use that severity when they FAIL. For untagged items, the AI assigns severity based on how exploitable or disruptive the gap is and states its reasoning.
|
||
|
||
---
|
||
|
||
## Section 0: Pre-flight
|
||
|
||
Goal: Make sure the review tools are available before relying on them.
|
||
|
||
- [ ] `trivy --version`
|
||
- [ ] `bandit --version` (Python projects)
|
||
- [ ] `gitleaks version`
|
||
- [ ] `curl` and `openssl` available
|
||
|
||
**AI Action:** If a tool is missing, tell the user and ask whether to install it or skip the checks that depend on it. Skipped checks are reported as WARN, not PASS.
|
||
|
||
---
|
||
|
||
## Section 1: Feature & Improvement Review
|
||
|
||
Goal: Catch missing functionality before users find it.
|
||
|
||
- [ ] Does the service have a health check endpoint (e.g. `/health` or `/ping`)?
|
||
- [ ] Does the app container have its own Compose `healthcheck` wired to that endpoint (not just the database)?
|
||
- [ ] Is `restart: unless-stopped` (or `always`) set on every service?
|
||
- [ ] Are all intended routes/endpoints implemented and reachable?
|
||
- [ ] Is there a meaningful error response for bad input (not raw stack traces)?
|
||
- [ ] Are there any obvious UX gaps or incomplete flows in the UI (if applicable)?
|
||
- [ ] Is there logging in place to capture errors and key events?
|
||
- [ ] Are there any TODO/FIXME/HACK comments in the code that indicate unfinished work?
|
||
- [ ] Does the service handle its own startup failures gracefully (exits cleanly, logs reason)?
|
||
|
||
### 1a. Ntfy Admin Notifications
|
||
|
||
Goal: Ensure the super admin is alerted to significant events without having to monitor logs manually. Ntfy is required for all web apps (see `preferences/python-web.md`).
|
||
|
||
- [ ] `[HIGH]` Is Ntfy integrated into the application?
|
||
- [ ] Are admin-relevant events triggering Ntfy notifications?
|
||
- [ ] `[HIGH]` Is the Ntfy topic access-protected (token or user auth on the self-hosted instance) so others can't subscribe to admin alerts?
|
||
- [ ] Are Ntfy calls wrapped in `try/except` so a failed notification can't crash the app?
|
||
- [ ] Ntfy URL, topic, and token come from `.env` — not hardcoded
|
||
|
||
**If Ntfy is NOT implemented**, flag as FAIL and recommend the following events for notification coverage based on what the app does:
|
||
|
||
| Event | Severity | Why it matters |
|
||
|---|---|---|
|
||
| Successful admin login | High | Detect unauthorized admin access |
|
||
| Failed admin login (threshold reached) | High | Brute-force indicator |
|
||
| New user registration | Medium | Visibility into who is joining |
|
||
| User account deletion | Medium | Audit trail for removals |
|
||
| Role/permission escalation | High | Privilege change could indicate compromise |
|
||
| Password reset requested | Medium | Could indicate account takeover attempt |
|
||
| Rate limit triggered | Medium | Abuse or misconfigured client |
|
||
| API key created or revoked | High | Credential lifecycle event |
|
||
| Service startup / crash recovery | Medium | Unexpected restarts need awareness |
|
||
| High error rate (e.g. 5xx spike) | High | App health degrading in production |
|
||
| Large data export initiated | Medium | Data exfiltration risk indicator |
|
||
| Config or environment change detected | High | Unplanned changes should be visible |
|
||
|
||
**AI Action:** Search the codebase for Ntfy integration (look for `ntfy`, `ntfy.sh`, or HTTP POST calls to a notification endpoint). If none found, list the above recommended events as FAIL items and ask the user whether to implement before go-live.
|
||
|
||
### 1b. Uptime Kuma Monitoring
|
||
|
||
Goal: Get alerted when the service is down — Ntfy runs inside the app, so it can't report its own outage.
|
||
|
||
- [ ] `[HIGH]` An Uptime Kuma monitor exists for this service, pointed at the public health endpoint URL (through NPM), not the internal container address
|
||
- [ ] The monitor has a notification configured (e.g. Ntfy) so downtime actually alerts someone
|
||
- [ ] Monitor uses an HTTP(s) status or keyword check, not just a ping/TCP check
|
||
- [ ] Certificate expiry notification is enabled on the monitor
|
||
|
||
**AI Action:** Ask the user to confirm the monitor is set up — the AI can't see Uptime Kuma directly.
|
||
|
||
---
|
||
|
||
**AI Action:** List any gaps found with file and line references. Ask the user whether to fix now or defer.
|
||
|
||
---
|
||
|
||
## Section 2: Performance Review
|
||
|
||
Goal: Ensure the service won't collapse under real load.
|
||
|
||
- [ ] Are database queries using indexes on columns used in WHERE/JOIN/ORDER BY clauses?
|
||
- [ ] Are N+1 query patterns present (loop that fires a query per item)?
|
||
- [ ] Is connection pooling configured for the database?
|
||
- [ ] Are large responses paginated?
|
||
- [ ] Are any blocking operations (file I/O, external API calls) being done synchronously in an async context?
|
||
- [ ] Are static assets (if any) being served through Nginx, not the app?
|
||
- [ ] Is there any unbounded data being loaded into memory (e.g. `SELECT *` with no limit)?
|
||
- [ ] Are background tasks or scheduled jobs using a proper queue/worker model (not threading hacks)?
|
||
- [ ] Is Gzip/Brotli compression enabled in Nginx for text responses?
|
||
|
||
**AI Action:** Flag any issues with specific file references. Suggest fixes. Ask user to confirm or defer.
|
||
|
||
---
|
||
|
||
## Section 3: Security Audit
|
||
|
||
Goal: Do not put a vulnerable service on the internet. Be thorough.
|
||
|
||
### 3a. Secrets & Credentials
|
||
- [ ] `[CRITICAL]` No hardcoded passwords, tokens, API keys, or secrets in any source file
|
||
- [ ] `[CRITICAL]` `.env` file is in `.gitignore` and not committed
|
||
- [ ] `[CRITICAL]` No secrets anywhere in git history (a previously committed `.env` is still exposed even if deleted later — rotate any secret found)
|
||
- [ ] `[HIGH]` `.dockerignore` exists and excludes `.env`, `.git`, and other local-only files so they aren't baked into the image
|
||
- [ ] `.env.example` exists with placeholder values only
|
||
- [ ] `[HIGH]` No secrets in Docker Compose files (use `env_file` or environment variable references, not literal values)
|
||
- [ ] `[HIGH]` No secrets in Nginx config files
|
||
|
||
### 3b. Authentication & Authorization
|
||
- [ ] `[CRITICAL]` All non-public endpoints require authentication
|
||
- [ ] Authentication tokens/sessions have an expiry
|
||
- [ ] `[HIGH]` Password hashing uses bcrypt, argon2, or scrypt — not MD5/SHA1
|
||
- [ ] `[CRITICAL]` There is no default admin password that ships with the service
|
||
- [ ] Role/permission checks exist if the app has multiple access levels
|
||
- [ ] `[HIGH]` Failed login attempts are rate-limited or account-locked after N failures
|
||
- [ ] Login and password-reset responses don't reveal whether an account exists (same message and similar timing for valid/invalid users)
|
||
- [ ] Logout and password change invalidate existing sessions/tokens
|
||
|
||
### 3c. Input Validation & Injection
|
||
- [ ] All user input is validated server-side (not just client-side)
|
||
- [ ] `[CRITICAL]` SQL queries use parameterized statements or ORM — no string concatenation
|
||
- [ ] `[HIGH]` File upload paths are sanitized — no path traversal possible
|
||
- [ ] `[HIGH]` HTML output is escaped to prevent XSS (or a framework handles this automatically)
|
||
- [ ] Redirects only go to allowed/relative URLs — no open redirect
|
||
- [ ] `[HIGH]` JSON deserialization does not allow arbitrary object instantiation
|
||
|
||
### 3d. HTTP & Nginx Security Headers
|
||
Verify the Nginx config for the proxy host includes:
|
||
- [ ] `X-Frame-Options: DENY` or `SAMEORIGIN`
|
||
- [ ] `X-Content-Type-Options: nosniff`
|
||
- [ ] `X-XSS-Protection: 0` or omitted — the legacy `1; mode=block` value is deprecated and can introduce vulnerabilities; CSP replaces it
|
||
- [ ] `Referrer-Policy: strict-origin-when-cross-origin`
|
||
- [ ] `Content-Security-Policy` header defined (even if broad to start)
|
||
- [ ] `Permissions-Policy` restricts browser features the app doesn't use (e.g. `camera=(), microphone=(), geolocation=()`)
|
||
- [ ] `[MEDIUM]` `Strict-Transport-Security` (HSTS) with `max-age` >= 31536000
|
||
- [ ] Server version header suppressed (`server_tokens off`)
|
||
- [ ] Unnecessary HTTP methods disabled (e.g. TRACE, DELETE if not used)
|
||
|
||
### 3e. TLS / HTTPS
|
||
- [ ] `[CRITICAL]` TLS certificate is valid and not self-signed for production
|
||
- [ ] `[HIGH]` HTTP traffic redirects to HTTPS (not served in parallel)
|
||
- [ ] TLS 1.0 and 1.1 disabled — only TLS 1.2+ allowed
|
||
- [ ] Weak cipher suites disabled
|
||
- [ ] Certificate expiry is monitored (NPM auto-renews, but verify it's configured)
|
||
|
||
### 3f. Docker & Container Security
|
||
- [ ] `[HIGH]` Containers do not run as root (check `user:` in Compose or Dockerfile `USER` instruction)
|
||
- [ ] `[HIGH]` No container has `privileged: true` unless there is a documented reason
|
||
- [ ] `[HIGH]` No unnecessary host volume mounts (especially `/var/run/docker.sock` unless intentional)
|
||
- [ ] Container images are not using `latest` tag in production
|
||
- [ ] `[CRITICAL]` Docker socket is not exposed to the external network
|
||
- [ ] Resource limits (`mem_limit`, `cpus`) are set on containers
|
||
|
||
### 3g. Network & Exposure
|
||
- [ ] `[CRITICAL]` Only port 80/443 are exposed publicly — no app ports (e.g. 8000, 3000) directly open to internet
|
||
- [ ] NPM proxy host has access list or basic auth if the service is internal-only
|
||
- [ ] Rate limiting is configured in Nginx or the app for API endpoints
|
||
- [ ] `[HIGH]` The service does not expose an admin panel (e.g. `/admin`, `/dashboard`) without additional auth
|
||
- [ ] `[HIGH]` API docs (`/docs`, `/redoc`, `/openapi.json` in FastAPI) are disabled or behind auth in production
|
||
- [ ] `[HIGH]` Debug mode is off (`debug=False`, no auto-reload, no verbose tracebacks returned to clients)
|
||
- [ ] `[CRITICAL]` Database ports (3306, 5432, 6379) are NOT exposed beyond the Docker network
|
||
- [ ] SSH is not running inside any container
|
||
|
||
### 3h. Database Access
|
||
- [ ] `[HIGH]` App connects with a dedicated MySQL user limited to its own database — not `root`
|
||
- [ ] App DB user has only the privileges it needs (no `GRANT`, `SUPER`, or global privileges)
|
||
- [ ] MySQL root password is separate from the app user's password and stored only in `.env`
|
||
|
||
### 3i. Dependency & Supply Chain
|
||
- [ ] Dependencies are pinned to specific versions (not `*` or `latest`)
|
||
- [ ] Known CVEs in dependencies? (run `trivy fs .` or `pip-audit` / `npm audit`)
|
||
- [ ] No abandoned or unmaintained packages with known issues
|
||
- [ ] Docker base images are from official/verified sources
|
||
|
||
### 3j. Request & Session Hardening
|
||
- [ ] `[HIGH]` Session/auth cookies set `Secure`, `HttpOnly`, and `SameSite=Strict` or `Lax`
|
||
- [ ] `[HIGH]` CSRF protection is in place for state-changing (non-GET) endpoints, or the API is stateless/token-authenticated (not cookie-authenticated) and doesn't need it
|
||
- [ ] `[HIGH]` CORS policy restricts `Access-Control-Allow-Origin` to known origins — no wildcard `*` combined with `Access-Control-Allow-Credentials: true`
|
||
- [ ] File uploads enforce a max size limit and restrict allowed file types/extensions
|
||
|
||
### 3k. Backups & Recovery
|
||
- [ ] `[HIGH]` Database/data volumes have an automated backup process (cron, script, or handled by the DB image)
|
||
- [ ] Backups are stored outside the container (separate volume, external disk, or remote target) so they survive the container/host being wiped
|
||
- [ ] A restore from backup has been tested at least once, not just assumed to work
|
||
- [ ] Backup files containing sensitive data are encrypted or access-restricted
|
||
- [ ] Live `.env` and any config that only exists in Portainer/Dockhand or NPM are backed up somewhere retrievable (they aren't in git)
|
||
|
||
### 3l. Logging & Observability
|
||
- [ ] Logs are rotated or size-capped (Docker `logging.driver`/`max-size`, or app-level rotation) so they can't fill disk over time
|
||
- [ ] Logs persist outside the container (volume-mounted or shipped elsewhere) so they survive a container restart/crash
|
||
- [ ] The app trusts `X-Forwarded-For`/`X-Real-IP` from the proxy correctly, so rate limiting and audit logs record the real client IP, not NPM's
|
||
- [ ] `[HIGH]` Logs never contain passwords, tokens, session IDs, API keys, or full request bodies with sensitive fields
|
||
|
||
**AI Action:** Run the following tools if available:
|
||
- `bandit -r . -ll` — Python static security analysis
|
||
- `trivy fs . --severity HIGH,CRITICAL` — dependency and filesystem CVE scan
|
||
- `gitleaks git .` — secret scan across full git history (older gitleaks versions: `gitleaks detect --source .`)
|
||
- `docker scout cves <image>` or `trivy image <image>` — container image vulnerability scan
|
||
|
||
Report all FAIL/WARN findings. Do not proceed to Section 4 until CRITICAL and HIGH issues are resolved.
|
||
|
||
---
|
||
|
||
## Section 4: External Verification
|
||
|
||
Goal: Confirm the live service matches what the config says. Sections 1–3 check files — this checks what the internet actually sees.
|
||
|
||
Only start this section once Sections 1–3 have no unresolved CRITICAL/HIGH findings. Enable the NPM proxy host, then run these checks immediately. If any CRITICAL/HIGH check fails, disable the proxy host until it's fixed.
|
||
|
||
- [ ] `[HIGH]` Security headers from 3d are present in the live response
|
||
- [ ] `[HIGH]` HTTP → HTTPS redirect works on the public URL
|
||
- [ ] `[MEDIUM]` SSL Labs test (https://www.ssllabs.com/ssltest/) scores A or better
|
||
- [ ] securityheaders.com (https://securityheaders.com) scores A or better, or gaps are documented
|
||
- [ ] `[CRITICAL]` External port scan shows only 80/443 open — run from outside the LAN (e.g. phone hotspot) or an online port scanner, not from inside the network
|
||
- [ ] `[HIGH]` `/docs`, `/redoc`, `/openapi.json`, and any admin paths return 401/403/404 from the public URL
|
||
- [ ] Error pages don't leak stack traces or server/framework versions (request a nonexistent path and a malformed request)
|
||
- [ ] Uptime Kuma monitor shows the service as up
|
||
|
||
**AI Action:** Run:
|
||
- `curl -sI https://<public-url>` — verify headers
|
||
- `curl -sI http://<public-url>` — verify redirect
|
||
- `curl -s -o /dev/null -w "%{http_code}" https://<public-url>/docs` (repeat for `/redoc`, `/openapi.json`, admin paths)
|
||
- `echo | openssl s_client -connect <host>:443 -servername <host> 2>/dev/null | openssl x509 -noout -dates -issuer`
|
||
|
||
Ask the user to run the SSL Labs, securityheaders.com, and external port scan checks and paste the results.
|
||
|
||
---
|
||
|
||
## Section 5: Go-Live Decision
|
||
|
||
- [ ] Is a rollback or "what to do if this breaks/is compromised" step documented (README or runbook) for this service?
|
||
|
||
After all sections are complete:
|
||
|
||
- List all unresolved FINDs grouped by severity: **CRITICAL / HIGH / MEDIUM / LOW**
|
||
- **CRITICAL or HIGH unresolved = NO GO.** These must be fixed before external access.
|
||
- **MEDIUM/LOW unresolved** = user decides whether to defer with documented acceptance
|
||
- Provide a final summary:
|
||
- Total checks: X
|
||
- Passed: X
|
||
- Failed (critical): X
|
||
- Failed (non-critical): X
|
||
- Deferred: X
|
||
- **Recommendation: GO / NO GO / GO WITH CONDITIONS**
|
||
|
||
---
|
||
|
||
## Section 6: Baseline & Handoff
|
||
|
||
Goal: Record what went live so the `service-update` maintenance playbook has something to compare against. Only complete after a GO or GO WITH CONDITIONS decision.
|
||
|
||
- [ ] Go-live version tagged in Gitea (e.g. `v1.0.0`) and pushed
|
||
- [ ] `Release-Notes/v1.0.md` created
|
||
- [ ] Running image tags/digests recorded in the report
|
||
- [ ] Deferred items recorded in the report with the reason for deferring
|
||
- [ ] Report saved to `reports/golive-<YYYY-MM-DD>.md` in the project directory — the first maintenance review uses it as its prior report
|