Files
homeschool/reports/maintenance-2026-09-24.md
T
derekcandClaude Opus 5.5 5a2510059d Fix timer action 500s after FastAPI upgrade
FastAPI 0.141 raises when response serialization hits a lazy-load
(current_block.subject.options) in async context, where 0.115 silently
skipped it. Eager-load Subject.options in session get/timer and the TV
dashboard queries. Pin pydantic 2.13.5.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
2026-09-24 23:51:10 -07:00

16 KiB
Raw Permalink Blame History

Maintenance Review — 2026-09-24

  • Service: Homeschool Dashboard (homeschool.chns.tech)
  • Host: docker-20-3 (app), NPM on Docker-22-1 (172.16.22.21) → docker-20-3:8054
  • Repo HEAD: 8e92ae6 (main, clean, in sync with origin). No git tags exist.
  • Prior report: none. This is the first maintenance review, and no go-live report exists.
  • Playbook: service-update.md

Rollback point (running images)

Container Image Image ID Built
homeschool_db mysql:8.0.40 sha256:6c55ddbef969… 2024-10-14
homeschool_backend homeschool-backend sha256:f72d140376b1… 2026-03-23 08:43
homeschool_frontend homeschool-frontend sha256:406c3c0faba3… 2026-03-23 08:29

App images are only tagged homeschool-*:latest (local build). A rebuild overwrites them, so re-tag them before any update: docker tag homeschool-backend homeschool-backend:rollback-20260924, and the same for the frontend.


Section 1: Baseline & Drift

Check Result Notes
Tree clean and pushed PASS
Running version matches main FAIL The backend code matches HEAD. The frontend bundle predates 8e92ae6: the live JS still has the old >=-15 meeting catch-up window, not >=-300. The fix from Mar 31 was never deployed.
Live stack matches repo compose USER Containers are named per the repo compose. Confirm no edits were made in the Portainer UI.
NPM proxy config documented USER README has an NPM section. Confirm the live proxy host matches it.
.env vs .env.example WARN All variables are set. DOCS_ENABLED is in both files but is not passed to the container in docker-compose.yml, so it has no effect. This fails safe (docs stay off), but README says setting it enables docs.
Rollback digests recorded PASS See table above. No release tags.

Section 2: Runtime Health

Check Result Notes
All containers healthy WARN Only db has a healthcheck. backend and frontend show only Up.
Restart count WARN backend RestartCount=4. It crash-loops about 4 times after every host reboot (09-06, 09-08, 09-13, 09-17): Can't connect to MySQL server on 'db' → startup failed. It recovers on its own. On reboot, the restart: policy ignores depends_on: service_healthy.
Public health endpoint PASS https://homeschool.chns.tech/api/health → 200
Logs (30 days) WARN Only the reboot crash loops above. 11× 429 on /ws/… (see the rate-limit finding in 5b). ~2.7k 404/405 from WordPress scanners (harmless).
CPU/memory headroom WARN db 421 MiB / 512 MiB (82%). backend 14%, frontend 6%.
Disk space PASS / 45% used (51 GB free)
Log rotation FAIL No logging: in compose and no /etc/docker/daemon.json. json-file logs grow without limit.
Reclaimable Docker space WARN Host-wide: 5.1 GB build cache and 6.6 GB unused images can be reclaimed.

Section 3: Functional Check (user to confirm)

The app has no scheduled/cron jobs. Its background work is WebSocket broadcasts and startup migrations only.

  • Login and logout. The access token expires after 30 minutes and should auto-refresh through the cookie.
  • Dashboard: start a day session, then start, pause, resume and complete a block.
  • TV view /tv/<token> updates live over WebSocket.
  • Schedules: create/edit a template. Admin: add a child or subject.
  • Logs view loads.
  • Super admin login at /super-admin/login → user list. The ntfy alert arrives and shows the real client IP.
  • DB connectivity: PASS. The only DB errors are the startup race on reboot.

Section 4: Backups & Recovery — CRITICAL

Check Result Notes
Recent DB backup FAIL (CRITICAL) No backup job found: no user crontab, no cron.d entry, no systemd timer, no backup container, no dump files. The only copy of the data is the homeschool_mysql_data volume on this host. USER: confirm whether VM-level snapshots or backups exist outside this host.
Backup size / integrity FAIL No backups exist.
Stored off-host FAIL
Upload volumes N/A No upload volumes. The DB is the only state.
.env backed up USER
Retention FAIL
Test restore in last 6 months FAIL

Section 5: Security Posture

5a. TLS & Headers

Check Result Notes
Cert valid >14 days PASS Let's Encrypt, expires 2026-12-19 (86 days). Confirm auto-renew is on in NPM.
HTTP → HTTPS redirect USER Port 80 timed out from inside the LAN (hairpin/split DNS). Run curl -sI http://homeschool.chns.tech from outside.
Security headers PASS XFO, XCTO, Referrer-Policy, CSP and HSTS are all present.
Server version suppressed PASS server: openresty (no version).

5b. Exposure

Check Result Notes
Only 80/443 public WARN 8054 is published on 0.0.0.0 of docker-20-3 (needed for the NPM host). Confirm a firewall limits it to 172.16.22.21 so the LAN can't bypass TLS. DB/backend ports are not published.
Rate limiting per client FAIL (HIGH) Frontend nginx sees every request as coming from the NPM host (20,450 of 20,451 log lines = 172.16.22.21). There is no real_ip config, so limit_req keyed on $binary_remote_addr is global. The login limit (5/min) applies to all users combined, and one brute-forcer can lock everyone out. The TV/WS limit (10/min) is shared too, which caused today's 429s on /ws/128699. Fix: add set_real_ip_from 172.16.22.21; real_ip_header X-Real-IP; (or X-Forwarded-For).
Spoofable client IP in alerts WARN auth.py:62,86 and admin.py:26 take the first X-Forwarded-For entry, which the client controls. An attacker can forge the IP shown in ntfy alerts. Use the entry added by the trusted proxy (the last one) or X-Real-IP.
No unreviewed containers/ports PASS 3 services, same as compose.

5c. Accounts & Secrets

Check Result Notes
Account review USER The DB query was not run: the tool permission policy blocked it. Run the query below.
Secret rotation WARN .env last modified 2026-03-22 (about 6 months). Secrets have never been rotated since then.
Failed-login spikes PASS 0 auth rate-limit hits. 54× 401 are almost all /api/auth/refresh (expired sessions). But the per-IP limits are broken (5b), so this signal is unreliable.
docker exec -it homeschool_db sh -c 'mysql -u"$MYSQL_USER" -p"$MYSQL_PASSWORD" "$MYSQL_DATABASE" -e "SELECT id, email, is_active, created_at FROM users ORDER BY id;"'

5d. Container Hardening

Check Result Notes
Non-root / no privileged / no sock PASS / WARN backend runs as appuser. The frontend nginx master runs as root (stock image; consider nginxinc/nginx-unprivileged). None are privileged, and there is no docker.sock mount.
Resource limits PASS mem_limit and cpus are set on all three.

Section 6: Vulnerabilities & Currency

Scanned with trivy image, HIGH/CRITICAL only.

Image CRIT HIGH Key fixable items
homeschool-backend (OS, debian 13.4) 19 716 openssl/libssl3 3.5.5-1~deb13u1 → deb13u2+, perl-base, gzip, util-linux, linux-libc-dev. Rebuild with an updated base and apt-get upgrade. Drop libmariadb-dev/gcc from runtime (aiomysql is pure Python). That removes the no-fix mariadb CRITs.
homeschool-backend (Python) 1 11 anyio 4.12.1→4.14.2 (CRIT), starlette 0.38.6 (via fastapi 0.115.0), PyJWT 2.12.0→2.13.0, cryptography 46.0.5→48.0.1+, python-multipart 0.0.22→0.0.27+, Mako 1.3.10→1.3.12 (via alembic)
homeschool-frontend (alpine 3.23.3) 2 50 libssl3/libcrypto3 3.5.5→3.5.8, curl, libexpat, musl, zlib. Rebuild on the latest nginx alpine.
mysql:8.0.40 0 130 Many el9 package fixes
Check Result Notes
Image CVEs FAIL (HIGH) See above. A rebuild fixes most of them.
App dependency CVEs FAIL (HIGH) Python list above. Frontend: no package-lock.json, so npm audit can't run and each build resolves transitive deps fresh. Known: axios 1.7.7 has advisories fixed in ≥1.12. vite 5.4.10 is build-only.
Runtimes EOL FAIL MySQL 8.0 reached EOL in April 2026. Move to 8.4 LTS. Node 20 reached EOL on 2026-04-30 (build stage only). Move to Node 22/24 LTS. Python 3.12 is supported. nginx 1.29 mainline is superseded by 1.30.x stable.
Abandoned deps WARN passlib 1.7.4 is unmaintained (last release 2020) and forces bcrypt==3.2.2. Consider using bcrypt directly.
Pinning PASS / WARN Images and top-level deps are pinned. There is no npm lockfile (above).

Section 7: Applying Updates — HIGH findings resolved

This section was first blocked on backups. The user then asked for the HIGH findings to be fixed the same day, so a one-off pre-update backup was taken.

Check Result Notes
Fresh backup before update PASS mysqldump → ~/backups/homeschool/homeschool-pre-mysql84-20260924-2327.sql.gz (14 tables). Volume copy → homeschool_mysql_data_pre84_20260924 (187 MB). Both are on-host only.
Rollback point recorded PASS homeschool-backend:rollback-20260924, homeschool-frontend:rollback-20260924, plus the table at the top
Major-version release notes PASS MySQL 8.0 → 8.4 in-place upgrade is supported. The app user already uses caching_sha2.
Pinned versions PASS All bumped versions are pinned exactly.
Tested against a copy of real data PASS Dump restored into a throwaway mysql:8.4.11. The new backend was smoke-tested there: register, login, children, subjects, schedules, logs, super admin, forged/alg=none JWT rejected, refresh. Caught a regression: unpinned PyMySQL floated to 1.2.3 and broke SQLAlchemy 2.0.35 ping. It is now pinned to 1.1.2.
Real-IP rate limiting tested PASS Test nginx: client A was limited after its burst, client B was unaffected, and a forged XFF prefix resolved to the true client.
Deploy PASS* The MySQL data upgrade (80040 → 80411) took about 90s, longer than the healthcheck window, so compose aborted the dependent containers. The user started them manually. Fixed by adding start_period: 180s to the db healthcheck.
Sections 2/3 re-run PASS (partial) All containers up, backend logs clean, public /api/health 200. Live bundle now contains 8e92ae6. nginx logs real client IPs. The user confirmed the site is back up.
Version tag / release notes PASS v1.1.0, Release-Notes/v1.1.md
Old images cleanup DEFERRED Keep the rollback tags and volume copy until the update has been stable for about a week. Then remove them: docker rmi homeschool-{backend,frontend}:rollback-20260924 mysql:8.0.40 and docker volume rm homeschool_mysql_data_pre84_20260924.

Changes applied

  • frontend/nginx.conf: set_real_ip_from the NPM host and Cloudflare ranges, real_ip_header X-Forwarded-For, real_ip_recursive on
  • backend/requirements.txt:
    • fastapi 0.115.0 → 0.141.1, starlette → 1.7.0 (pinned)
    • anyio 4.15.1 (pinned), PyJWT 2.12.0 → 2.15.0, cryptography 46.0.5 → 50.0.1
    • python-multipart 0.0.22 → 0.0.32, alembic 1.13.3 → 1.20.0, Mako 1.4.3 (pinned), PyMySQL 1.1.2 (pinned)
  • backend/Dockerfile: python 3.12.13 → 3.12.14-slim. Dropped gcc / libmysqlclient-dev / pkg-config (not needed). Added apt-get upgrade.
  • frontend/Dockerfile: nginx 1.29.6 → 1.30.5-alpine (stable) plus apk upgrade
  • docker-compose.yml: mysql 8.0.40 → 8.4.11. db healthcheck start_period: 180s.

New running images

Container Image Image ID
homeschool_db mysql:8.4.11 sha256:ee241324a55f…
homeschool_backend homeschool-backend sha256:51b49afb4ec4…
homeschool_frontend homeschool-frontend sha256:f9a737699cda…

Post-deploy regression

After the HIGH deploy, POST /api/sessions/{id}/timer returned 500 on every call: 8 failures, 0 successes, starting 06:43 UTC. FastAPI 0.141 raises on a failed lazy-load of current_block.subject.options during response serialization, where 0.115 silently skipped it. The first smoke test didn't cover timer actions. Fixed in v1.1.0 by eager-loading Subject.options in sessions.py and dashboard.py. The smoke test now covers the full timer lifecycle, the TV dashboard with an active block, schedule/subject writes, and every parameterless GET route (from the OpenAPI spec).

Post-update scan (trivy HIGH/CRITICAL)

Image Before After
backend OS 19 C / 716 H 0 C / 44 H (no fixes available)
backend Python 1 C / 11 H 0
frontend 2 C / 50 H 0
mysql 0 C / 130 H 8.4.11: 4 H curl (OS), 6 H in bundled Python tools, 1 C / 21 H Go stdlib in gosu. All are upstream in Oracle's latest image and accepted until the next mysql 8.4.x release.

Section 8: Documentation

Check Result Notes
README accurate WARN DOCS_ENABLED is documented but not wired into compose (Section 1).
Rollback runbook FAIL README has "Stopping and Restarting" but no backup, restore or rollback procedure.
Deferred items recorded PASS This report.

Section 9: Summary

CRITICAL

  1. No database backups. No backup job, no dump files, and nothing stored off-host has been found.

HIGH (all resolved 2026-09-24; see Section 7)

  1. Rate limiting is global, not per client. Fixed with real_ip.
  2. Fixable HIGH/CRITICAL CVEs in all images and Python deps. Fixed by the rebuild and dependency bumps.
  3. MySQL 8.0 is EOL. Upgraded to 8.4.11 LTS.

MEDIUM (all resolved 2026-09-24, v1.1.0)

  1. The frontend is not deployed at HEAD. Resolved by the rebuild.
  2. No Docker log rotation. Added x-logging anchor (json-file, 10m × 3) to all services.
  3. The backend crash-loops on reboot, and there are no backend/frontend healthchecks. _wait_for_db() in main.py retries for up to 180s. Tested: the backend started before its DB, logged retries, and came up with 0 restarts. Healthchecks were added to both services.
  4. Node 20 EOL and no lockfile. Now node:24.21.0-alpine with package-lock.json and npm ci. axios → 1.20.0, vite → 5.4.21. Remaining npm audit items (vite ≤6.4.2 / esbuild) are dev-server-only and accepted; the fix needs Vite 6.4+.
  5. No rollback/restore runbook or release tags. README now has "Backup, Restore & Rollback" and documents real-IP proxy trust. Added Release-Notes/v1.1.md and tagged v1.1.0.
  6. Client IP in ntfy alerts can be spoofed. New app/utils/client_ip.py reads X-Real-IP, which nginx sets after real_ip resolution. Tested: a forged XFF is ignored.

LOW

  1. db memory is at 82% of its 512 MiB limit.
  2. DOCS_ENABLED is not passed to the container (README mismatch).
  3. passlib is unmaintained. The frontend nginx runs as root.
  4. Secrets are about 6 months old and have never been rotated.
  5. About 11 GB of Docker space can be reclaimed host-wide.

Pending user input

Portainer/NPM drift, functional smoke tests, ntfy test, off-host/VM backups, HTTP→HTTPS redirect from outside, firewall on 8054, NPM auto-renew, and the user account review.

Counts

  • Total checks: 58
  • Passed: 19
  • Failed (critical/high): 7. That is backups (counted once, CRITICAL), rate limiting, image CVEs, dependency CVEs, runtime EOL, log rotation and frontend drift.
  • Failed (non-critical) / WARN: 18
  • Pending user confirmation: 14
  • Deferred: 1 (old image cleanup)
  • Updates applied: 9 (HIGH: real-IP rate limiting, Python deps, base images, MySQL 8.4. MEDIUM: log rotation, DB-wait + healthchecks, Node 24 + lockfile, runbook + release tag, trusted client IP)

Overall: NEEDS ATTENTION. All HIGH and MEDIUM findings are resolved. The CRITICAL no-backups finding is still open: the pre-update dump and volume copy are on-host only, with no schedule.