Fix MEDIUM findings from 2026-09-24 maintenance review

- Backend retries DB connection at startup (up to 180s) so host reboots
  no longer crash-loop it; add backend and frontend healthchecks
- Docker log rotation (json-file 10m x 3) on all services
- ntfy alerts use X-Real-IP (set by nginx after real_ip resolution)
  instead of the client-controlled first X-Forwarded-For entry
- Frontend build on Node 24 LTS with package-lock.json + npm ci;
  axios 1.20.0, vite 5.4.21
- README: backup/restore/rollback runbook, real-IP proxy trust notes
- Release-Notes/v1.1.md; version 1.1.0

Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
derekcandClaude Opus 5.5 committed 2026-09-24 23:44:25 -07:00
1 parent 3170c7f4eb
commit 4130467b22
11 files changed
+1851 -20

No files matched your search

+71 -1
View File
@@ -49,7 +49,7 @@ A self-hosted web app for managing homeschool schedules, tracking daily learning
| Frontend server | nginx (Docker) | | Frontend server | nginx (Docker) |
| Backend API | FastAPI (Python 3.12) | | Backend API | FastAPI (Python 3.12) |
| Real-time | WebSockets via FastAPI | | Real-time | WebSockets via FastAPI |
| Database | MySQL 8 | | Database | MySQL 8.4 LTS |
| ORM | SQLAlchemy 2.0 (async) | | ORM | SQLAlchemy 2.0 (async) |
| Auth | JWT — PyJWT + passlib/bcrypt | | Auth | JWT — PyJWT + passlib/bcrypt |
| Orchestration | Docker Compose | | Orchestration | Docker Compose |
@@ -374,6 +374,8 @@ Nginx enforces the following rate limits (per IP):
Requests over the limit receive a `429 Too Many Requests` response. Requests over the limit receive a `429 Too Many Requests` response.
Limits are keyed on the real client IP. `frontend/nginx.conf` trusts the reverse proxy and Cloudflare (`set_real_ip_from`) and walks `X-Forwarded-For` past them. **If the NPM host IP changes, or you front the app with a different proxy/CDN, update those `set_real_ip_from` lines.** Otherwise every request appears to come from the proxy and all clients share one limit. The backend reads the resolved IP from `X-Real-IP` for ntfy alerts.
### Login Lockout ### Login Lockout
After 5 consecutive failed login attempts, a parent account is locked for 15 minutes. The lock clears automatically after the cooldown, or immediately when a super admin resets the user's password. The lockout threshold and duration are configured in `backend/app/routers/auth.py` (`_LOGIN_MAX_ATTEMPTS`, `_LOGIN_LOCKOUT_MINUTES`). After 5 consecutive failed login attempts, a parent account is locked for 15 minutes. The lock clears automatically after the cooldown, or immediately when a super admin resets the user's password. The lockout threshold and duration are configured in `backend/app/routers/auth.py` (`_LOGIN_MAX_ATTEMPTS`, `_LOGIN_LOCKOUT_MINUTES`).
@@ -411,3 +413,71 @@ docker compose down -v
# Restart without rebuilding # Restart without rebuilding
docker compose up docker compose up
``` ```
---
## Backup, Restore & Rollback
All persistent state is in the MySQL volume (`homeschool_mysql_data`). The live `.env` is not in git, so keep a copy somewhere safe.
### Back up the database
```bash
mkdir -p ~/backups/homeschool
docker exec homeschool_db sh -c 'exec mysqldump -uroot -p"$MYSQL_ROOT_PASSWORD" \
--single-transaction --routines --triggers --events --databases "$MYSQL_DATABASE"' \
| gzip > ~/backups/homeschool/homeschool-$(date +%Y%m%d-%H%M).sql.gz
```
Copy the dump **off the Docker host**. A backup on the same disk doesn't survive a host failure.
### Restore the database from a dump
```bash
docker compose stop backend frontend
zcat ~/backups/homeschool/<file>.sql.gz | \
docker exec -i homeschool_db sh -c 'exec mysql -uroot -p"$MYSQL_ROOT_PASSWORD"'
docker compose start backend frontend
```
The dump includes `CREATE DATABASE`/`DROP TABLE` statements, so it replaces the current tables.
### Before an update: record a rollback point
```bash
# App images are built locally and tagged :latest, so a rebuild overwrites them
docker tag homeschool-backend homeschool-backend:rollback-$(date +%Y%m%d)
docker tag homeschool-frontend homeschool-frontend:rollback-$(date +%Y%m%d)
# Take a database backup (above). For a MySQL version upgrade, also copy the volume,
# because MySQL can't be downgraded in place:
docker compose stop
docker volume create homeschool_mysql_data_pre_<date>
docker run --rm -v homeschool_mysql_data:/from:ro -v homeschool_mysql_data_pre_<date>:/to \
alpine sh -c 'cp -a /from/. /to/'
```
### Roll back an app update
```bash
docker tag homeschool-backend:rollback-<date> homeschool-backend:latest
docker tag homeschool-frontend:rollback-<date> homeschool-frontend:latest
docker compose up -d --no-build
```
If the update changed the schema, restore the pre-update dump too (above).
### Roll back a MySQL version upgrade
1. `git checkout` the previous `image: mysql:…` line in `docker-compose.yml`.
2. `docker compose stop`.
3. Copy the pre-upgrade volume back over `homeschool_mysql_data`:
```bash
docker run --rm -v homeschool_mysql_data_pre_<date>:/from:ro -v homeschool_mysql_data:/to \
alpine sh -c 'rm -rf /to/* && cp -a /from/. /to/'
```
Or recreate the volume and restore the dump.
4. `docker compose up -d`.
### Releases
Deployed versions are tagged `vX.Y.Z` in git, with notes in `Release-Notes/vX.Y.md`. To see what's running, compare `docker image inspect homeschool-backend --format '{{.Created}}'` to the tag date.
+36
View File
@@ -0,0 +1,36 @@
# v1.1
## v1.1.0 — 2026-09-24
Maintenance release from the 2026-09-24 review (`reports/maintenance-2026-09-24.md`).
### Security
- Rate limits now apply to each client. nginx resolves the real client IP through Cloudflare → NPM (`set_real_ip_from` + `real_ip_recursive`). Before this, every request looked like it came from NPM, so all users shared one login limit.
- ntfy alerts show the real client IP (`X-Real-IP` from nginx). They no longer use the first `X-Forwarded-For` entry, which the client controls.
- Python dependency updates:
- PyJWT 2.15.0 (fixes an auth bypass)
- anyio 4.15.1
- fastapi 0.141.1 / starlette 1.7.0
- cryptography 50.0.1
- python-multipart 0.0.32
- alembic 1.20.0 / Mako 1.4.3
- Base images:
- python 3.12.14-slim, with `apt-get upgrade` and without the unused gcc/mysqlclient headers
- nginx 1.30.5-alpine (stable), with `apk upgrade`
- Node 24 LTS for the build stage
- axios 1.20.0 and vite 5.4.21. Frontend dependencies are now locked in `package-lock.json` and installed with `npm ci`.
### Platform
- MySQL 8.0.40 (EOL) → **8.4.11 LTS**. The data dictionary is upgraded in place, and **you can't downgrade in place**. Restore the pre-upgrade volume copy or dump instead; see README "Backup, Restore & Rollback".
### Reliability
- The backend retries the DB connection for up to 180s at startup. Before this it crash-looped after host reboots, because the restart policy ignores `depends_on`.
- Healthchecks on backend (`/api/health`) and frontend. The db healthcheck has `start_period: 180s`.
- Docker log rotation on all services: json-file, 10 MB × 3.
### Fixes
- The meeting alert catch-up window fix (`8e92ae6`) is now deployed.
### Known / accepted
- npm audit: vite ≤6.4.2 and esbuild advisories affect only the **dev server**, which production never runs (it serves the static build). Fixing them fully needs Vite 6.4+/7.
- mysql:8.4.11 has upstream findings in Oracle's image: curl, bundled Python tools, and gosu's Go stdlib.
+19
View File
@@ -1,5 +1,7 @@
import asyncio
import logging import logging
import random import random
import time
from contextlib import asynccontextmanager from contextlib import asynccontextmanager
from fastapi import FastAPI, WebSocket, WebSocketDisconnect from fastapi import FastAPI, WebSocket, WebSocketDisconnect
@@ -67,8 +69,25 @@ async def _add_index_if_missing(conn, index_name: str, table: str, column: str):
raise raise
async def _wait_for_db(timeout: float = 180, interval: float = 2) -> None:
"""Retry until MySQL accepts connections. On host reboot Docker's restart
policy ignores depends_on, so the backend can start before the DB is up."""
deadline = time.monotonic() + timeout
while True:
try:
async with engine.connect() as conn:
await conn.execute(text("SELECT 1"))
return
except OperationalError as e:
if time.monotonic() >= deadline:
raise
logger.info("Database not ready (%s), retrying in %ss", e.orig, interval)
await asyncio.sleep(interval)
@asynccontextmanager @asynccontextmanager
async def lifespan(app: FastAPI): async def lifespan(app: FastAPI):
await _wait_for_db()
# Create tables on startup (Alembic handles migrations in prod, this is a safety net) # Create tables on startup (Alembic handles migrations in prod, this is a safety net)
async with engine.begin() as conn: async with engine.begin() as conn:
await conn.run_sync(Base.metadata.create_all) await conn.run_sync(Base.metadata.create_all)
+2 -1
View File
@@ -9,6 +9,7 @@ from app.auth.jwt import create_admin_token, create_access_token, hash_password
from app.config import get_settings from app.config import get_settings
from app.dependencies import get_db, get_admin_user from app.dependencies import get_db, get_admin_user
from app.models.user import User from app.models.user import User
from app.utils.client_ip import client_ip
from app.utils.ntfy import notify from app.utils.ntfy import notify
router = APIRouter(prefix="/api/admin", tags=["admin"]) router = APIRouter(prefix="/api/admin", tags=["admin"])
@@ -23,7 +24,7 @@ async def admin_login(body: dict, request: Request):
logger.warning("Failed super-admin login attempt for username=%s", username) logger.warning("Failed super-admin login attempt for username=%s", username)
raise HTTPException(status_code=status.HTTP_401_UNAUTHORIZED, detail="Invalid admin credentials") raise HTTPException(status_code=status.HTTP_401_UNAUTHORIZED, detail="Invalid admin credentials")
token = create_admin_token({"sub": "admin"}) token = create_admin_token({"sub": "admin"})
ip = request.headers.get("X-Forwarded-For", request.client.host if request.client else "unknown").split(",")[0].strip() ip = client_ip(request)
ua = request.headers.get("User-Agent", "unknown") ua = request.headers.get("User-Agent", "unknown")
await notify( await notify(
title="Homeschool Dashboard Super Admin Login", title="Homeschool Dashboard Super Admin Login",
+3 -2
View File
@@ -18,6 +18,7 @@ from app.models.user import User
from app.models.subject import Subject from app.models.subject import Subject
from app.schemas.auth import LoginRequest, RegisterRequest, TokenResponse from app.schemas.auth import LoginRequest, RegisterRequest, TokenResponse
from app.schemas.user import UserOut from app.schemas.user import UserOut
from app.utils.client_ip import client_ip
from app.utils.ntfy import notify from app.utils.ntfy import notify
router = APIRouter(prefix="/api/auth", tags=["auth"]) router = APIRouter(prefix="/api/auth", tags=["auth"])
@@ -59,7 +60,7 @@ async def register(body: RegisterRequest, response: Response, request: Request,
refresh = create_refresh_token({"sub": str(user.id)}) refresh = create_refresh_token({"sub": str(user.id)})
response.set_cookie(REFRESH_COOKIE, refresh, **COOKIE_OPTS) response.set_cookie(REFRESH_COOKIE, refresh, **COOKIE_OPTS)
ip = request.headers.get("X-Forwarded-For", request.client.host if request.client else "unknown").split(",")[0].strip() ip = client_ip(request)
ua = request.headers.get("User-Agent", "unknown") ua = request.headers.get("User-Agent", "unknown")
await notify( await notify(
title="Homeschool Dashboard New User Registered", title="Homeschool Dashboard New User Registered",
@@ -83,7 +84,7 @@ async def login(body: LoginRequest, response: Response, request: Request, db: As
raise HTTPException(status_code=401, detail="Invalid credentials") raise HTTPException(status_code=401, detail="Invalid credentials")
now = datetime.now(timezone.utc).replace(tzinfo=None) now = datetime.now(timezone.utc).replace(tzinfo=None)
ip = request.headers.get("X-Forwarded-For", request.client.host if request.client else "unknown").split(",")[0].strip() ip = client_ip(request)
ua = request.headers.get("User-Agent", "unknown") ua = request.headers.get("User-Agent", "unknown")
if user.locked_until and user.locked_until > now: if user.locked_until and user.locked_until > now:
+10
View File
@@ -0,0 +1,10 @@
from fastapi import Request
def client_ip(request: Request) -> str:
"""Real client IP as resolved by the frontend nginx (real_ip module).
X-Forwarded-For is client-controlled and must not be trusted here; nginx
sets X-Real-IP from $remote_addr after walking past trusted proxies.
"""
return request.headers.get("X-Real-IP") or (request.client.host if request.client else "unknown")
+20
View File
@@ -1,3 +1,9 @@
x-logging: &default-logging
driver: json-file
options:
max-size: "10m"
max-file: "3"
services: services:
db: db:
image: mysql:8.4.11 image: mysql:8.4.11
@@ -20,6 +26,7 @@ services:
start_period: 180s start_period: 180s
mem_limit: 512m mem_limit: 512m
cpus: 1.0 cpus: 1.0
logging: *default-logging
backend: backend:
build: ./backend build: ./backend
@@ -41,8 +48,15 @@ services:
condition: service_healthy condition: service_healthy
networks: networks:
- homeschool_net - homeschool_net
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/api/health', timeout=3)"]
interval: 30s
timeout: 5s
retries: 3
start_period: 180s
mem_limit: 512m mem_limit: 512m
cpus: 1.0 cpus: 1.0
logging: *default-logging
frontend: frontend:
build: ./frontend build: ./frontend
@@ -54,8 +68,14 @@ services:
- backend - backend
networks: networks:
- homeschool_net - homeschool_net
healthcheck:
test: ["CMD", "wget", "-q", "-O", "/dev/null", "http://127.0.0.1/"]
interval: 30s
timeout: 5s
retries: 3
mem_limit: 128m mem_limit: 128m
cpus: 0.5 cpus: 0.5
logging: *default-logging
networks: networks:
homeschool_net: homeschool_net:
+3 -3
View File
@@ -1,10 +1,10 @@
# Stage 1: Build Vue.js app # Stage 1: Build Vue.js app
FROM node:20.20.1-alpine AS builder FROM node:24.21.0-alpine AS builder
WORKDIR /app WORKDIR /app
COPY package.json package-lock.json* ./ COPY package.json package-lock.json ./
RUN npm install RUN npm ci
COPY . . COPY . .
RUN npm run build RUN npm run build
+1674
View File
File diff suppressed because it is too large. Load diff
+3 -3
View File
@@ -1,6 +1,6 @@
{ {
"name": "homeschool-frontend", "name": "homeschool-frontend",
"version": "1.0.0", "version": "1.1.0",
"private": true, "private": true,
"scripts": { "scripts": {
"dev": "vite", "dev": "vite",
@@ -8,13 +8,13 @@
"preview": "vite preview" "preview": "vite preview"
}, },
"dependencies": { "dependencies": {
"axios": "1.7.7", "axios": "1.20.0",
"pinia": "2.2.4", "pinia": "2.2.4",
"vue": "3.5.12", "vue": "3.5.12",
"vue-router": "4.4.5" "vue-router": "4.4.5"
}, },
"devDependencies": { "devDependencies": {
"@vitejs/plugin-vue": "5.1.4", "@vitejs/plugin-vue": "5.1.4",
"vite": "5.4.10" "vite": "5.4.21"
} }
} }
+10 -10
View File
@@ -134,7 +134,7 @@ This section was first blocked on backups. The user then asked for the HIGH find
| Real-IP rate limiting tested | PASS | Test nginx: client A was limited after its burst, client B was unaffected, and a forged XFF prefix resolved to the true client. | | Real-IP rate limiting tested | PASS | Test nginx: client A was limited after its burst, client B was unaffected, and a forged XFF prefix resolved to the true client. |
| Deploy | PASS* | The MySQL data upgrade (80040 → 80411) took about 90s, longer than the healthcheck window, so compose aborted the dependent containers. The user started them manually. Fixed by adding `start_period: 180s` to the db healthcheck. | | Deploy | PASS* | The MySQL data upgrade (80040 → 80411) took about 90s, longer than the healthcheck window, so compose aborted the dependent containers. The user started them manually. Fixed by adding `start_period: 180s` to the db healthcheck. |
| Sections 2/3 re-run | PASS (partial) | All containers up, backend logs clean, public `/api/health` 200. Live bundle now contains `8e92ae6`. nginx logs real client IPs. The user confirmed the site is back up. | | Sections 2/3 re-run | PASS (partial) | All containers up, backend logs clean, public `/api/health` 200. Live bundle now contains `8e92ae6`. nginx logs real client IPs. The user confirmed the site is back up. |
| Version tag / release notes | DEFERRED | The repo has no tagging or `Release-Notes/` convention yet. | | Version tag / release notes | PASS | `v1.1.0`, `Release-Notes/v1.1.md` |
| Old images cleanup | DEFERRED | Keep the rollback tags and volume copy until the update has been stable for about a week. Then remove them: `docker rmi homeschool-{backend,frontend}:rollback-20260924 mysql:8.0.40` and `docker volume rm homeschool_mysql_data_pre84_20260924`. | | Old images cleanup | DEFERRED | Keep the rollback tags and volume copy until the update has been stable for about a week. Then remove them: `docker rmi homeschool-{backend,frontend}:rollback-20260924 mysql:8.0.40` and `docker volume rm homeschool_mysql_data_pre84_20260924`. |
### Changes applied ### Changes applied
@@ -184,13 +184,13 @@ This section was first blocked on backups. The user then asked for the HIGH find
3. ~~Fixable HIGH/CRITICAL CVEs in all images and Python deps.~~ Fixed by the rebuild and dependency bumps. 3. ~~Fixable HIGH/CRITICAL CVEs in all images and Python deps.~~ Fixed by the rebuild and dependency bumps.
4. ~~MySQL 8.0 is EOL.~~ Upgraded to 8.4.11 LTS. 4. ~~MySQL 8.0 is EOL.~~ Upgraded to 8.4.11 LTS.
### MEDIUM ### MEDIUM (all resolved 2026-09-24, v1.1.0)
5. ~~The frontend is not deployed at HEAD.~~ Resolved by the rebuild. 5. ~~The frontend is not deployed at HEAD.~~ Resolved by the rebuild.
6. No Docker log rotation. 6. ~~No Docker log rotation.~~ Added `x-logging` anchor (json-file, 10m × 3) to all services.
7. The backend crash-loops on every host reboot (DB startup race). The db `start_period` does **not** fix this: on reboot, Docker's restart policy ignores `depends_on`. It needs connect-retry in the backend startup. backend and frontend have no healthchecks. 7. ~~The backend crash-loops on reboot, and there are no backend/frontend healthchecks.~~ `_wait_for_db()` in `main.py` retries for up to 180s. Tested: the backend started before its DB, logged retries, and came up with 0 restarts. Healthchecks were added to both services.
8. Node 20 (build stage) is EOL, and there is no `package-lock.json`. 8. ~~Node 20 EOL and no lockfile.~~ Now node:24.21.0-alpine with `package-lock.json` and `npm ci`. axios → 1.20.0, vite → 5.4.21. Remaining npm audit items (vite ≤6.4.2 / esbuild) are dev-server-only and accepted; the fix needs Vite 6.4+.
9. There is no rollback/restore runbook, and there are no release tags. 9. ~~No rollback/restore runbook or release tags.~~ README now has "Backup, Restore & Rollback" and documents real-IP proxy trust. Added `Release-Notes/v1.1.md` and tagged `v1.1.0`.
10. Client IP in ntfy alerts can be spoofed through `X-Forwarded-For`. 10. ~~Client IP in ntfy alerts can be spoofed.~~ New `app/utils/client_ip.py` reads `X-Real-IP`, which nginx sets after real_ip resolution. Tested: a forged XFF is ignored.
### LOW ### LOW
11. db memory is at 82% of its 512 MiB limit. 11. db memory is at 82% of its 512 MiB limit.
@@ -208,7 +208,7 @@ Portainer/NPM drift, functional smoke tests, ntfy test, off-host/VM backups, HTT
- Failed (critical/high): 7. That is backups (counted once, CRITICAL), rate limiting, image CVEs, dependency CVEs, runtime EOL, log rotation and frontend drift. - Failed (critical/high): 7. That is backups (counted once, CRITICAL), rate limiting, image CVEs, dependency CVEs, runtime EOL, log rotation and frontend drift.
- Failed (non-critical) / WARN: 18 - Failed (non-critical) / WARN: 18
- Pending user confirmation: 14 - Pending user confirmation: 14
- Deferred: 2 (release tag, old image cleanup) - Deferred: 1 (old image cleanup)
- Updates applied: 4 (real-IP rate limiting, Python deps, base images, MySQL 8.4) - Updates applied: 9 (HIGH: real-IP rate limiting, Python deps, base images, MySQL 8.4. MEDIUM: log rotation, DB-wait + healthchecks, Node 24 + lockfile, runbook + release tag, trusted client IP)
**Overall: NEEDS ATTENTION.** All HIGH findings are resolved. The CRITICAL no-backups finding is still open: the pre-update dump and volume copy are on-host only, with no schedule. **Overall: NEEDS ATTENTION.** All HIGH and MEDIUM findings are resolved. The CRITICAL no-backups finding is still open: the pre-update dump and volume copy are on-host only, with no schedule.