Fix MEDIUM findings from 2026-09-24 maintenance review

- Backend retries DB connection at startup (up to 180s) so host reboots
  no longer crash-loop it; add backend and frontend healthchecks
- Docker log rotation (json-file 10m x 3) on all services
- ntfy alerts use X-Real-IP (set by nginx after real_ip resolution)
  instead of the client-controlled first X-Forwarded-For entry
- Frontend build on Node 24 LTS with package-lock.json + npm ci;
  axios 1.20.0, vite 5.4.21
- README: backup/restore/rollback runbook, real-IP proxy trust notes
- Release-Notes/v1.1.md; version 1.1.0

Co-Authored-By: Claude Opus 5.5 <[email protected]>
This commit is contained in:
derekcandClaude Opus 5.5 committed 2026-09-24 23:44:25 -07:00
1 parent 3170c7f4eb
commit 4130467b22
11 files changed
+1851 -20

No files matched your search

+71 -1
View File
@@ -49,7 +49,7 @@ A self-hosted web app for managing homeschool schedules, tracking daily learning
| Frontend server | nginx (Docker) |
| Backend API | FastAPI (Python 3.12) |
| Real-time | WebSockets via FastAPI |
| Database | MySQL 8 |
| Database | MySQL 8.4 LTS |
| ORM | SQLAlchemy 2.0 (async) |
| Auth | JWT — PyJWT + passlib/bcrypt |
| Orchestration | Docker Compose |
@@ -374,6 +374,8 @@ Nginx enforces the following rate limits (per IP):
Requests over the limit receive a `429 Too Many Requests` response.
Limits are keyed on the real client IP. `frontend/nginx.conf` trusts the reverse proxy and Cloudflare (`set_real_ip_from`) and walks `X-Forwarded-For` past them. **If the NPM host IP changes, or you front the app with a different proxy/CDN, update those `set_real_ip_from` lines.** Otherwise every request appears to come from the proxy and all clients share one limit. The backend reads the resolved IP from `X-Real-IP` for ntfy alerts.
### Login Lockout
After 5 consecutive failed login attempts, a parent account is locked for 15 minutes. The lock clears automatically after the cooldown, or immediately when a super admin resets the user's password. The lockout threshold and duration are configured in `backend/app/routers/auth.py` (`_LOGIN_MAX_ATTEMPTS`, `_LOGIN_LOCKOUT_MINUTES`).
@@ -411,3 +413,71 @@ docker compose down -v
# Restart without rebuilding
docker compose up
```
---
## Backup, Restore & Rollback
All persistent state is in the MySQL volume (`homeschool_mysql_data`). The live `.env` is not in git, so keep a copy somewhere safe.
### Back up the database
```bash
mkdir -p ~/backups/homeschool
docker exec homeschool_db sh -c 'exec mysqldump -uroot -p"$MYSQL_ROOT_PASSWORD" \
--single-transaction --routines --triggers --events --databases "$MYSQL_DATABASE"' \
| gzip > ~/backups/homeschool/homeschool-$(date +%Y%m%d-%H%M).sql.gz
```
Copy the dump **off the Docker host**. A backup on the same disk doesn't survive a host failure.
### Restore the database from a dump
```bash
docker compose stop backend frontend
zcat ~/backups/homeschool/<file>.sql.gz | \
docker exec -i homeschool_db sh -c 'exec mysql -uroot -p"$MYSQL_ROOT_PASSWORD"'
docker compose start backend frontend
```
The dump includes `CREATE DATABASE`/`DROP TABLE` statements, so it replaces the current tables.
### Before an update: record a rollback point
```bash
# App images are built locally and tagged :latest, so a rebuild overwrites them
docker tag homeschool-backend homeschool-backend:rollback-$(date +%Y%m%d)
docker tag homeschool-frontend homeschool-frontend:rollback-$(date +%Y%m%d)
# Take a database backup (above). For a MySQL version upgrade, also copy the volume,
# because MySQL can't be downgraded in place:
docker compose stop
docker volume create homeschool_mysql_data_pre_<date>
docker run --rm -v homeschool_mysql_data:/from:ro -v homeschool_mysql_data_pre_<date>:/to \
alpine sh -c 'cp -a /from/. /to/'
```
### Roll back an app update
```bash
docker tag homeschool-backend:rollback-<date> homeschool-backend:latest
docker tag homeschool-frontend:rollback-<date> homeschool-frontend:latest
docker compose up -d --no-build
```
If the update changed the schema, restore the pre-update dump too (above).
### Roll back a MySQL version upgrade
1. `git checkout` the previous `image: mysql:…` line in `docker-compose.yml`.
2. `docker compose stop`.
3. Copy the pre-upgrade volume back over `homeschool_mysql_data`:
```bash
docker run --rm -v homeschool_mysql_data_pre_<date>:/from:ro -v homeschool_mysql_data:/to \
alpine sh -c 'rm -rf /to/* && cp -a /from/. /to/'
```
Or recreate the volume and restore the dump.
4. `docker compose up -d`.
### Releases
Deployed versions are tagged `vX.Y.Z` in git, with notes in `Release-Notes/vX.Y.md`. To see what's running, compare `docker image inspect homeschool-backend --format '{{.Created}}'` to the tag date.
+36
View File
@@ -0,0 +1,36 @@
# v1.1
## v1.1.0 — 2026-09-24
Maintenance release from the 2026-09-24 review (`reports/maintenance-2026-09-24.md`).
### Security
- Rate limits now apply to each client. nginx resolves the real client IP through Cloudflare → NPM (`set_real_ip_from` + `real_ip_recursive`). Before this, every request looked like it came from NPM, so all users shared one login limit.
- ntfy alerts show the real client IP (`X-Real-IP` from nginx). They no longer use the first `X-Forwarded-For` entry, which the client controls.
- Python dependency updates:
- PyJWT 2.15.0 (fixes an auth bypass)
- anyio 4.15.1
- fastapi 0.141.1 / starlette 1.7.0
- cryptography 50.0.1
- python-multipart 0.0.32
- alembic 1.20.0 / Mako 1.4.3
- Base images:
- python 3.12.14-slim, with `apt-get upgrade` and without the unused gcc/mysqlclient headers
- nginx 1.30.5-alpine (stable), with `apk upgrade`
- Node 24 LTS for the build stage
- axios 1.20.0 and vite 5.4.21. Frontend dependencies are now locked in `package-lock.json` and installed with `npm ci`.
### Platform
- MySQL 8.0.40 (EOL) → **8.4.11 LTS**. The data dictionary is upgraded in place, and **you can't downgrade in place**. Restore the pre-upgrade volume copy or dump instead; see README "Backup, Restore & Rollback".
### Reliability
- The backend retries the DB connection for up to 180s at startup. Before this it crash-looped after host reboots, because the restart policy ignores `depends_on`.
- Healthchecks on backend (`/api/health`) and frontend. The db healthcheck has `start_period: 180s`.
- Docker log rotation on all services: json-file, 10 MB × 3.
### Fixes
- The meeting alert catch-up window fix (`8e92ae6`) is now deployed.
### Known / accepted
- npm audit: vite ≤6.4.2 and esbuild advisories affect only the **dev server**, which production never runs (it serves the static build). Fixing them fully needs Vite 6.4+/7.
- mysql:8.4.11 has upstream findings in Oracle's image: curl, bundled Python tools, and gosu's Go stdlib.
+19
View File
@@ -1,5 +1,7 @@
import asyncio
import logging
import random
import time
from contextlib import asynccontextmanager
from fastapi import FastAPI, WebSocket, WebSocketDisconnect
@@ -67,8 +69,25 @@ async def _add_index_if_missing(conn, index_name: str, table: str, column: str):
raise
async def _wait_for_db(timeout: float = 180, interval: float = 2) -> None:
"""Retry until MySQL accepts connections. On host reboot Docker's restart
policy ignores depends_on, so the backend can start before the DB is up."""
deadline = time.monotonic() + timeout
while True:
try:
async with engine.connect() as conn:
await conn.execute(text("SELECT 1"))
return
except OperationalError as e:
if time.monotonic() >= deadline:
raise
logger.info("Database not ready (%s), retrying in %ss", e.orig, interval)
await asyncio.sleep(interval)
@asynccontextmanager
async def lifespan(app: FastAPI):
await _wait_for_db()
# Create tables on startup (Alembic handles migrations in prod, this is a safety net)
async with engine.begin() as conn:
await conn.run_sync(Base.metadata.create_all)
+2 -1
View File
@@ -9,6 +9,7 @@ from app.auth.jwt import create_admin_token, create_access_token, hash_password
from app.config import get_settings
from app.dependencies import get_db, get_admin_user
from app.models.user import User
from app.utils.client_ip import client_ip
from app.utils.ntfy import notify
router = APIRouter(prefix="/api/admin", tags=["admin"])
@@ -23,7 +24,7 @@ async def admin_login(body: dict, request: Request):
logger.warning("Failed super-admin login attempt for username=%s", username)
raise HTTPException(status_code=status.HTTP_401_UNAUTHORIZED, detail="Invalid admin credentials")
token = create_admin_token({"sub": "admin"})
ip = request.headers.get("X-Forwarded-For", request.client.host if request.client else "unknown").split(",")[0].strip()
ip = client_ip(request)
ua = request.headers.get("User-Agent", "unknown")
await notify(
title="Homeschool Dashboard Super Admin Login",
+3 -2
View File
@@ -18,6 +18,7 @@ from app.models.user import User
from app.models.subject import Subject
from app.schemas.auth import LoginRequest, RegisterRequest, TokenResponse
from app.schemas.user import UserOut
from app.utils.client_ip import client_ip
from app.utils.ntfy import notify
router = APIRouter(prefix="/api/auth", tags=["auth"])
@@ -59,7 +60,7 @@ async def register(body: RegisterRequest, response: Response, request: Request,
refresh = create_refresh_token({"sub": str(user.id)})
response.set_cookie(REFRESH_COOKIE, refresh, **COOKIE_OPTS)
ip = request.headers.get("X-Forwarded-For", request.client.host if request.client else "unknown").split(",")[0].strip()
ip = client_ip(request)
ua = request.headers.get("User-Agent", "unknown")
await notify(
title="Homeschool Dashboard New User Registered",
@@ -83,7 +84,7 @@ async def login(body: LoginRequest, response: Response, request: Request, db: As
raise HTTPException(status_code=401, detail="Invalid credentials")
now = datetime.now(timezone.utc).replace(tzinfo=None)
ip = request.headers.get("X-Forwarded-For", request.client.host if request.client else "unknown").split(",")[0].strip()
ip = client_ip(request)
ua = request.headers.get("User-Agent", "unknown")
if user.locked_until and user.locked_until > now:
+10
View File
@@ -0,0 +1,10 @@
from fastapi import Request
def client_ip(request: Request) -> str:
"""Real client IP as resolved by the frontend nginx (real_ip module).
X-Forwarded-For is client-controlled and must not be trusted here; nginx
sets X-Real-IP from $remote_addr after walking past trusted proxies.
"""
return request.headers.get("X-Real-IP") or (request.client.host if request.client else "unknown")
+20
View File
@@ -1,3 +1,9 @@
x-logging: &default-logging
driver: json-file
options:
max-size: "10m"
max-file: "3"
services:
db:
image: mysql:8.4.11
@@ -20,6 +26,7 @@ services:
start_period: 180s
mem_limit: 512m
cpus: 1.0
logging: *default-logging
backend:
build: ./backend
@@ -41,8 +48,15 @@ services:
condition: service_healthy
networks:
- homeschool_net
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://localhost:8000/api/health', timeout=3)"]
interval: 30s
timeout: 5s
retries: 3
start_period: 180s
mem_limit: 512m
cpus: 1.0
logging: *default-logging
frontend:
build: ./frontend
@@ -54,8 +68,14 @@ services:
- backend
networks:
- homeschool_net
healthcheck:
test: ["CMD", "wget", "-q", "-O", "/dev/null", "http://127.0.0.1/"]
interval: 30s
timeout: 5s
retries: 3
mem_limit: 128m
cpus: 0.5
logging: *default-logging
networks:
homeschool_net:
+3 -3
View File
@@ -1,10 +1,10 @@
# Stage 1: Build Vue.js app
FROM node:20.20.1-alpine AS builder
FROM node:24.21.0-alpine AS builder
WORKDIR /app
COPY package.json package-lock.json* ./
RUN npm install
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build
+1674
View File
File diff suppressed because it is too large. Load diff
+3 -3
View File
@@ -1,6 +1,6 @@
{
"name": "homeschool-frontend",
"version": "1.0.0",
"version": "1.1.0",
"private": true,
"scripts": {
"dev": "vite",
@@ -8,13 +8,13 @@
"preview": "vite preview"
},
"dependencies": {
"axios": "1.7.7",
"axios": "1.20.0",
"pinia": "2.2.4",
"vue": "3.5.12",
"vue-router": "4.4.5"
},
"devDependencies": {
"@vitejs/plugin-vue": "5.1.4",
"vite": "5.4.10"
"vite": "5.4.21"
}
}
+10 -10
View File
@@ -134,7 +134,7 @@ This section was first blocked on backups. The user then asked for the HIGH find
| Real-IP rate limiting tested | PASS | Test nginx: client A was limited after its burst, client B was unaffected, and a forged XFF prefix resolved to the true client. |
| Deploy | PASS* | The MySQL data upgrade (80040 → 80411) took about 90s, longer than the healthcheck window, so compose aborted the dependent containers. The user started them manually. Fixed by adding `start_period: 180s` to the db healthcheck. |
| Sections 2/3 re-run | PASS (partial) | All containers up, backend logs clean, public `/api/health` 200. Live bundle now contains `8e92ae6`. nginx logs real client IPs. The user confirmed the site is back up. |
| Version tag / release notes | DEFERRED | The repo has no tagging or `Release-Notes/` convention yet. |
| Version tag / release notes | PASS | `v1.1.0`, `Release-Notes/v1.1.md` |
| Old images cleanup | DEFERRED | Keep the rollback tags and volume copy until the update has been stable for about a week. Then remove them: `docker rmi homeschool-{backend,frontend}:rollback-20260924 mysql:8.0.40` and `docker volume rm homeschool_mysql_data_pre84_20260924`. |
### Changes applied
@@ -184,13 +184,13 @@ This section was first blocked on backups. The user then asked for the HIGH find
3. ~~Fixable HIGH/CRITICAL CVEs in all images and Python deps.~~ Fixed by the rebuild and dependency bumps.
4. ~~MySQL 8.0 is EOL.~~ Upgraded to 8.4.11 LTS.
### MEDIUM
### MEDIUM (all resolved 2026-09-24, v1.1.0)
5. ~~The frontend is not deployed at HEAD.~~ Resolved by the rebuild.
6. No Docker log rotation.
7. The backend crash-loops on every host reboot (DB startup race). The db `start_period` does **not** fix this: on reboot, Docker's restart policy ignores `depends_on`. It needs connect-retry in the backend startup. backend and frontend have no healthchecks.
8. Node 20 (build stage) is EOL, and there is no `package-lock.json`.
9. There is no rollback/restore runbook, and there are no release tags.
10. Client IP in ntfy alerts can be spoofed through `X-Forwarded-For`.
6. ~~No Docker log rotation.~~ Added `x-logging` anchor (json-file, 10m × 3) to all services.
7. ~~The backend crash-loops on reboot, and there are no backend/frontend healthchecks.~~ `_wait_for_db()` in `main.py` retries for up to 180s. Tested: the backend started before its DB, logged retries, and came up with 0 restarts. Healthchecks were added to both services.
8. ~~Node 20 EOL and no lockfile.~~ Now node:24.21.0-alpine with `package-lock.json` and `npm ci`. axios → 1.20.0, vite → 5.4.21. Remaining npm audit items (vite ≤6.4.2 / esbuild) are dev-server-only and accepted; the fix needs Vite 6.4+.
9. ~~No rollback/restore runbook or release tags.~~ README now has "Backup, Restore & Rollback" and documents real-IP proxy trust. Added `Release-Notes/v1.1.md` and tagged `v1.1.0`.
10. ~~Client IP in ntfy alerts can be spoofed.~~ New `app/utils/client_ip.py` reads `X-Real-IP`, which nginx sets after real_ip resolution. Tested: a forged XFF is ignored.
### LOW
11. db memory is at 82% of its 512 MiB limit.
@@ -208,7 +208,7 @@ Portainer/NPM drift, functional smoke tests, ntfy test, off-host/VM backups, HTT
- Failed (critical/high): 7. That is backups (counted once, CRITICAL), rate limiting, image CVEs, dependency CVEs, runtime EOL, log rotation and frontend drift.
- Failed (non-critical) / WARN: 18
- Pending user confirmation: 14
- Deferred: 2 (release tag, old image cleanup)
- Updates applied: 4 (real-IP rate limiting, Python deps, base images, MySQL 8.4)
- Deferred: 1 (old image cleanup)
- Updates applied: 9 (HIGH: real-IP rate limiting, Python deps, base images, MySQL 8.4. MEDIUM: log rotation, DB-wait + healthchecks, Node 24 + lockfile, runbook + release tag, trusted client IP)
**Overall: NEEDS ATTENTION.** All HIGH findings are resolved. The CRITICAL no-backups finding is still open: the pre-update dump and volume copy are on-host only, with no schedule.
**Overall: NEEDS ATTENTION.** All HIGH and MEDIUM findings are resolved. The CRITICAL no-backups finding is still open: the pre-update dump and volume copy are on-host only, with no schedule.