brackt/docs/infrastructure-roadmap.md
Chris Parsons 164c2adfb7
All checks were successful
🚀 Deploy / 🧪 Test (push) Successful in 2m39s
🚀 Deploy / ʦ🔍 Typecheck & Lint (push) Successful in 1m23s
🚀 Deploy / 🐳 Build (push) Successful in 1m10s
🚀 Deploy / 🚀 Deploy (push) Successful in 12s
Update infrastructure roadmap: Phase C uses Redis adapter not sticky sessions
19 room-targeted socket emit callsites need cross-instance fan-out — the
Redis adapter handles them all automatically. Drops sticky sessions and
the custom drain script in favour of compose update_config order:start-first.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-08 01:09:00 -07:00

188 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Infrastructure Roadmap for brackt.com
## Context
brackt.com runs as a **single Node.js process** on Vultr, fronted by Traefik, backed by managed Postgres via pgbouncer.
Based on the original discovery:
- **No HA fire.** Drafts are infrequent enough that a 30-min outage isn't catastrophic *yet*.
- **Zero-downtime deploys are the immediate pain.** Connections drop on every release.
- **Low concurrency today** (<5 live drafts, <20 in 6mo) capacity is not the constraint.
- **Avoid Redis** unless clearly justified; prefer Postgres-backed.
The roadmap sequences work by **risk-adjusted value**: cheap, reversible wins first.
---
## Current production topology
As of Phase A (merged 2026-06-08):
- **Single `brackt` container** running on Vultr, image `sjc.vultrcr.com/chrisparsons/brackt:latest`.
- **One-shot `migrate` container** runs migrations before brackt starts.
- **Traefik** fronts the app TLS via Let's Encrypt, apex + www redirect.
- **External Postgres** via Vultr's built-in pgbouncer (`PGBOUNCER=true`).
- **No in-process `setInterval` for timers or snapshots** draft timer is deadline-based (`picksExpiresAt`); snapshots/standings/simulations are driven by external Forgejo Actions cron jobs hitting authenticated HTTP endpoints.
- **No observability stack.** Logs live in `docker logs brackt`.
- **Still present:** `container_name: brackt` (blocks replicas), no healthcheck, `:latest` tag (painful rollbacks). These are Phase B.
---
## Roadmap
### ✅ Phase 1 — Deadline-based draft timer (complete)
The 1-second `setInterval` in `server/timer.ts` has been replaced with per-season `setTimeout` calls keyed to `picksExpiresAt` (a `timestamptz` on `draftTimers`). On process restart, active deadlines are recovered from the DB and rescheduled.
**Result:** zero per-second DB writes or socket emits during a draft. The server wakes only at pick deadlines. All callsites updated: `make-pick`, `adjust-time-bank`, draft loader, `useDraftSocketEvents`.
---
### ✅ Phase A — External HTTP cron jobs (complete, merged 2026-06-08)
Replaced the last in-process `setInterval` (`server/snapshots.ts` 24h loop) with Forgejo Actions scheduled workflows hitting authenticated endpoints.
**New endpoints** (all behind `X-Cron-Secret` header):
| Endpoint | Schedule | What it does |
|---|---|---|
| `POST /admin/jobs/run-daily-snapshots` | `5 0 * * *` | Daily fantasy standings snapshots for active/draft seasons |
| `POST /admin/jobs/sync-and-simulate` | `0 */2 * * *` | Syncs standings from external APIs; simulates only when standings changed |
| `GET /healthz` | Traefik / Docker healthcheck | 200 `{ok:true}` when DB reachable, 503 otherwise |
**Schema additions** (migration 0118): `standingsLastChangedAt`, `lastSimulatedAt` on `sportsSeasons`.
**Change detection:** `syncStandings()` compares `gamesPlayed` + `leagueRank` against current DB rows before upsert; sets `standingsLastChangedAt` only on real change. Simulation skipped if nothing changed.
**Setup required on production:** add `CRON_SECRET` to Forgejo repo secrets and to the server `.env`.
---
### 🔜 Phase B — Production hardening + deploy automation (next, ~1 day)
Three blockers for multi-replica work, plus automating compose file delivery.
#### 1. Automate compose file delivery via deploy workflow
Currently `.production-info/docker-compose-production.yml` is a manual reference the server's `~/brackt/docker-compose.yml` drifts silently. Fix: add an `appleboy/scp-action` step to the deploy job that copies the repo's compose file to the server on every deploy.
**Deploy workflow changes:**
- Add checkout step to the deploy job (needed to access the file)
- Add SCP step: copy `.production-info/docker-compose-production.yml` `~/brackt/docker-compose.yml`
- Build with both `:latest` and `:{sha}` tags; pass `IMAGE_TAG=${{ github.sha }}` to the SSH step
- SSH step exports `IMAGE_TAG` before running `docker compose up`
#### 2. Move secrets to a server-side `.env` file
Currently secrets reach the container via unknown means (no `env_file:` in compose, no `environment:` entries for DB credentials). Make this explicit: add `env_file: .env` to the brackt service; document all required vars in `.production-info/.env.production.example`.
Non-secret config (NODE_ENV, PGBOUNCER, APP_URL, etc.) stays inline in the compose file. Secrets (DATABASE_URL, BETTER_AUTH_SECRET, CRON_SECRET, RESEND_API_KEY, Cloudinary, Discord, Google, Sentry, Turnstile) go in the server `.env`.
#### 3. Compose file fixes
- **Remove `container_name: brackt`** Docker rejects duplicate names; this alone blocks `replicas: 2`.
- **Add healthcheck** using `/healthz` (already shipped in Phase A):
```yaml
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:3000/healthz"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
```
- **Pin image tag** to `${CONTAINER_REGISTRY}/brackt:${IMAGE_TAG:-latest}` — rollback becomes `IMAGE_TAG=<sha> docker compose up`.
**Files:**
- `.production-info/docker-compose-production.yml` — remove container_name, add healthcheck, parameterise image tag, add env_file
- `.production-info/.env.production.example` — template for the server-side secrets file
- `.forgejo/workflows/deploy.yml` — add checkout + SCP step, build dual tags, pass IMAGE_TAG
**Verification:**
- `curl https://brackt.com/healthz` → 200 `{ok:true}`
- Deploy with a SHA tag; confirm rollback by setting `IMAGE_TAG=<previous-sha>` and rerunning compose
- `docker compose up --scale brackt=2` no longer fails with "name already in use"
---
### 🔜 Phase C — Zero-downtime deploys via Traefik sticky sessions (~23 days)
Requires Phase B (no `container_name`, healthcheck, pinned tag).
Run 2 web containers behind Traefik, with the **Redis Socket.IO adapter** handling cross-instance fan-out. Rolling deploy uses Docker Compose's `update_config` to bring up the new replica before stopping the old one.
**Why Redis (not sticky sessions + polling/LISTEN/NOTIFY):** there are 19 room-targeted `getSocketIO().to(draft-${seasonId}).emit()` callsites across routes, models, and the timer. With sticky sessions, two different browsers watching the same draft (phone + laptop) can land on different instances — every one of those 19 emits needs to reach both. Polling introduces unacceptable latency for live draft events (pause, force-pick, timer). LISTEN/NOTIFY would require instrumenting all 19 callsites. The `@socket.io/redis-adapter` intercepts all of them automatically with zero callsite changes, and lets us drop sticky sessions entirely (simpler deploy). Redis is a single ~5MB container — at this point it's clearly justified.
**Compose changes:**
```yaml
redis:
image: redis:7-alpine
restart: unless-stopped
networks:
- internal
# On the brackt service:
deploy:
replicas: 2
update_config:
parallelism: 1
delay: 10s
order: start-first # new replica healthy before old one stops
environment:
- REDIS_URL=redis://redis:6379
# (remove sticky session labels — not needed with Redis adapter)
```
**App change** (`server/socket.ts`): install `@socket.io/redis-adapter` + `redis`, attach the adapter when `REDIS_URL` is set. Gate on env var so dev works without Redis.
**Rolling deploy:** `docker compose up -d` with `order: start-first` handles it automatically — no custom drain script needed.
**Files:** `.production-info/docker-compose-production.yml`, `server/socket.ts`, `package.json`.
**Verification:** start two instances locally, open two draft-room tabs, confirm pick events reach both. Rolling-deploy in production; clients see at most one brief reconnect.
---
### Phase D — Optional: split tick into its own process (~1 week, not urgent)
After Phase 1, the "tick" is an idle process that wakes only at pick deadlines. It can stay co-located with web indefinitely. Only revisit if you want to deploy web without touching timer code at all.
**If/when:** extract `server/tick.ts` as a standalone entrypoint, run as a second compose service, enforce singleton via Postgres advisory lock.
---
## Sequencing summary
| Phase | Status | Effort | Key win |
|---|---|---|---|
| 1 | ✅ Done | — | Deadline-based timer; eliminates 1s tick |
| A | ✅ Done | — | External cron; automated standings sync + sim; /healthz |
| B | 🔜 Next | ~1d | Compose delivered by CI; pinned tags; healthcheck; unblocks C |
| C | Blocked on B | ~2d | Zero-downtime rolling deploys via replicas + Redis adapter |
| D | Optional | ~1wk | Deploy web without restarting timer process |
**Total to zero-downtime deploys: ~3 days of focused work** (Phases B + C).
---
## Open questions
- **Where do logs/metrics go?** Still unresolved. Options: structured JSON to stdout + `docker logs | jq` (zero effort), hosted (Better Stack / Axiom free tier, ~half day), or Grafana + Loki sidecars (~1 day). Pick one before adding more observability instrumentation.
- **Pgbouncer pool mode:** confirm current pool is transaction mode. Add a second session-mode pool before adopting `LISTEN/NOTIFY` or pg-boss.
---
## Critical files
- `server/timer.ts` — deadline-based scheduler; `picksExpiresAt` on `draftTimers`
- `server/socket.ts` — emits `{ expiresAt, timeRemaining }` on `timer-pick-started`
- `app/routes/healthz.ts` — DB ping; used by Docker healthcheck and Traefik
- `app/routes/admin/jobs.run-daily-snapshots.ts` — daily snapshot cron endpoint
- `app/routes/admin/jobs.sync-and-simulate.ts` — standings sync + conditional sim cron endpoint
- `app/lib/cron-auth.ts` — `requireCronSecret()` helper for cron endpoints
- `app/services/standings-sync/index.ts` — `syncStandings()` returns `{ changed: boolean }`
- `.production-info/docker-compose-production.yml` — source of truth for production compose (Phase B: delivered by CI)
- `.production-info/.env.production.example` — template for server-side secrets (Phase B)
- `.forgejo/workflows/deploy.yml` — CI/CD pipeline (Phase B: add SCP + dual image tags)
- `.forgejo/workflows/daily-snapshots.yml` — daily snapshot cron (Forgejo Actions)
- `.forgejo/workflows/sync-and-simulate.yml` — 2h standings sync + sim cron (Forgejo Actions)