# Infrastructure Roadmap for brackt.com ## Context brackt.com runs as a **single Node.js process** on Vultr, fronted by Traefik, backed by managed Postgres via pgbouncer. Based on the original discovery: - **No HA fire.** Drafts are infrequent enough that a 30-min outage isn't catastrophic *yet*. - **Zero-downtime deploys are the immediate pain.** Connections drop on every release. - **Low concurrency today** (<5 live drafts, <20 in 6mo) — capacity is not the constraint. - **Avoid Redis** unless clearly justified; prefer Postgres-backed. The roadmap sequences work by **risk-adjusted value**: cheap, reversible wins first. --- ## Current production topology As of Phase A (merged 2026-06-08): - **Single `brackt` container** running on Vultr, image `sjc.vultrcr.com/chrisparsons/brackt:latest`. - **One-shot `migrate` container** runs migrations before brackt starts. - **Traefik** fronts the app — TLS via Let's Encrypt, apex + www redirect. - **External Postgres** via Vultr's built-in pgbouncer (`PGBOUNCER=true`). - **No in-process `setInterval` for timers or snapshots** — draft timer is deadline-based (`picksExpiresAt`); snapshots/standings/simulations are driven by external Forgejo Actions cron jobs hitting authenticated HTTP endpoints. - **No observability stack.** Logs live in `docker logs brackt`. - **Still present:** `container_name: brackt` (blocks replicas), no healthcheck, `:latest` tag (painful rollbacks). These are Phase B. --- ## Roadmap ### ✅ Phase 1 — Deadline-based draft timer (complete) The 1-second `setInterval` in `server/timer.ts` has been replaced with per-season `setTimeout` calls keyed to `picksExpiresAt` (a `timestamptz` on `draftTimers`). On process restart, active deadlines are recovered from the DB and rescheduled. **Result:** zero per-second DB writes or socket emits during a draft. The server wakes only at pick deadlines. All callsites updated: `make-pick`, `adjust-time-bank`, draft loader, `useDraftSocketEvents`. --- ### ✅ Phase A — External HTTP cron jobs (complete, merged 2026-06-08) Replaced the last in-process `setInterval` (`server/snapshots.ts` 24h loop) with Forgejo Actions scheduled workflows hitting authenticated endpoints. **New endpoints** (all behind `X-Cron-Secret` header): | Endpoint | Schedule | What it does | |---|---|---| | `POST /admin/jobs/run-daily-snapshots` | `5 0 * * *` | Daily fantasy standings snapshots for active/draft seasons | | `POST /admin/jobs/sync-and-simulate` | `0 */2 * * *` | Syncs standings from external APIs; simulates only when standings changed | | `GET /healthz` | Traefik / Docker healthcheck | 200 `{ok:true}` when DB reachable, 503 otherwise | **Schema additions** (migration 0118): `standingsLastChangedAt`, `lastSimulatedAt` on `sportsSeasons`. **Change detection:** `syncStandings()` compares `gamesPlayed` + `leagueRank` against current DB rows before upsert; sets `standingsLastChangedAt` only on real change. Simulation skipped if nothing changed. **Setup required on production:** add `CRON_SECRET` to Forgejo repo secrets and to the server `.env`. --- ### 🔜 Phase B — Production hardening + deploy automation (next, ~1 day) Three blockers for multi-replica work, plus automating compose file delivery. #### 1. Automate compose file delivery via deploy workflow Currently `.production-info/docker-compose-production.yml` is a manual reference — the server's `~/brackt/docker-compose.yml` drifts silently. Fix: add an `appleboy/scp-action` step to the deploy job that copies the repo's compose file to the server on every deploy. **Deploy workflow changes:** - Add checkout step to the deploy job (needed to access the file) - Add SCP step: copy `.production-info/docker-compose-production.yml` → `~/brackt/docker-compose.yml` - Build with both `:latest` and `:{sha}` tags; pass `IMAGE_TAG=${{ github.sha }}` to the SSH step - SSH step exports `IMAGE_TAG` before running `docker compose up` #### 2. Move secrets to a server-side `.env` file Currently secrets reach the container via unknown means (no `env_file:` in compose, no `environment:` entries for DB credentials). Make this explicit: add `env_file: .env` to the brackt service; document all required vars in `.production-info/.env.production.example`. Non-secret config (NODE_ENV, PGBOUNCER, APP_URL, etc.) stays inline in the compose file. Secrets (DATABASE_URL, BETTER_AUTH_SECRET, CRON_SECRET, RESEND_API_KEY, Cloudinary, Discord, Google, Sentry, Turnstile) go in the server `.env`. #### 3. Compose file fixes - **Remove `container_name: brackt`** — Docker rejects duplicate names; this alone blocks `replicas: 2`. - **Add healthcheck** using `/healthz` (already shipped in Phase A): ```yaml healthcheck: test: ["CMD", "curl", "-f", "http://localhost:3000/healthz"] interval: 30s timeout: 5s retries: 3 start_period: 10s ``` - **Pin image tag** to `${CONTAINER_REGISTRY}/brackt:${IMAGE_TAG:-latest}` — rollback becomes `IMAGE_TAG= docker compose up`. **Files:** - `.production-info/docker-compose-production.yml` — remove container_name, add healthcheck, parameterise image tag, add env_file - `.production-info/.env.production.example` — template for the server-side secrets file - `.forgejo/workflows/deploy.yml` — add checkout + SCP step, build dual tags, pass IMAGE_TAG **Verification:** - `curl https://brackt.com/healthz` → 200 `{ok:true}` - Deploy with a SHA tag; confirm rollback by setting `IMAGE_TAG=` and rerunning compose - `docker compose up --scale brackt=2` no longer fails with "name already in use" --- ### 🔜 Phase C — Zero-downtime deploys via Traefik sticky sessions (~2–3 days) Requires Phase B (no `container_name`, healthcheck, pinned tag). Run 2 web containers behind Traefik, with the **Redis Socket.IO adapter** handling cross-instance fan-out. Rolling deploy uses Docker Compose's `update_config` to bring up the new replica before stopping the old one. **Why Redis (not sticky sessions + polling/LISTEN/NOTIFY):** there are 19 room-targeted `getSocketIO().to(draft-${seasonId}).emit()` callsites across routes, models, and the timer. With sticky sessions, two different browsers watching the same draft (phone + laptop) can land on different instances — every one of those 19 emits needs to reach both. Polling introduces unacceptable latency for live draft events (pause, force-pick, timer). LISTEN/NOTIFY would require instrumenting all 19 callsites. The `@socket.io/redis-adapter` intercepts all of them automatically with zero callsite changes, and lets us drop sticky sessions entirely (simpler deploy). Redis is a single ~5MB container — at this point it's clearly justified. **Compose changes:** ```yaml redis: image: redis:7-alpine restart: unless-stopped networks: - internal # On the brackt service: deploy: replicas: 2 update_config: parallelism: 1 delay: 10s order: start-first # new replica healthy before old one stops environment: - REDIS_URL=redis://redis:6379 # (remove sticky session labels — not needed with Redis adapter) ``` **App change** (`server/socket.ts`): install `@socket.io/redis-adapter` + `redis`, attach the adapter when `REDIS_URL` is set. Gate on env var so dev works without Redis. **Rolling deploy:** `docker compose up -d` with `order: start-first` handles it automatically — no custom drain script needed. **Files:** `.production-info/docker-compose-production.yml`, `server/socket.ts`, `package.json`. **Verification:** start two instances locally, open two draft-room tabs, confirm pick events reach both. Rolling-deploy in production; clients see at most one brief reconnect. --- ### Phase D — Optional: split tick into its own process (~1 week, not urgent) After Phase 1, the "tick" is an idle process that wakes only at pick deadlines. It can stay co-located with web indefinitely. Only revisit if you want to deploy web without touching timer code at all. **If/when:** extract `server/tick.ts` as a standalone entrypoint, run as a second compose service, enforce singleton via Postgres advisory lock. --- ## Sequencing summary | Phase | Status | Effort | Key win | |---|---|---|---| | 1 | ✅ Done | — | Deadline-based timer; eliminates 1s tick | | A | ✅ Done | — | External cron; automated standings sync + sim; /healthz | | B | 🔜 Next | ~1d | Compose delivered by CI; pinned tags; healthcheck; unblocks C | | C | Blocked on B | ~2d | Zero-downtime rolling deploys via replicas + Redis adapter | | D | Optional | ~1wk | Deploy web without restarting timer process | **Total to zero-downtime deploys: ~3 days of focused work** (Phases B + C). --- ## Open questions - **Where do logs/metrics go?** Still unresolved. Options: structured JSON to stdout + `docker logs | jq` (zero effort), hosted (Better Stack / Axiom free tier, ~half day), or Grafana + Loki sidecars (~1 day). Pick one before adding more observability instrumentation. - **Pgbouncer pool mode:** confirm current pool is transaction mode. Add a second session-mode pool before adopting `LISTEN/NOTIFY` or pg-boss. --- ## Critical files - `server/timer.ts` — deadline-based scheduler; `picksExpiresAt` on `draftTimers` - `server/socket.ts` — emits `{ expiresAt, timeRemaining }` on `timer-pick-started` - `app/routes/healthz.ts` — DB ping; used by Docker healthcheck and Traefik - `app/routes/admin/jobs.run-daily-snapshots.ts` — daily snapshot cron endpoint - `app/routes/admin/jobs.sync-and-simulate.ts` — standings sync + conditional sim cron endpoint - `app/lib/cron-auth.ts` — `requireCronSecret()` helper for cron endpoints - `app/services/standings-sync/index.ts` — `syncStandings()` returns `{ changed: boolean }` - `.production-info/docker-compose-production.yml` — source of truth for production compose (Phase B: delivered by CI) - `.production-info/.env.production.example` — template for server-side secrets (Phase B) - `.forgejo/workflows/deploy.yml` — CI/CD pipeline (Phase B: add SCP + dual image tags) - `.forgejo/workflows/daily-snapshots.yml` — daily snapshot cron (Forgejo Actions) - `.forgejo/workflows/sync-and-simulate.yml` — 2h standings sync + sim cron (Forgejo Actions)