Two tightly-coupled bugs visible after the PUBLIC_BASE_URL fix landed:
1. Trigger URLs now correctly show :22784 (the host-exposed nginx port),
but opening one in the browser rendered the SPA's NotFound page instead
of hitting the backend trigger handler.
Root cause: dev.nginx only had `location /` → frontend_dev. No location
block for `/c/`, `/k/`, or even `/api/`. Every request went to the SPA;
only `/api/*` happened to work because vite's internal proxy caught it
on its way through frontend_dev → canary. /c/* and /k/* (the trigger
paths) just fell through to the SPA fallback.
Fix: add `upstream canary { server canary:8080 }` to nginx.conf and
three new location blocks in dev.nginx for /api/, /c/, /k/. Each one
proxy_passes directly to the canary upstream with X-Real-IP +
X-Forwarded-For + X-Forwarded-Proto set so middleware.RealIP() resolves
the actual visitor IP rather than the docker bridge. Mirrors what
spec §12.3 prescribes for prod nginx.
2. Manage URL in responses still showed :8080 even after PUBLIC_BASE_URL
was set to :22784. Two separate config keys: canary.base_url
(PUBLIC_BASE_URL → trigger_url) and canary.manage_url (CANARY_MANAGE_URL
→ manage_url). The latter has no env-name overlap with the spec's
PUBLIC_BASE_URL, so it kept its default literal "http://localhost:8080".
Fix: in config.Load(), after unmarshal, fall manage_url back to base_url
when it's empty or still equals the bare default. Operators setting one
env var (PUBLIC_BASE_URL) now get both URLs computed off it. Explicit
CANARY_MANAGE_URL still wins if set — preserves the prod ability to
serve manage at a different domain than triggers. Extracted the literal
to a `defaultCanaryBaseURL` const so the default and the fallback
sentinel reference the same string.
After this + the prior PUBLIC_BASE_URL alias, a fresh token from the form
should produce trigger_url + manage_url both pointing at
http://localhost:${NGINX_HOST_PORT}/... — clickable, reachable, and triggers
hitting /c/ now route to canary and record events.
Operator action: just dev-down && just dev-up to rebuild nginx + canary
containers with the new config.
Spec §12.4, .env.example, .env.development, and dev.compose.yml all use the
name PUBLIC_BASE_URL for the externally-reachable URL stamped into trigger/
manage links. But the koanf env-key map in config.go only routed
CANARY_BASE_URL → canary.base_url. So setting PUBLIC_BASE_URL had no effect,
the default fallback "http://localhost:8080" silently won, and every issued
token came back with trigger_url/manage_url like
"http://localhost:8080/c/<id>" regardless of where canary was actually reachable.
Visible to the operator as: trigger/manage URLs in the SPECIMEN ISSUED dossier
pointing at :8080 (where canary doesn't even listen externally — host port is
nginx ${NGINX_HOST_PORT}, defaulting to randomised 22784).
Fix is one line: add PUBLIC_BASE_URL alongside CANARY_BASE_URL in the env-key
map. Both env names now route to canary.base_url. Backwards-compatible —
existing CANARY_BASE_URL deployments keep working.
Operator action: set `PUBLIC_BASE_URL=http://localhost:${NGINX_HOST_PORT}` in
.env.development (e.g. http://localhost:22784) and restart canary. URLs will
then reflect the actual reachable address.
- internal/geoip: Service over oschwald/geoip2-golang v2 with City db
enrichment (Country/Region/Subdivision/City) and a NopService
factory for environments without the mmdb file. Nil-safe Lookup
returns a zero geoip.Lookup on any failure path (nil receiver,
nil reader, empty/malformed IP, reader error, !rec.HasData).
- internal/geoip: unexported cityReader interface allows whitebox
fake-based unit testing without bundling a real .mmdb fixture;
extractLookup is a pure function for synthetic-record coverage.
- config: GeoIPConfig{Path} koanf section; GEOLITE_PATH env binding
matching compose.yml; default /data/GeoLite2-City.mmdb.
- deps: github.com/oschwald/geoip2-golang/v2 v2.1.0.
ASN/ASNOrg fields are reserved on geoip.Lookup but stay zero with a
City-only mmdb; populating them is a future-extension concern (ASN
db is not fetched by scripts/init.sh per design §12.5).
- internal/middleware/operator_bearer.go: returns 404 on miss (NOT 401)
per spec §11.5; subtle.ConstantTimeCompare for timing safety; rejects
when configured token is empty (defense-in-depth even though wire-up
skips mount in that case).
- Table-driven test covers: missing header, empty header, wrong scheme
(Basic), case-sensitive scheme (lowercase bearer), no payload,
no-space-after-Bearer, equal-length wrong token, different-length
wrong token, correct token, empty-config rejects both empty and
populated headers, extra-whitespace rejection. Also asserts no
WWW-Authenticate header on 404 (endpoint-hiding intent).
- config.OperatorConfig + OPERATOR_TOKEN env mapping; default empty
(mount is skipped when empty).
Swaps the Phase-9 directEventRecorder/mysqlEventRecorder shims for
real event.Service.Record via an eventRecorderAdapter. Constructs
notify.Service with telegram + webhook senders. Spawns the retention
loop goroutine off the main wait group; gracefulShutdown now also
notifySvc.Wait()s for in-flight notifications to drain.
Closes deferred Phase-9 items 2 (real event recorder), 3 (nested
create-tier rate limits), 5 (FingerprintRecorder wiring), 7 (drop
required tag on TurnstileResp).
- cmd/canary/main.go:
* buildEventStack constructs telegram + webhook senders and registers
them on notify.Service. event.Service wired with eventRepo (Store +
StatusWriter), tokenRepo (TokenIncrementer), redis client, notifier.
* eventRecorderAdapter and a single mysqlEventRecorder both delegate
to event.Service.Record(ctx, t.NotifyInfo(), evt) — one path,
same dedup + notify semantics for HTTP triggers and mysql triggers.
* fingerprintRecorderAdapter wraps event.Repository.AttachFingerprint
with the configured window (defaults to 5 min) so the slow-redirect
fingerprint POST writes through.
* spawnRetentionLoop runs event.Service.RunRetentionLoop on the
shared wait group; configured interval + per-token limit gate it.
* /api/tokens POST gains nested create-tier limits per spec §11.3:
LimitCreateMin (5/min) + LimitCreateHour (20/hr) chained ahead of
Turnstile, with their own fingerprint-keyed Redis namespaces so
they don't share buckets with the read-tier limiter.
* Shutdown order: server → telemetry → redis → db, then
notifySvc.Wait() to drain in-flight goroutines, then wg.Wait()
for retention loop + mysql listener.
- internal/config/config.go:
* NotifyConfig: dedup_ttl, send_timeout, max_tries, max_elapsed,
initial_interval, retention_interval, retention_limit,
webhook_hmac_secret, telegram_api_base, fingerprint_window.
* RateLimitConfig: create_min_rate / burst, create_hour_rate / burst.
* Defaults + env mappings (NOTIFY_*, RATE_LIMIT_CREATE_*, plus the
spec'd WEBHOOK_HMAC_SECRET).
- internal/token/dto.go: drop the `validate:"required"` tag on
cf_turnstile_response. The middleware (TurnstileVerify) is the sole
authority — it bypasses on empty SecretKey for dev mode and would
reject empty-token submissions in prod via the Verifier's own
ErrEmptyToken path. The validator tag was forcing tests to pass
a placeholder string in dev mode which conflicted with the bypass.
Phase 9 tasks 9.7 + 9.8 + closes Phase 8 task 8.5 deferral.
config.Config additions:
- Canary {BaseURL, ManageURL} — public URLs for trigger
+ manage redirects
- Turnstile {SecretKey, SiteKey} — Cloudflare Turnstile
(empty SecretKey = dev bypass)
- MySQL {Enabled, Addr, PublicHost, — fake TCP listener config
PublicPort}
cmd/canary/main.go full wire-up per spec §6.3 + §11.1:
- Global middleware: RequestID → Logger → Recovery → SecurityHeaders
- /healthz: no further middleware (router-mounted directly)
- /c/{id}, /c/{id}/fingerprint, /k/{id}, /k/{id}/*: mounted at
router root, OUTSIDE /api subrouter, so they skip CORS + rate
limit (spec §11.1 line 1552 — we want to record EVERY hit)
- /api subrouter: CORS → RateLimit (KeyByFingerprint, FailOpen=true)
- /api/tokens/types: no further middleware (read-only public list)
- /api/tokens POST: + TurnstileVerify middleware
- mysql goroutine: spawned via spawnMySQLListener iff
cfg.MySQL.Enabled; uses mysqlTokenLookup + mysqlEventRecorder
adapters bridging mysql.{TokenLookup,EventRecorder} to
token.Service + event.Repository
Adapter types in cmd/canary/main.go:
- registryAdapter: bridges registry.Registry (map type) to
token.Registry interface (Get method)
- directEventRecorder: bridges token.EventRecorder to
event.Repository.Insert + token.Repository.IncrementTriggerCount
(synchronous; Phase 10 swaps in event.Service for dedup + async
notify)
- mysqlTokenLookup, mysqlEventRecorder: same shape for the TCP
mysql handler
webbug refactor (task 9.8):
- Delete local realIP/lastNonEmptyXFF/optionalHeader copies (~30 LOC)
- Use middleware.RealIP and middleware.OptionalHeader instead
- Tests still pass (the IP-precedence suite in webbug_test.go
exercises middleware.RealIP behavior end-to-end via the Trigger
contract)
Pre-rollup gate clean at HEAD:
- go build ./... + go vet ./... clean
- go test -race -timeout=60s ./... all pass
- go test -tags=integration -race -timeout=300s ./internal/token/...
./internal/event/... all pass
- golangci-lint run ./... → 0 issues
- grep -rn "//nolint" --include="*.go" returns nothing
DEFERRED (not in Phase 9 scope):
- Phase 9 task 9.9 integration test (full POST /api/tokens against
testcontainers) — defer to Phase 10's broader integration suite
along with event.Service + notify.Service end-to-end coverage
- Phase 10 brings: event.Service with dedup + async notify, swaps
in for the directEventRecorder/mysqlEventRecorder adapters
- Phase 13 brings: geoip.Service wired into event recording
Pre-audit gate (`golangci-lint run`) at the start of Phase 2 surfaced 15
issues that should have been zeroed before Phase 1 rollup. Per the
fix-in-phase rule (no backlog rot), clearing everything before Phase 2
audit agents run on a green tree.
Config:
- .golangci.yml: migrate `issues.exclude-rules` and `issues.exclude-dirs`
to v2 syntax (`linters.exclusions.{rules,paths}`); test-file funlen/
dupl/goconst exclusion now actually applies under golangci-lint v2.10
- .golangci.yml: add G706 to gosec excludes — false-positive log-injection
reports on slog structured-logging call sites (slog separates message
from kv args, immune to log-line injection by construction)
errcheck (5):
- cmd/canary/main.go: `_ = telemetry.Shutdown(...)` → log on error
- internal/token/repository.go: `defer stmt.Close()` → log on close error
- internal/event/repository.go: same as token
- internal/testutil/postgres.go: `_ = pgContainer.Terminate(...)` and
`_ = db.Close()` → t.Logf on cleanup error (use distinct err names to
avoid govet shadow)
funlen (1):
- cmd/canary/main.go: split `run` into `run` + `initTelemetry` +
`mountRouter` + `gracefulShutdown` (was 55 statements, now under
the 50 cap; helpers are individually well under)
govet shadow (1):
- cmd/canary/main.go: migrations check switched from `if err :=` to
`if err =` (reuses outer err whose value was already consumed) — the
inner short-decl was lexically shadowing without intent
- cmd/canary/main.go: select-case `err` renamed to `startErr` to avoid
the same shadow pattern
golines (6, all auto-fixed by `golangci-lint run --fix`):
- main.go, config.go, telemetry.go, event/repository.go, token/dto.go,
token/repository.go — long lines wrapped to 80-col
Verified post-fix: build OK, vet OK, unit tests PASS under -race,
integration tests PASS under -tags=integration -race (12s testcontainers),
golangci-lint reports 0 issues.