Commit Graph

2 Commits

Author SHA1 Message Date
CarterPerez-dev d9dd59db2a fix: close audit-pass-1 remaining MAJOR + quick MINOR (Findings 14, 19, 20, 28, 29)
- Geo.hs (Finding 14): geoAsnCounts is now bounded by
  defaultGeoAsnCountCap = 200_000. capAsnCounts is called inside
  bumpAsnCounter; on overflow, the entry with the oldest
  awWindowStart is evicted via Data.List.minimumBy + Data.Ord.comparing.
  Closes the unbounded-Map memory growth path. Uses minimumBy and
  comparing imports added to the existing module.

- ML/Middleware.hs, WAF/Engine.hs, Honeypot.hs (Finding 19): em
  dashes (\x2014) in user-facing response bodies and the generated
  robots.txt comment replaced with ASCII hyphens, per the project's
  guardrail-safe terminology rule.

- aenebris.cabal (Finding 20): copyright field updated from
  '2025 Carter Perez' to '2026 AngelaMos' to match the file headers.

- ML/IForest.hs (Finding 28): pathLength now respects a hard
  maxIForestDepth = 64 cutoff. Beyond that depth the function
  returns currentDepth + c(1) (= currentDepth) without further
  recursion, so a pathological tree built by a buggy fitter cannot
  blow the stack. Standard iForest depth is ceil(log2(256)) = 8, so
  the cap leaves >>8 generous headroom.

- ML/Model.hs (Finding 29): validateCategoricalNode now also
  rejects categorical thresholds that are not whole numbers. A
  threshold of 1.5 would silently floor to 1 today; with this
  change it is reported as a clear validation error instead.
  Real LightGBM never writes non-integer cat indices, so this is
  defense-in-depth against malformed exporters.

Build clean, 358 examples passing, 0 failures.
2026-04-29 02:15:14 -04:00
CarterPerez-dev 989635e7b0 feat: add Aenebris.ML.IForest for anomaly scoring
- New Aenebris.ML.IForest module: pure Isolation Forest scorer
  implementing the Liu et al. 2008 ICDM formula
  anomaly_score = 2^(-E[h(x)] / c(n))
  where E[h(x)] is the average path length across iTrees and
  c(n) = 2*H(n-1) - 2(n-1)/n  (expected unsuccessful BST search depth).
  ITree ADT (ITreeLeaf size | ITreeSplit featIdx threshold left right)
  with leaf-size c(n) correction added to traversal depth at every
  leaf, matching the original paper. Score-only mode (accept pre-
  trained forests from Python sklearn or similar); fitting is deferred
  to a later phase. Default constants from Liu et al.: 100 trees,
  256 subsample size, max depth ceil(log2(256)) = 8.

- 24 tests covering: harmonicNumber edge cases (H(0), H(1), H(2),
  H(3), large-n asymptotic), normalizationConstant for n in {0, 1,
  2, 256}, single-split and deep-tree path length traversal,
  scoreIForest edge cases (empty forest, subsample 0, subsample 1),
  shorter-path-equals-higher-score invariant, default constants,
  Euler-Mascheroni precision.

- aenebris.cabal: expose Aenebris.ML.IForest in the library stanza.
- test/Spec.hs: import Expectation explicitly from Test.Hspec for the
  shouldBeApprox helper.

- 325 total examples passing, 0 GHC warnings on the new module.

The escalation-gate composition with the calibrated GBDT score
(per docs/research/phase-2.5-ml-synthesis.md correction over the
0.8/0.2 folk-wisdom blend) is intentionally NOT in this module --
it belongs in the downstream ML.Engine module that wires Loader +
Inference + Calibration + IForest into a single decision pipeline.
2026-04-29 00:42:18 -04:00