Commit Graph

5 Commits

Author SHA1 Message Date
CarterPerez-dev 989635e7b0 feat: add Aenebris.ML.IForest for anomaly scoring
- New Aenebris.ML.IForest module: pure Isolation Forest scorer
  implementing the Liu et al. 2008 ICDM formula
  anomaly_score = 2^(-E[h(x)] / c(n))
  where E[h(x)] is the average path length across iTrees and
  c(n) = 2*H(n-1) - 2(n-1)/n  (expected unsuccessful BST search depth).
  ITree ADT (ITreeLeaf size | ITreeSplit featIdx threshold left right)
  with leaf-size c(n) correction added to traversal depth at every
  leaf, matching the original paper. Score-only mode (accept pre-
  trained forests from Python sklearn or similar); fitting is deferred
  to a later phase. Default constants from Liu et al.: 100 trees,
  256 subsample size, max depth ceil(log2(256)) = 8.

- 24 tests covering: harmonicNumber edge cases (H(0), H(1), H(2),
  H(3), large-n asymptotic), normalizationConstant for n in {0, 1,
  2, 256}, single-split and deep-tree path length traversal,
  scoreIForest edge cases (empty forest, subsample 0, subsample 1),
  shorter-path-equals-higher-score invariant, default constants,
  Euler-Mascheroni precision.

- aenebris.cabal: expose Aenebris.ML.IForest in the library stanza.
- test/Spec.hs: import Expectation explicitly from Test.Hspec for the
  shouldBeApprox helper.

- 325 total examples passing, 0 GHC warnings on the new module.

The escalation-gate composition with the calibrated GBDT score
(per docs/research/phase-2.5-ml-synthesis.md correction over the
0.8/0.2 folk-wisdom blend) is intentionally NOT in this module --
it belongs in the downstream ML.Engine module that wires Loader +
Inference + Calibration + IForest into a single decision pipeline.
2026-04-29 00:42:18 -04:00
CarterPerez-dev e5c2f59675 feat: add Aenebris.ML.Inference and Aenebris.ML.Calibration
- New Aenebris.ML.Inference module: pure tree-walk over the unified Tree
  SoA produced by Aenebris.ML.Loader. Mirrors LightGBM C++ NumericalDecision
  exactly: NaN remapped to 0 unless MissingType=NaN; MissingType=Zero/NaN
  routes via default-left flag; '<=' predicate (not '<'); categorical
  bitmap test using cat_boundaries slice; average_output divisor for RF
  mode; sigmoid link for binary logistic objective. kZeroThreshold = 1e-35
  matches LightGBM's IsZero exactly. Defensive against NaN/Inf/negative
  on categorical paths. 35 tests covering leaf-only, numerical splits,
  missing-value semantics, categorical routing, multi-tree sums, sigmoid
  saturation, average_output, regression vs binary objectives, and
  end-to-end against the disk fixture.

- New Aenebris.ML.Calibration module: pure calibration of binary
  classifier outputs. Calibrator ADT (NoCalibrator | PlattCalibrator a b
  | IsotonicCalibrator bp). Platt scaling uses Newton's method with
  damped line search on Platt-1999 smoothed targets (N+ + 1)/(N+ + 2)
  and 1/(N- + 2); numerically stable softplus-based log-loss. Isotonic
  regression uses Pool Adjacent Violators on a stack with tie pre-
  grouping, sorted (raw, calibrated) breakpoints, linear interpolation,
  out-of-range clamping. Sub-2-sample input returns NoCalibrator. 28
  tests covering all three calibrator constructors, fit edge cases,
  Platt convergence on logistic data, isotonic correction of non-
  monotone data, tie handling, all-same-label degenerate cases, and
  lookup behavior including clamping and interpolation.

- aenebris.cabal: expose both new modules in the library stanza.

- 301 total examples passing, 0 GHC warnings on either new module.
2026-04-29 00:35:07 -04:00
CarterPerez-dev a302741b1e feat: add Aenebris.ML.Loader for LightGBM v4 model parsing
- New Aenebris.ML.Loader module: pure ByteString -> Either ParseError Ensemble
  parser for LightGBM v4 plain-text models. Header / tree-blocks / trailing-
  section state machine; converts split-arrays + leaf-arrays into the unified
  Tree SoA via unifyChild + complement; rejects non-v4, multi-class, and
  is_linear=1; extracts embedded sigmoid scale and bare average_output flag;
  runs validateEnsemble before returning.
- Frozen test fixtures under test/fixtures/ml/ (tiny_lgbm_v4.txt,
  stump_lgbm_v4.txt) hand-crafted from the v4 spec annotated example. CI runs
  against these with zero Python dependency.
- scripts/regen_lgbm_fixture.py: uv-runnable PEP 723 script that retrains a
  tiny binary GBDT and writes to a separate
  test/fixtures/ml/regen_sample_v4.txt so the frozen fixtures stay pristine.
- 30 new mlLoaderSpec tests covering happy path, header rejection, sigmoid
  extraction, average_output, tree-level rejection, feature_names with '=',
  and ParseError reporting. 238 total examples passing, 0 GHC warnings on
  Loader.
2026-04-28 17:35:33 -04:00
CarterPerez-dev ef315b072b chore: add demos for projects, update haskell-reverse-proxy modules, refresh siem assets
- Add DEMO.md and screenshots for bug-bounty-platform, hash-cracker,
  linux-cis-hardening-auditor, simple-port-scanner, simple-vulnerability-scanner,
  systemd-persistence-scanner, base64-tool, caesar-cipher, dns-lookup,
  metadata-scrubber-tool, network-traffic-analyzer, siem-dashboard
- Link DEMO.md from project READMEs
- Add .gitignore entries for DEMO-TRACKER and simple-port-scanner build dir
- Restructure haskell-reverse-proxy with DDoS, Fingerprint, ML, RateLimit, WAF,
  Geo, and Honeypot modules; drop superseded research docs and old Makefile
- Refresh siem-dashboard dashboard.png and rename alerts.png to alert-detail.png
2026-04-26 23:12:48 -04:00
CarterPerez-dev 9397b12a4a add learn folder, move hidden files, update root readme, add learn folder, add clang tidy, and format 2026-02-26 00:38:55 -05:00