Free IDE extension for risk-free vibe coding — keep secrets out of AI.
Benchmark methodology
The benchmark runs the same production classifier the API serves — not a research model. Reproduce every number from a clean checkout. Hover between the precision-first gate and the honest caveats below before you cite it.
The runner loads every JSONL row from the phase-9 public corpus and evaluates it with the production guard detector through lib/guard/analyze.ts. The corpus is synthetic and maintained in-repo; it is not an independent third-party dataset.
An attack counts as detected when it receives a protective action other than ALLOW. A false positive is counted when a benign control receives any protective action. Binary, reproducible, no human judgement in the loop.
Node version, OS platform, CPU model and core count, and memory are recorded in the report so results are reproducible. Latency is measured locally on CPU with no network call and no extra LLM round-trip.