Interactive scorecard
AI-Sec Tool Scorecard Builder
Pick 1–3 reviewed tools, then tune what matters for your context. Every dimension score comes from the documented review on this site — vendor documentation, source code and published evaluations — with the exact evidence sentence shown inline, so you can see why a tool ranks where it does, not just that it does.
Each dimension is scored 0–5 from the linked review, on published evidence rather than first-hand testing. 5 = strongest documented; 0 = effectively absent or a hard weakness. Higher is always better, including for cost/effort dimensions (5 = lowest cost / least effort). Scores reviewed 2026-08.
All tools & raw review scores
| Tool | Detection rate | Novel-attack resilience | Low false-positive cost | Latency fit | Integration effort | Deployment flexibility | Maintenance signal |
|---|---|---|---|---|---|---|---|
| Garak Apache 2.0 (open source) | 4 | 2 | 3 | 1 | 2 | 4 | 4 |
| Lakera Guard Commercial (SaaS; enterprise self-host) | 4 | 3 | 4 | 3 | 5 | 4 | 4 |
| Guardrails AI Apache 2.0 (open source) | 3 | 2 | 3 | 3 | 4 | 5 | 4 |
| PyRIT MIT (open source) | 4 | 3 | 3 | 3 | 4 | 4 | 5 |
| Rebuff Apache 2.0 (open source) | 4 | 2 | 3 | 3 | 3 | 5 | 3 |
| Arize Phoenix Apache 2.0 (open source) | 3 | 2 | 2 | 2 | 4 | 5 | 4 |