OWASP LLM Top 10 (2025)
LLM Security Leaderboard
Every model is probed with the same versioned attack corpus, at temperature 0, and scored against the OWASP LLM Top 10 (2025). Reproducible, auditable, deterministic — not a one-off blog test.
Temperature 0
fixed sampling
Corpus 2026-09-17.1
immutable + hashed
Holdout-verified
anti-overfit
OWASP LLM Top 10
2025 control set
Illustrative demo data.
These scores render the leaderboard design before live measurement runs are published. They are not real measurements.
N/A controls are excluded from the composite and never penalize. PASS/PARTIAL/FAIL are per OWASP LLM Top 10 (2025).
Methodology
Deterministic, N/A-aware scoring against the OWASP LLM Top 10 (2025). Every model is probed with the same versioned attack corpus at temperature 0 with a fixed system prompt and max tokens. A control is scored only where the use-case template marks it applicable; non-applicable controls never penalize. The public score is computed over the public corpus plus a private holdout set to guard against benchmark overfitting. Each result is auditable to a timestamped, versioned ledger.
Updated 2026-09-17. Corpus 2026-09-17.1.
Want your app scored?
The public leaderboard tests models under controlled conditions. Our full AI / LLM audit runs the complete Top 10 against your actual application through a use-case template matched to how it is built.
Get the full AI / LLM audit