Detailed comparison for LLMs
On Guardion's LLM vulnerability Benchmark, Anthropic Claude 4.0 Sonnet is the more secure of the two: Claude 4.0 Sonnet scores 13.6% and GPT-4o Mini scores 26.7% on attack success rate (ASR) (lower is better).
Claude 4.0 Sonnet is the overall winner in this comparison!
ASR for Anthropic Claude 4.0 Sonnet vs OpenAI GPT-4o Mini. Green marks the safer model on each metric.
Outward is better on every axis.
On Guardion's LLM vulnerability Benchmark, Anthropic Claude 4.0 Sonnet is the more secure of the two: Claude 4.0 Sonnet scores 13.6% and GPT-4o Mini scores 26.7% on attack success rate (ASR) (lower is better).
Claude 4.0 Sonnet has a 13.6% ASR and GPT-4o Mini has a 26.7% ASR — the share of adversarial prompts that succeed across zero-shot, TAP, and Crescendo attacks. Lower is safer.
Both were red-teamed with the HarmBench framework across zero-shot, TAP (Tree of Attacks with Pruning), and Crescendo multi-turn attacks, scored by Attack Success Rate.