Detailed comparison for LLMs
On Guardion's LLM vulnerability Benchmark, Cohere Command R is the more secure of the two: Command R scores 32.9% and GPT-4.1 scores 34.9% on attack success rate (ASR) (lower is better).
Command R is the overall winner in this comparison!
ASR for Cohere Command R vs OpenAI GPT-4.1. Green marks the safer model on each metric.
Outward is better on every axis.
On Guardion's LLM vulnerability Benchmark, Cohere Command R is the more secure of the two: Command R scores 32.9% and GPT-4.1 scores 34.9% on attack success rate (ASR) (lower is better).
Command R has a 32.9% ASR and GPT-4.1 has a 34.9% ASR — the share of adversarial prompts that succeed across zero-shot, TAP, and Crescendo attacks. Lower is safer.
Both were red-teamed with the HarmBench framework across zero-shot, TAP (Tree of Attacks with Pruning), and Crescendo multi-turn attacks, scored by Attack Success Rate.