Detailed comparison for LLMs
On Guardion's LLM vulnerability Benchmark, OpenAI GPT-5.1 Codex is the more secure of the two: Llama 4 Behemoth scores 25.0% and GPT-5.1 Codex scores 8.5% on attack success rate (ASR) (lower is better). One or both scores are estimated from public safety evaluations pending a Guardion benchmark run.
GPT-5.1 Codex is the overall winner in this comparison!
ASR for Meta Llama 4 Behemoth vs OpenAI GPT-5.1 Codex. Green marks the safer model on each metric. Only the overall score is available for estimated models.