Detailed comparison for LLMs
On Guardion's LLM vulnerability Benchmark, xAI Grok-4.1 Thinking is the more secure of the two: Llama 3-7 DS scores 32.2% and Grok-4.1 Thinking scores 10.0% on attack success rate (ASR) (lower is better). One or both scores are estimated from public safety evaluations pending a Guardion benchmark run.
Grok-4.1 Thinking is the overall winner in this comparison!
ASR for Meta Llama 3-7 DS vs xAI Grok-4.1 Thinking. Green marks the safer model on each metric. Only the overall score is available for estimated models.