Mistral Mixtral 8x7B vs OpenAI GPT OSS 120B

Detailed comparison for LLMs

MistralOpenAI

On Guardion's LLM vulnerability Benchmark, OpenAI GPT OSS 120B is the more secure of the two: Mixtral 8x7B scores 38.0% and GPT OSS 120B scores 15.0% on attack success rate (ASR) (lower is better). One or both scores are estimated from public safety evaluations pending a Guardion benchmark run.

Head-to-Head Overview

GPT OSS 120B is the overall winner in this comparison!

Attack Success Rate (lower is safer)

ASR for Mistral Mixtral 8x7B vs OpenAI GPT OSS 120B. Green marks the safer model on each metric. Only the overall score is available for estimated models.

Overall (ASR)

Mixtral 8x7B
38.0%
GPT OSS 120B
15.0%

TAP Attack Method (ASR)

Mixtral 8x7B
66.2%
GPT OSS 120B
100.0%

Crescendo Attack Method (ASR)

Mixtral 8x7B
24.3%
GPT OSS 120B
100.0%

Zero-Shot (ASR)

Mixtral 8x7B
23.6%
GPT OSS 120B
100.0%

Key Highlights

  • OpenAI GPT OSS 120B has a lower Overall (ASR).
  • Mistral Mixtral 8x7B has a lower TAP Attack Method (ASR).
  • Mistral Mixtral 8x7B has a lower Crescendo Attack Method (ASR).
  • Mistral Mixtral 8x7B has a lower Zero-Shot (ASR).

Security Profile

Outward is better on every axis.

OverallTAPCrescendoZero-Shot
Mixtral 8x7B
GPT OSS 120B
Full security profile
Mistral Mixtral 8x7B
Full security profile
OpenAI GPT OSS 120B

Frequently asked questions

Is Mistral Mixtral 8x7B or OpenAI GPT OSS 120B more secure?

On Guardion's LLM vulnerability Benchmark, OpenAI GPT OSS 120B is the more secure of the two: Mixtral 8x7B scores 38.0% and GPT OSS 120B scores 15.0% on attack success rate (ASR) (lower is better). One or both scores are estimated from public safety evaluations pending a Guardion benchmark run.

What is the attack success rate (ASR) of Mixtral 8x7B vs GPT OSS 120B?

Mixtral 8x7B has a 38.0% ASR and GPT OSS 120B has a 15.0% ASR — the share of adversarial prompts that succeed across zero-shot, TAP, and Crescendo attacks. Lower is safer.

How were Mixtral 8x7B and GPT OSS 120B tested?

Both were red-teamed with the HarmBench framework across zero-shot, TAP (Tree of Attacks with Pruning), and Crescendo multi-turn attacks, scored by Attack Success Rate.

Related Comparisons