Gemma 4 E4BvsQwen3.5-Omni-Flash
Across 6 shared benchmarks, Qwen3.5-Omni-Flash leads overall: Gemma 4 E4B wins 1, Qwen3.5-Omni-Flash wins 5, with 0 ties and an average score difference of -15.38.
Gemma 4 E4B
DeepMind · 2026-04-02 · Multimodal model
Qwen3.5-Omni-Flash
阿里巴巴 · 2026-03-30 · Multimodal model
Gemma 4 E4B1 win(17%)(83%)5 winsQwen3.5-Omni-Flash
Benchmark scores
Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.
Agent Level Benchmark
Qwen3.5-Omni-Flash 2/2| Benchmark | Gemma 4 E4B | Qwen3.5-Omni-Flash | Diff |
|---|---|---|---|
| τ²-Bench - Telecom | 26225 / 264Normal (With Tools) | 84.5094 / 264Normal (With Tools) | -58.50 |
| Terminal Bench Hard | 7.60190 / 244Normal (With Tools) | 8.30186 / 244Normal (With Tools) | -0.70 |
General Evaluation
Qwen3.5-Omni-Flash 1/1| Benchmark | Gemma 4 E4B | Qwen3.5-Omni-Flash | Diff |
|---|---|---|---|
| GPQA Diamond | 54.90399 / 462Normal (No Tools) | 74.20289 / 462Normal (No Tools) | -19.30 |
General Knowledge
Qwen3.5-Omni-Flash 1/1| Benchmark | Gemma 4 E4B | Qwen3.5-Omni-Flash | Diff |
|---|---|---|---|
| HLE | 4.80488 / 563Normal (No Tools) · Text only | 7.60427 / 563Normal (No Tools) · Text only | -2.80 |
Instruction Following
Gemma 4 E4B 1/1| Benchmark | Gemma 4 E4B | Qwen3.5-Omni-Flash | Diff |
|---|---|---|---|
| IF Bench | 40.50214 / 282Normal (No Tools) | 38230 / 282Normal (No Tools) | +2.50 |
Multimodal Understanding
Qwen3.5-Omni-Flash 1/1| Benchmark | Gemma 4 E4B | Qwen3.5-Omni-Flash | Diff |
|---|---|---|---|
| MMMU-Pro | 51.20197 / 227Normal (No Tools) | 64.70154 / 227Normal (No Tools) | -13.50 |
Specs
| Field | Gemma 4 E4B | Qwen3.5-Omni-Flash |
|---|---|---|
| Publisher | DeepMind | 阿里巴巴 |
| Release date | 2026-04-02 | 2026-03-30 |
| Model type | Multimodal model | Multimodal model |
| Architecture | Dense | Dense |
| Parameters | 8B | Not available |
| Context length | 128K | 256K |
| Max output | 8K | 8K |
Summary
- Gemma 4 E4Bleads in:Instruction Following (1/1)
- Qwen3.5-Omni-Flashleads in:Agent Level Benchmark (2/2), General Evaluation (1/1), General Knowledge (1/1), Multimodal Understanding (1/1)
On average across the 6 shared benchmarks, Qwen3.5-Omni-Flash scores 15.38 higher.
Largest single-benchmark gap: τ²-Bench - Telecom — Gemma 4 E4B 26 vs Qwen3.5-Omni-Flash 84.50 (-58.50).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.