DataLearner logo

Gemma 4 E4BvsQwen3.5-Omni-Flash

Across 6 shared benchmarks, Qwen3.5-Omni-Flash leads overall: Gemma 4 E4B wins 1, Qwen3.5-Omni-Flash wins 5, with 0 ties and an average score difference of -15.38.

DeepMind
Gemma 4 E4B

DeepMind · 2026-04-02 · Multimodal model

阿里巴巴
Qwen3.5-Omni-Flash

阿里巴巴 · 2026-03-30 · Multimodal model

Gemma 4 E4B1 win(17%)(83%)5 winsQwen3.5-Omni-Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Agent Level Benchmark

Qwen3.5-Omni-Flash 2/2
BenchmarkGemma 4 E4BQwen3.5-Omni-FlashDiff
τ²-Bench - Telecom26225 / 264Normal (With Tools)84.5094 / 264Normal (With Tools)-58.50
Terminal Bench Hard7.60190 / 244Normal (With Tools)8.30186 / 244Normal (With Tools)-0.70

General Evaluation

Qwen3.5-Omni-Flash 1/1
BenchmarkGemma 4 E4BQwen3.5-Omni-FlashDiff
GPQA Diamond54.90399 / 462Normal (No Tools)74.20289 / 462Normal (No Tools)-19.30

General Knowledge

Qwen3.5-Omni-Flash 1/1
BenchmarkGemma 4 E4BQwen3.5-Omni-FlashDiff
HLE4.80488 / 563Normal (No Tools) · Text only7.60427 / 563Normal (No Tools) · Text only-2.80

Instruction Following

Gemma 4 E4B 1/1
BenchmarkGemma 4 E4BQwen3.5-Omni-FlashDiff
IF Bench40.50214 / 282Normal (No Tools)38230 / 282Normal (No Tools)+2.50

Multimodal Understanding

Qwen3.5-Omni-Flash 1/1
BenchmarkGemma 4 E4BQwen3.5-Omni-FlashDiff
MMMU-Pro51.20197 / 227Normal (No Tools)64.70154 / 227Normal (No Tools)-13.50

Specs

FieldGemma 4 E4BQwen3.5-Omni-Flash
PublisherDeepMind阿里巴巴
Release date2026-04-022026-03-30
Model typeMultimodal modelMultimodal model
ArchitectureDenseDense
Parameters8BNot available
Context length128K256K
Max output8K8K

Summary

  • Gemma 4 E4Bleads in:Instruction Following (1/1)
  • Qwen3.5-Omni-Flashleads in:Agent Level Benchmark (2/2), General Evaluation (1/1), General Knowledge (1/1), Multimodal Understanding (1/1)

On average across the 6 shared benchmarks, Qwen3.5-Omni-Flash scores 15.38 higher.

Largest single-benchmark gap: τ²-Bench - Telecom — Gemma 4 E4B 26 vs Qwen3.5-Omni-Flash 84.50 (-58.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.