Post

HN
Hacker News (Newest)

Show HN: I benchmarked Gemma 4 E2B โ€“ the 2B model beat the 12B on multi-turn

Google's newest 2B model tested across 10 enterprise task suites against Gemma 2 2B, Gemma 3 4B, Gemma 4 E4B, and Gemma 3 12B. Run locally on Apple Silicon.

5 Gemma models (2B, 3-4B, E2B, E4B, 12B) ยท 10 enterprise test suites ยท ~120 test cases ยท Apple Silicon (MPS) ยท temperature 0.0 ยท deterministic runs ยท local inference via Hugging Face Transformers

After last month's deep dive on Gemma 4 E4B, I had to ask the obvious follow-up: what about its smaller sibling? Google released Gemma 4 E2B alongside E4B โ€” a 2-billion parameter model positioned as the entry point to the new architecture. Half the parameters, half the memory, presumably half the capability.

The pitch from Google is that the Gemma 4 architecture improvements aren't just about raw scale โ€” they should propagate down to the smallest variants. So I rebuilt the test harness, added the new model to the registry, and ran all ten enterprise suites against it. Then I compared the results against the Gemma 2 2B baseline (the previous-generation 2B model) and the rest of the Gemma family.

The results are surprising in ways I did not expect.

Same enterprise-relevant task suites as the E4B writeup, plus three more I added since: