
Google Gemma 4 26B A4B is an open-weight Mixture-of-Experts language model from Google's Gemma family, built for developers who want near-frontier quality at a price they can run locally. The 26B is total parameters on disk; A4B is the roughly 3.88 billion parameters activated per token, so it behaves like a 4B dense model at inference. Community MLX work pushed straight-in-memory inference about 130 percent faster on Apple Silicon, which Google itself rounded to 2x.
Core Features
- MoE architecture with 128 fine-grained experts and top-8 routing per token, activating about 3.88B params.
- Released under Google's Gemma open-weight terms, so you can self-host and fine-tune.
- Runs in roughly 14 GB at 4-bit, fitting a 24 to 32 GB MacBook Pro or workstation.
- Benefits from community MLX kernels: about 573 tokens per second decode and 7,000 plus tokens per second prefill on tuned Mac setups.
- Available through Ollama, llama.cpp, MLX, and mlx-lm with quantization-aware training preserved reasoning.
Use Cases
- Privacy-sensitive local chat and note summarization on an M-series Mac.
- Offline coding assist where you do not want prompts leaving the machine.
- Benchmarking and research on small active-parameter MoE behavior.
Pricing
Gemma 4 26B A4B is free to download and run under Google's Gemma open-weight license; the only cost is the hardware you already own. Cloud inference is available through Google and partner endpoints at standard token rates.
Pros and Cons
- Pro: Open weights and small active parameters make it one of the few frontier-class MoEs you can run on a single Mac.
- Con: It still lags the largest dense models on hard reasoning, so it is a local workhorse rather than a flagship.
Our Take: Best for Mac and edge developers who want a real open MoE without a data center; the trade-off is that it trails the biggest dense models on the hardest reasoning, so treat it as a local workhorse, not a flagship replacement.
Browse related models in our AI Engine/Model category.




