
IBM Granite 4.2-8B is an open-weight reasoning model from IBM that adds step-by-step chain-of-thought thinking, tool calling, and 128K context for enterprises and developers who want to run agents on their own hardware under a permissive license. Released on August 25, 2026, it is the mid-size member of IBM's three-model Granite 4.2 family, sitting between a 3B edge model and a 30B flagship.
Core Features
- Native reasoning: built-in chain-of-thought with full, low-effort, and off modes so you can trade latency for depth per query.
- Long context: 128K tokens natively, extendable to 512K, for long documents and multi-step agentic workflows.
- Reasoning-augmented tool calling: the model decides which tools to invoke and why, producing more reliable function calls.
- Apache 2.0 license with cryptographically signed weights, deployable on Ollama, vLLM, SGLang, and LM Studio.
- Multilingual coverage across 70-plus languages, with 25-plus tested in depth.
Use Cases
- Enterprise support agents where data must stay inside the customer's VPC.
- On-prem coding and terminal agents that navigate codebases without sending code off-site.
- Cost-sensitive production inference where a small dense model beats a large API bill.
Pricing
The 8B weights are free under Apache 2.0 for self-hosting. Hosted, Replicate runs the 8B at $0.06 per million input tokens and $0.25 per million output tokens; IBM watsonx.ai uses pay-as-you-go per-million-token pricing plus a Standard plan around $1,050 per month.
Our Take
Best for regulated teams that need a transparent, licensable open model they fully control; the trade-off is that Granite still trails Qwen on raw coding benchmarks, so pick it for trust and licensing over leaderboard peaks. Compare with other AI models and engines.




