IBM Granite 4.2-8B - Open Reasoning Model for Enterprise Agents

IBM Granite 4.2-8B is an open-weight reasoning model from IBM that adds step-by-step chain-of-thought thinking, tool calling, and 128K context for enterprises and developers who want to run agents on their own hardware under a permissive license. Released on August 25, 2026, it is the mid-size member of IBM's three-model Granite 4.2 family, sitting between a 3B edge model and a 30B flagship.

Core Features

  • Native reasoning: built-in chain-of-thought with full, low-effort, and off modes so you can trade latency for depth per query.
  • Long context: 128K tokens natively, extendable to 512K, for long documents and multi-step agentic workflows.
  • Reasoning-augmented tool calling: the model decides which tools to invoke and why, producing more reliable function calls.
  • Apache 2.0 license with cryptographically signed weights, deployable on Ollama, vLLM, SGLang, and LM Studio.
  • Multilingual coverage across 70-plus languages, with 25-plus tested in depth.

Use Cases

  • Enterprise support agents where data must stay inside the customer's VPC.
  • On-prem coding and terminal agents that navigate codebases without sending code off-site.
  • Cost-sensitive production inference where a small dense model beats a large API bill.

Pricing

The 8B weights are free under Apache 2.0 for self-hosting. Hosted, Replicate runs the 8B at $0.06 per million input tokens and $0.25 per million output tokens; IBM watsonx.ai uses pay-as-you-go per-million-token pricing plus a Standard plan around $1,050 per month.

Our Take

Best for regulated teams that need a transparent, licensable open model they fully control; the trade-off is that Granite still trails Qwen on raw coding benchmarks, so pick it for trust and licensing over leaderboard peaks. Compare with other AI models and engines.

FacebookXWhatsAppEmail