Granite 4.2: IBM bets on agents you can run on your own hardware

20 September 202622 views

The new generation of Granite models advances two ideas at once: autonomous agentic scenarios and predictable deployment within the enterprise perimeter. Against the backdrop of growing interest in local LLMs, this looks like an attempt to offer businesses an alternative to external APIs — with control over data and infrastructure.

Granite 4.2: IBM bets on agents you can run on your own hardware

Why IBM moved toward agents in the first place

An enterprise client rarely buys a model for the sake of flashy demos. They need software to actually do something: triage tickets, reconcile invoices, call internal APIs, write summaries of documents. That is exactly the niche agentic logic occupies — where the model doesn't just answer a question but decides for itself which tool to call next.

Granite 4.2 in this story is not "just another big model" but IBM's attempt to round out its lineup so that it performs equally confidently in the cloud and in a customer's air-gapped environment. The bet is on a combination of three things: open weights, modest hardware requirements, and predictable behavior when calling external functions.

What the Granite lineup actually is

IBM has maintained a family of open models under the common name Granite for years now. Over that time a recognizable philosophy has taken shape: models are released in several sizes, oriented primarily toward code, working with tables, extracting data from documents, and enterprise scenarios, rather than competing on general chat benchmarks.

Size matters

The lineup is built as a ladder — from very small variants that fit on a single GPU or even run on a CPU, up to mid-range models. The point is that for every task you can pick the minimally sufficient size. For an agent that mostly routes requests and calls functions, a gigantic model is often overkill: it's more expensive at inference and slower to respond.

Hybrid architecture

In the fourth generation of the family, IBM moved to a hybrid scheme where some layers are built on Mamba and others on classic transformer attention. The practical payoff here is prosaic: long context is processed more cheaply, memory is used more economically, and throughput on the same hardware is higher. For agentic scenarios this is critical, because the context is constantly being loaded with tool call results, document fragments, and step history.

A license with no surprises

Granite models are distributed under the Apache 2.0 license. For a corporate lawyer that's boring news, and that's precisely the good news: you can fine-tune, embed into a commercial product, deploy inside your perimeter, and not report to the vendor how much the product earned.

The main argument: you can keep this in-house

Cloud APIs are convenient right up to the moment data falls under regulatory restrictions. A bank, a clinic, an industrial enterprise, or a government agency often physically cannot send some documents outside. That's where the conversation about "your own hardware" begins.

What you actually need to run it

The key here is not top-tier accelerators but a reasonable minimum. Small Granite versions come up on a single consumer card, in a container on a server without a GPU, or on a developer's laptop. Mid-range sizes already require several cards or quantization. Deployment usually goes through standard tools like vLLM or Ollama, which removes the problem of being locked into a vendor's exclusive stack.

Security and predictability

The second layer of the argument is control. When the model lives in your environment, you decide which logs to write, which data goes into the prompt, what gets cached, and how long it's stored. For agents this is especially important: they can perform actions, not just generate text, and each such action should ideally pass through your own rules and audit.

IBM here sells not just the model but also the surrounding tooling: the watsonx platform, a set of "guardian" models for filtering unwanted content, and tools for fine-tuning on your own data. The model in this scheme is a detail — albeit a central one.

What an agentic scenario looks like in practice

The idea of an agent is simple: the model receives a task, decides what data it's missing, calls the needed tool, reads the response, and continues until it reaches a result. The difference between "just a chat" and an agent is the presence of a loop and permissions.

Typical roles these models are tuned for:

  • Router. Parses an incoming request and decides which scenario to hand it to.
  • Extractor. Pulls fields from invoices, contracts, statements, and puts them into a structure.
  • Executor. Calls APIs of internal systems: creates a ticket, updates a status, sends an email.
  • Reviewer. Checks the result of the previous step and decides whether it can be accepted.

A small model at each of these steps is often more advantageous than one big one: cheaper, faster, easier to test, and — importantly — easier to replace if the quality stops being acceptable.

What to look at and where the limitations are

Open models are not a magic pill, and an honest conversation about limitations is more useful than marketing.

Size versus reasoning quality

Small models are economical, but on complex multi-step tasks they still lag behind large ones. The compromise is usually sought architecturally: split roles among several small models, add deterministic checks in code, and leave the final decision to a human.

The ecosystem around it

Open weights mean the model moves easily between frameworks — from LangChain to custom orchestrators. The flip side: there are fewer ready-made "off-the-shelf" solutions for a specific industry than with proprietary platforms, and part of the integration work falls on the customer's team.

Hardware and total cost of ownership

Your own hardware means not only savings on API calls but also capital expenses, electricity, cooling, and the people who maintain it all. For a small company, the cloud will almost always be cheaper; for a large one with sensitive data, the opposite is true — and that threshold shifts depending on volumes.

Evaluating quality

The main trap is measuring an agent with general benchmarks. A working agent is tested on its own scenarios: a set of real tasks with a known correct answer, resilience to garbage input, and behavior when a called tool returns an error. Without such a set, any number from a release note remains just a number from a release note.

What follows from this

IBM's response to the current state of the market looks logical. While some chase peak quality in the cloud, others fill the niche where control, cost per token, and the ability to deploy everything inside the perimeter matter more. Open weights plus modest resource requirements plus a focus on tool calling — that's exactly the combination an enterprise customer tired of getting approval to send data outside needs.

The key question is not whether the new version will beat the leaders on general tests. The question is whether the team has the discipline to build a proper agentic loop around the model: with limited permissions, audit, tests, and a clear rollback. If so, keeping such an agent on your own hardware becomes a perfectly workable strategy rather than a compromise.

Frequently asked questions

Granite 4.2: IBM bets on agents you can run on your own hardware