Every AI Model on Clawployees Now Has an IQ

Dan

Ask someone who is not an AI engineer to pick a model, and watch what happens. The dropdown says anthropic/claude-sonnet-4-5, moonshotai/kimi-k3, openrouter/z-ai/glm-5.2. Every one of those strings is a fact about routing, and none of them is a fact about the model. Is the first one smarter than the third? Cheaper? Better at judgement calls? The name does not say.

So people pick the one they have heard of, or the one sitting at the top of the list, and then wonder why their agent behaves the way it does. Choosing the brain your AI colleague thinks with should not require a subscription to a benchmark newsletter.

Every model on Clawployees now carries an IQ.

Human average is 100

The IQ scale is one of the very few numbers a non-specialist can read instantly. 100 is average. 130 is unusual. You already knew that before you opened this page, which is the entire reason we used it.

To be clear about what it is: a language model does not have an IQ in any literal, human sense. What the badge gives you is a single comparable figure, on a scale you already understand, for how much raw capability a model brings to a job. Higher numbers reason better and handle more ambiguity. Lower numbers are faster and cost less.

What the models actually score

Here is a cross-section of the fleet, from the current frontier down to the small open-weight models people run on their own hardware.

ModelMakerIQBand
Claude Opus 5 (max effort)Anthropic138Thinker
GPT-5.6 Sol (max)OpenAI136Thinker
Kimi K3 (max)Moonshot AI135Thinker
Claude Sonnet 5 (max effort)Anthropic130Thinker
GLM-5.2 (max)Z AI128All rounder
Gemini 3.6 Flash (high)Google127All rounder
MiniMax-M3MiniMax120All rounder
GLM-5.1 (reasoning)Z AI117All rounder
Claude 4.5 Haiku (reasoning)Anthropic108Doer
Mistral Medium 3.5Mistral108Doer
gpt-oss-120b (high)OpenAI103Doer
Llama 4 MaverickMeta95Doer

Scores as of August 2026. Where a model can be run at more than one reasoning setting, the table names the setting that was tested, because the same model genuinely performs differently depending on how hard you let it think.

Two things usually surprise people looking at this for the first time.

The first is how tight the top is. The gap between the best model in the world and the fourth best is eight points, which on this scale is about half a standard deviation. The frontier is a crowd, not a leader.

The second is how good the cheap models have become. A 108 is not a weak model. It is a capable one that will handle the overwhelming majority of what a working agent does all day, at a fraction of the cost of the model above it. The interesting decision is rarely "what is the best model". It is "how much of this job actually needs the expensive brain".

Thinker, All rounder, Doer

Some people do not want a number at all. They want a word. So each score also carries a plain-language band:

  • Thinker (130 and above): handles judgement calls and open ended work. Reach for these when the task is ambiguous, the stakes are high, or nobody can write down the rules in advance.
  • All rounder (110 to 129): a good general choice for most jobs. If you are not sure, start here.
  • Doer (below 110): fast and cheap. Best on clear, repeatable tasks where the shape of the work is known: triage, tagging, routing, drafting from a template, answering the same twenty questions.

This began as two bands, smart models against execution models, and two turned out to be wrong. A straight split drops every capable mid-range model into a bucket whose implicit label is "does no thinking", which is both false about the models and insulting to the customer who picked one. Someone running a Doer on a high-volume, well-defined job has made a good decision, not a compromise.

A practical pattern: give the agent that talks to your customers a Thinker, and give the agent that files, tags, and routes a Doer. Most teams do not need one model. They need the right one per job, which is much easier to reason about once every option carries a number.

The number holds still

A model's score does not drift because a competitor shipped something. If you picked an All rounder in March, it is still an All rounder in September, and it will not be quietly demoted overnight because of a press release from a company you do not use. A model's number changes when the model is re-tested, and not otherwise.

That matters more than it sounds. A number that moves on its own is a number you stop trusting, and a rating nobody trusts is worse than no rating at all.

When we do not know, we say nothing

If we do not have a score for a model, it gets no badge. Not a dash, not a zero, not an estimate, not a plausible guess based on its neighbours. Nothing at all.

This is deliberate, and it is the rule we care most about. A wrong number here would be wrong at the exact moment you are making a decision, and you would have no way of knowing. So the badge stays silent unless the figure behind it is real. A missing badge means we do not know yet, and it means precisely that.

Where you will see it

Everywhere a model appears: the model picker, the agent creation wizard, the agent edit and detail screens, chat, templates, the inference catalogue, your own provider settings, analytics breakdowns, and archived agent snapshots. Hover any badge for the band and what it means.

What this is not

It is not a claim that models have minds. It is also not a promise that a 132 beats a 128 on your particular job.

Capability is one axis of several. A lower-scoring model with a huge context window, a much lower price, or better behaviour when calling your tools may well be the correct choice for the work in front of you, and the number will not tell you that. Treat it as the first question you ask about a model, not the last.

What it replaces is worse: a string of routing metadata, and a guess. For anyone whose job is not reading model cards, that is the difference between choosing and hoping.

Model intelligence scores are derived from public test results published by Artificial Analysis.

Related Articles