The Aggr-Score

Every answer we compare is scored 0–10 by an independent AI judge across five weighted dimensions. Here is exactly how — and where to stay skeptical.

When you ask aggrai a question, several models answer it independently. The Aggr-Score is a single 0–10 quality rating we attach to each answer, so you can see which one held up without reading every word of all of them. It is a judgement, not a measurement — designed to be transparent, so you can decide how much to trust it.

An independent judge

The models that answer your question do not score themselves. A separate model — Claude Haiku — reads every answer and rates each one against the same fixed rubric. It judges on the content of the answer, and applies identical criteria to every model in the comparison. That independence is the point: a fair comparison needs a judge with no stake in the result.

Five dimensions

Each answer is scored on five dimensions. Every dimension is itself the mean of three finer sub-criteria, then combined with the weights below. The weights reflect what we think matters most in a trustworthy answer — being right outranks being polished.

DimensionWeightWhat it measures
Accuracy30%Are the claims factually correct, with no fabrication or hallucination? The single most consequential dimension — and the only one with a hard safety cap (below).
Completeness25%Does it actually answer the question asked, and the intent behind it, rather than a narrower or adjacent version?
Calibration20%Does the answer's confidence match its evidence? Hedging where it should, committing where it can — epistemic honesty, not false certainty.
Clarity15%Is it well-structured and appropriately concise? No padding, no wall of text — the right length for the question.
Insight10%Does it add a non-obvious angle, useful framing, or a consideration the reader wouldn't have reached alone? The 'nice to have'.

How the number is calculated

Each dimension is scored 0–5. We take the weighted average of the five (Accuracy 30%, Completeness 25%, Calibration 20%, Clarity 15%, Insight 10%), which gives a 0–5 result, then double it to the familiar 0–10 scale you see on the card. So a genuinely strong answer across the board lands in the high 8s and 9s; a mixed one sits mid-scale.

The fatal-flaw cap

One rule overrides the arithmetic. If an answer's Accuracy is critically low — confidently wrong, or fabricating facts — its overall score is capped at 4.0 / 10, no matter how complete, clear, or insightful it is. A beautifully written wrong answer is not a good answer, and we won't let strong presentation launder a factual failure into a high score. When the cap is applied, the card is marked Limited so you know why the number is low.

The contribution bar is a different thing

Don't confuse the Aggr-Score with the "Where the summary came from" bar. The score rates each model's own answer. The contribution bar shows how much of each model's content ended up shaping aggrai's combined answer — the synthesis at the top. A model can score well and still contribute little to the synthesis if another said the same thing better, and vice-versa.

What the score is — and isn't

  • It's one judge's read, not ground truth. A capable model applying a consistent rubric is a strong signal, but it is still a judgement and can be wrong — especially on specialist or contested topics.
  • It's relative within a comparison. Scores are most useful for comparing the answers in front of you, not as an absolute grade of a model in general.
  • A truncated answer scores what was returned. If a provider cut an answer off at our length cap, its score reflects the partial text — that's marked Truncated, not a verdict on the model.
  • The judge can't see the future. On time-sensitive questions, an answer can be well-scored and still out of date. When we ground a question with live web search, the sources are shown so you can check.

We show the full per-dimension breakdown on every answer precisely so you can disagree with the headline number. The score is a fast way in — the answers, and your own judgement, are the point.

Curious which models we compare? See the model list, or read more in the docs.