Skip to content
ORIEL
Method

Method

Oriel does not run evaluations. It records and presents them. Everything here comes from Artificial Analysis — below is what the numbers mean and where to read them carefully.

Where the data comes from

Every figure comes from the public Artificial Analysis API, fetched once a day at 05:00 Beijing time (21:00 UTC) and committed straight into this site's repository. Oriel runs no benchmarks of its own and applies no weighting, correction, or recalibration — the numbers here are the numbers upstream published.

The intelligence index

The intelligence index is Artificial Analysis's composite of several benchmarks, currently v4.3. One thing matters above all: scores from different versions are not comparable. Upstream recalibrates every model when the version changes, so trend lines here break at version boundaries rather than joining into a smooth but meaningless curve. Score movements across a version change are excluded from the changes feed for the same reason.

How value is calculated

Value = mean of available benchmark scores ÷ price per 1M output tokens

This one is Oriel's own derivation, not upstream data. It is computed only when a model has at least two benchmark scores and an output price — with a single score the ratio is too noisy, and a cheap small model measured on one benchmark would top the list on nothing. It measures points per dollar. It does not tell you whether a model fits your work.

Coverage

Of 673 language models, only some are measured on any given metric: intelligence 664, coding 259, agentic 156, pricing 444, performance 333. A dash means not tested — not tested badly. Those are different claims, so selecting a dimension removes models without that metric from the board instead of leaving you to read several hundred dashes.

History

Daily snapshots start on Jul 23, 2026 and now cover 62 days. Earlier days were backfilled from this repository's git history — before that, each fetch simply overwrote the previous one. The record is still short: trend charts and the changes feed need weeks before they say much. Where there aren't enough points, they say so rather than drawing a misleading line.

Arena rankings

Media boards are Elo ratings from head-to-head human votes, with a 95% confidence interval. Where error bars overlap, the ranking gap carries no statistical meaning — 3rd and 7th place may be indistinguishable. The bars are drawn so you can see that, instead of being handed a clean column of ranks that hides it.

Known limits

The speech-to-text word error rate index is published to one decimal place, leaving many models tied at the same value; that board cannot actually order them, and the page marks how many are tied. Performance figures (throughput, latency) are medians measured at a particular time over a particular route, and move with provider load and your location. Prices cover base token billing only — no volume discounts, committed-use rates, or caching strategy.