MODELS
Pick the right model for the job.
Every model on rtrvr.ai, compared on the dimensions that matter for a browser agent: capability on complex sites, speed, what a task costs, and how much page it holds. Measured from production traffic through 31 August 2026.
GLM 5.3
Tops our lineup on the independent AA index at 59.5. The flagship to reach for when the page is genuinely hard.
GPT-5.6 Luna
The Gemini tiers push more tokens per second, but this is the quickest one that also scores — and it stalls least of anything we measure. Free on every plan.
DeepSeek V4.1 Flash
Dependable across long multi-step runs where nobody is timing each step, and since V4.1 it reads screenshots too.
GLM 5.3 Flash
Scores 57.5 on the AA index — ahead of every paid tier here. Writes clean plan code, thinks in few tokens, and reads screenshots, at the cheapest rates we publish.
Head to head
Every axis is measured, none of them by us marking our own homework.
Compare up to 5 models
- Capability
- AA index, across our lineup
- Speed
- AA output tok/s
- Affordability
- Cost per everyday task
- Context
- Window we serve
Capability is the Artificial Analysis Intelligence Index v4.1.1, read 31 August 2026; Speed is AA timing each first-party API. A full spoke means best of this lineup, not best available: the top model anywhere scores 63, and raw scores sit in the table below. AA tests at maximum reasoning effort, so capability is a ceiling and rtrvr's default setting resolves lower.
Which free model should I use?
All of these are free on every plan and in Free Mode. Pick by job, not by budget.
Quick back-and-forth — one question per page, timed drills
Both answer in seconds and, just as important, rarely stall. The pair to be on when you are waiting on every question.
Long multi-step runs — crawl a site, fill a sheet, work a list
The volume model — the most production-proven tier here, and since V4.1 it reads screenshots too. Its slow tail is the platform’s longest, which matters little when nobody is timing each step.
Pages with screenshots, PDFs, or charts
These read images natively. MiMo and Inkling cannot — on a scanned PDF they are working blind.
Huge pages and long documents — a million tokens of context
Both take a 1M-token context and run free. GLM 5.3 Flash is the stronger of the two and reads images; MiMo V2.5 has the deepest cache discount on repeat runs. MiMo V2.5 Pro is 1M too, on paid plans.
Stretching credits as far as they go
The cheapest rates on the platform, and it does not skimp: top AA score of the free tiers, clean plan code, and it reads images. No production latency record yet.
Every model, side by side
Per-task costs price the same three token bundles under each model’s current rates.
| Model | AA indexmax effort | Best for | $ / 1M tokensin · cached · out | Quickcredits/task | Everydaycredits/task | Long runcredits/task | SpeedAA tok/s | Plan roundours · p50 · p90 | Cache hit | Context | Vision |
|---|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek | |||||||||||
DeepSeek FlashFree Mode | 51.8 | Proven at volume · reads screenshots · long multi-step runs | 0.3 · 0.006 · 1.2 | 0.08 | 0.95 | 3.1 | 108 tok/s | 12.9s · 85.3s | 72% | 1M | ✓ |
DeepSeek Pro | 53.2 | Deep multi-step reasoning · ~10× Flash cost, +1.4 AA points | 1.32 · 0.044 · 3.96 | 0.32 | 4.0 | 13 | 54 tok/s | — | 60% | 128K | — |
| ZAI GLM | |||||||||||
GLM 5.3 FlashFree Mode | 57.5 | Code smarts · quick thinker · native vision | 0.075 · 0.015 · 0.25 | 0.02 | 0.26 | 0.99 | 44 tok/s | — | — | 1M | ✓ |
GLM 5.3 | 59.5 | Flagship open-weights · strong on hard pages | 1.4 · 0.26 · 4.4 | 0.34 | 4.8 | 18 | 78 tok/s | — | — | 1M | — |
| OpenAI | |||||||||||
GPT-5.6 LunaFree Mode | 52.3 | Fastest capable tier — quick loops, tool calls, vision | 0.22 · 0.022 · 1.32 | 0.07 | 0.80 | 2.9 | 129 tok/s | 10.4s · 43.8s | 27% | 400K | ✓ |
GPT-5.6 Terra | 56.6 | Heavy tool use on complex sites — premium tier | 2.2 · 0.22 · 13.2 | 0.73 | 8.0 | 29 | 122 tok/s | 24.4s · 57.4s | — | 400K | ✓ |
| Gemini | |||||||||||
Gemini Flash LiteFree Mode | 37.4 | Snappiest replies · quick lookups & summaries | 0.3 · 0.03 · 2.5 | 0.12 | 1.2 | 4.4 | 358 tok/s | — | 42% | 1M | ✓ |
Gemini Flash | 56.0 | Fastest output measured · media-heavy pages | 0.75 · 0.075 · 3.75 | 0.22 | 2.6 | 9.4 | 315 tok/s | — | 34% | 1M | ✓ |
| Thinking Machines | |||||||||||
Inkling SmallFree Mode | 41.2 | Compact open-weights reasoner · low cost | 0.3 · 0.06 · 1.2 | 0.08 | 1.1 | 4.1 | 48 tok/s | — | — | 256K | — |
Inkling | 42.3 | Open-weights reasoner · premium rates, mid-pack score | 1 · 0.17 · 4.05 | 0.27 | 3.5 | 13 | 77 tok/s | — | — | 256K | — |
| Xiaomi MiMo | |||||||||||
MiMo V2.5Free Mode | 38.0 | Budget pick for huge pages · slower output | 0.119 · 0.00238 · 0.238 | 0.02 | 0.34 | 1.1 | 63 tok/s | — | 70% | 1M | — |
MiMo V2.5 Pro | 42.9 | Deeper reasoning on huge pages | 0.3045 · 0.00252 · 0.609 | 0.06 | 0.86 | 2.7 | 39 tok/s | — | — | 1M | — |
1 credit = $0.01, priced on the token bundles below. AA index and Speed both come from Artificial Analysis, one source per column so each one ranks — the Artificial Analysis Intelligence Index v4.1.1 at maximum reasoning effort (a ceiling, not what a default-effort run scores), and first-party output speed. Round times and cache hits are our own production medians as of 31 August 2026, so they cover only the tiers we route enough traffic to measure; hover a speed to see what we clock through rtrvr.
What a “task” means here
The 25th percentile, median, and 75th percentile per-task token bundles from production traffic.
Quick lookup
One round — read a page, answer a question
- Rounds
- 1
- Input
- 1.5K tokens
- Cached
- 0%
- Output
- 300 tokens
Everyday task
A short plan plus an extraction or a few actions
- Rounds
- 1–2
- Input
- 50K tokens
- Cached
- 50%
- Output
- 1.5K tokens
Long agent run
Multi-step browsing with tool calls across pages
- Rounds
- 5+
- Input
- 250K tokens
- Cached
- 70%
- Output
- 6.5K tokens
Try them on your own tasks
Switch models any time from the composer in the Chrome extension or on rtrvr.ai/cloud — value tiers run free in Free Mode.