MODELS
Pick the right model for the job.
Every model on rtrvr.ai, compared on the dimensions that matter for a browser agent: how capable it is on complex sites, how fast it streams, what a task actually costs, and how much page it can hold. Numbers come from recent production traffic.
GPT-5.6 Terra
The strongest pick for complex sites and extensive tool calling. Premium rates — bring it in when the task earns it.
DeepSeek V4 Flash
The platform default and its most production-proven model: fast, cheap, and reliable across multi-step runs.
GLM 5.3 Flash
The lowest cost per task, with native vision. It reasons lightly — great for straightforward tasks, not the hardest pages.
Head to head
Capability is our editorial ranking from quality batteries and external benchmarks; speed, affordability, and context come from measured throughput and current credit rates.
Compare up to 5 models
Every model, side by side
Per-task costs price the same three canonical workloads under each model’s current rates, so columns compare apples to apples.
| Model | Best for | $ / 1M tokensin · cached · out | Quickcredits/task | Everydaycredits/task | Long runcredits/task | Speed | Cache hit | Context | Vision |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek | |||||||||
DeepSeek FlashFree Mode | Everyday default — fast, cheap, reliable | 0.112 · 0.0224 · 0.224 | 0.02 | 0.37 | 1.4 | 117 tok/s | 72% | 1M | — |
DeepSeek Pro | Deep multi-step reasoning · ~10× Flash cost | 1.32 · 0.044 · 3.96 | 0.32 | 4.0 | 13 | 116 tok/s | 60% | 128K | — |
| ZAI GLM | |||||||||
GLM 5.3 FlashFree Mode | Lowest cost · reads screenshots (vision) | 0.075 · 0.015 · 0.25 | 0.02 | 0.26 | 0.99 | ~110 tok/s | — | 1M | ✓ |
GLM 5.3 | Flagship open-weights · strong on hard pages | 1.4 · 0.26 · 4.4 | 0.34 | 4.8 | 18 | ~65 tok/s | — | 1M | — |
| OpenAI | |||||||||
GPT-5.6 LunaFree Mode | Frontier quality · quick tool calls · vision | 0.22 · 0.022 · 1.32 | 0.07 | 0.80 | 2.9 | 71 tok/s | 27% | 400K | ✓ |
GPT-5.6 Terra | Most capable — complex sites, heavy tool use | 2.2 · 0.22 · 13.2 | 0.73 | 8.0 | 29 | ~55 tok/s | — | 400K | ✓ |
| Gemini | |||||||||
Gemini Flash LiteFree Mode | Snappiest replies · quick lookups & summaries | 0.3 · 0.03 · 2.5 | 0.12 | 1.2 | 4.4 | 169 tok/s | 42% | 1M | ✓ |
Gemini Flash | Multimodal all-rounder · media-heavy pages | 0.75 · 0.075 · 3.75 | 0.22 | 2.6 | 9.4 | 102 tok/s | 34% | 1M | ✓ |
| Thinking Machines | |||||||||
Inkling SmallFree Mode | Compact open-weights reasoner · low cost | 0.3 · 0.06 · 1.2 | 0.08 | 1.1 | 4.1 | ~90 tok/s | — | 256K | — |
Inkling | Frontier open-weights reasoning · premium | 1 · 0.17 · 4.05 | 0.27 | 3.5 | 13 | ~60 tok/s | — | 256K | — |
| Xiaomi MiMo | |||||||||
MiMo V2.5Free Mode | Budget pick for huge pages · slower output | 0.119 · 0.00238 · 0.238 | 0.02 | 0.34 | 1.1 | 34 tok/s | 70% | 1M | — |
MiMo V2.5 Pro | Deeper reasoning on huge pages | 0.3045 · 0.00252 · 0.609 | 0.06 | 0.86 | 2.7 | ~30 tok/s | — | 1M | — |
1 credit = $0.01. Per-task credits price the same canonical token bundle under every model’s rates — see the task shapes below. Speed and cache-hit are medians from recent production traffic; ~values are estimates for newer tiers.
What a “task” means here
We aggregated every LLM round of every production task into per-task token bundles. These three shapes are that fleet’s 25th percentile, median, and 75th percentile.
Quick lookup
One round — read a page, answer a question
- Rounds
- 1
- Input
- 1.5K tokens
- Cached
- 0%
- Output
- 300 tokens
Everyday task
A short plan plus an extraction or a few actions
- Rounds
- 1–2
- Input
- 50K tokens
- Cached
- 50%
- Output
- 1.5K tokens
Long agent run
Multi-step browsing with tool calls across pages
- Rounds
- 5+
- Input
- 250K tokens
- Cached
- 70%
- Output
- 6.5K tokens
Try them on your own tasks
Switch models any time from the composer in the Chrome extension or on rtrvr.ai/cloud — value tiers run free in Free Mode.