AI MODELS / BROWSER AUTOMATION
AI models for
browser automation.
Research sites, collect data, and fill out forms with rtrvr. Compare model speed, cost, and results to choose one for your task.
Compare cost and features
Explore the chart, then compare every model in the table. Choose a view or filter by the features you need.
Model comparison
See how models compare on benchmark score, speed, cost, and context. See the full data table.
Artificial Analysis supplies the index and API speed, recorded 23 September 2026. Index: Artificial Analysis Intelligence Index v4.3.2, maximum reasoning effort. Cost uses the everyday sample task. Models without an AA score remain in the table below.
Model data
Every model and measurement, grouped by source. AA benchmarks: 23 September 2026. Historical rtrvr data: 31 August 2026. How we measure.
All 16 columns are shown. Scroll across or jump to a group below.
| Model | Estimated model credits | Price per million tokens · USD | Artificial Analysis | Historical rtrvr sample | Supported inputs | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Quick | Everyday | Long run | Input | Cached input | Output | Cache write | Indexmax effort | Speedtokens/sec | Plan, medianseconds | Plan, 90th %ileseconds | Cache hit | Output speedtokens/sec | Contexttokens | Reads images | |
DeepSeek FlashFree Mode Collecting data across pages; reads images | 0.08 | 0.90 | 3.0 | $0.285 | $0.006 | $1.14 | — | 39.5 | 232 | 12.9 | 85.3 | 72% | 77 | 1M | Yes |
DeepSeek Pro Tasks with longer reasoning | 0.25 | 3.2 | 11 | $1.056 | $0.035 | $3.168 | — | 36.0 | 67 | — | — | 60% | 116 | 128K | No |
Gemini Flash Pages with images and PDFs | 0.22 | 2.6 | 9.4 | $0.75 | $0.075 | $3.75 | — | Not rated | 315 | — | — | 34% | — | 1M | Yes |
Gemini Flash LiteFree Mode Quick lookups and summaries | 0.12 | 1.2 | 4.4 | $0.30 | $0.03 | $2.50 | — | 22.2 | 367 | — | — | 42% | 209 | 1M | Yes |
GPT-6 LunaFree Mode Quick questions, tool calls, and images | 0.03 | 0.35 | 1.3 | $0.10 | $0.01 | $0.50 | $0.125 | 37.3 | 162 | — | — | — | — | 1M | Yes |
GPT-6 Sol Complex sites and extensive tool use | 0.60 | 7.0 | 25 | $2.00 | $0.20 | $10.00 | $2.50 | 47.5 | 131 | — | — | — | — | 1M | Yes |
Inkling Text tasks with longer reasoning | 0.27 | 3.5 | 13 | $1.00 | $0.17 | $4.05 | — | 25.0 | 114 | — | — | — | — | 256K | No |
Inkling SmallFree Mode Low-cost text-based tasks | 0.08 | 1.1 | 4.1 | $0.30 | $0.06 | $1.20 | — | 27.8 | 199 | — | — | — | — | 256K | No |
MiMo V2.5Free Mode Long text with repeated context | 0.02 | 0.34 | 1.1 | $0.119 | $0.003 | $0.238 | — | Not rated | 63 | — | — | 70% | 34 | 1M | No |
MiMo V2.5 Pro Long documents and deeper reasoning | 0.06 | 0.86 | 2.7 | $0.305 | $0.003 | $0.609 | — | 26.0 | 48 | — | — | — | — | 1M | No |
GLM 5.3 Complex text-based tasks | 0.24 | 3.4 | 13 | $0.98 | $0.182 | $3.08 | — | 44.8 | 61 | — | — | — | — | 1M | No |
GLM 5.3 FlashFree Mode Code, tool calls, and screenshots | 0.02 | 0.32 | 1.2 | $0.09 | $0.018 | $0.30 | — | 41.8 | 64 | — | — | — | — | 1M | Yes |
Scores use Artificial Analysis Intelligence Index v4.3.2. AA tests at maximum reasoning effort; rtrvr uses a lower effort by default. Models without a score are marked “Not rated”.
How models perform on rtrvr
Speed and usage from real browser tasks in the extension, Cloud, and API. Updated hourly.
Share of agent calls
Every hour, across all browser-agent traffic- GPT-6 Luna 66%
- DeepSeek Flash 15%
- GLM 5.3 Flash 9%
- GPT-6 Sol 2%
- Other 8%
Model and provider data
All 12 columns are shown, including task outcomes. Scroll across or jump to a group below.
| Model · provider | Usage | Speed & wait times | Errors & cache | Task outcomes | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Shareof calls | First tokenmedian wait | First token90th percentile | Speedtokens/sec | Response timemedian | Errorsgeneration | Cache hit | Answeredact runs | Doneplan runs | Reviewed doneplan runs | Stuckact runs | |
GPT-6 Luna | 66% | 1.3 s | 3.0 s | 103 | 3.2 s | 0.2% | 85% | 97% | 64% | 44% | 3% |
OpenAIuser key | 72% | 1.2 s | 1.8 s | 85 | 2.4 s | 0.2% | 85% | 97% | 60% | 45% | 3% |
Amazon Bedrock | 21% | 2.7 s | 3.9 s | 136 | 3.7 s | 0.0% | 92% | — | — | — | — |
ChatGPTuser key | 7% | 1.6 s | 2.3 s | 98 | 5.0 s | 0.0% | 63% | — | 74% | 39% | — |
DeepSeek Flash | 15% | 3.6 s | 6.5 s | 107 | 5.6 s | 0.2% | 85% | 92% | 86% | 41% | 1% |
GMI Cloud | 100% | 3.6 s | 6.5 s | 107 | 5.6 s | 0.1% | 85% | 92% | 85% | 40% | 1% |
GLM 5.3 Flash | 9% | 3.8 s | 7.7 s | 72 | 7.1 s | 2.9% | 67% | 82% | 78% | 33% | 2% |
GMI Cloud | 63% | 5.1 s | 9.0 s | 28 | 7.5 s | 4.1% | 66% | 81% | — | — | 2% |
Azure | 37% | 2.2 s | 5.1 s | 87 | 6.2 s | 0.9% | 68% | — | 78% | 34% | — |
GPT-6 Soluser key | 2% | 1.8 s | 5.2 s | 47 | 12 s | 0.0% | 53% | — | 84% | 40% | — |
ChatGPT | 75% | 2.6 s | 5.3 s | 47 | 10 s | 0.0% | 53% | — | 83% | 41% | — |
Gemini Flashvia Google | <1% | — | — | 58 | 4.7 s | 0.0% | 2% | — | — | — | — |
Gemini Flash Litevia Googleuser key | <1% | — | — | 130 | 1.4 s | 0.0% | 42% | — | — | — | — |
Speed and wait times are medians from browser tasks. Errors count failed model responses; they exclude invalid API keys and exhausted quotas. We show a row after at least five people have used it, so provider shares may total less than 100%. A dash means too little data. “Answered” and “Done” reflect what the agent reported. We do not independently verify each result.
Compare providers over time
Choose a model below to compare its providers, hour by hour.
Output speed
Output tokens per second, median- OpenAIuser key85 tok/s
- Amazon Bedrock136 tok/s
- ChatGPTuser key98 tok/s
First token
Time to the first output token, median- OpenAIuser key1.2 s
- Amazon Bedrock2.7 s
- ChatGPTuser key1.6 s
Round time
Median time for one model response- OpenAIuser key2.4 s
- Amazon Bedrock3.7 s
- ChatGPTuser key5.0 s
Failed responses
Calls that failed to generate- OpenAIuser key0.2%
- Amazon Bedrock0.0%
- ChatGPTuser key0.0%
Which model should you try?
Start with one of these Free Mode models in the extension or Cloud. Free Mode includes sponsored cards and daily fair-use limits. See plans and limits
Answer questions about a page
Use GPT-6 Luna for questions and tool calls. Try Gemini Flash Lite when output speed matters most. Compare current response times in the live results above.
Collect data across many pages
Try DeepSeek Flash for collecting a dataset, filling a spreadsheet, or working through a list. It supports screenshots. Check live round times if you need a quick response at every step.
Read screenshots, PDFs, and charts
These models accept images. Choose one when the answer depends on a chart, screenshot, or scanned document.
Read long pages and documents
Both support a one-million-token context window. GLM 5.3 Flash also reads images. MiMo V2.5 offers a larger discount on cached input.
Spend fewer credits on model tokens
GLM 5.3 Flash has the lowest estimated model cost for the everyday sample workload. It also accepts images. Browser and proxy usage are billed separately on paid runs.
Where the numbers come from
Live results show how models run on rtrvr. Benchmarks and sample costs help you compare them.
Real browser tasks
We update these results hourly. Each model runs a different mix of tasks. The agent reports whether it finished; we do not independently check each result. We omit small samples.
Capability and output speed
Artificial Analysis supplies the index and first-party API speeds, recorded 23 September 2026. The index is a general benchmark, not a browser-task success rate.
Three sample tasks
These examples use token counts from real tasks through 31 August 2026. Estimates cover model tokens. Browser, proxy, and other usage can add to the total. Displayed rates use GMI Cloud for GLM, DeepSeek, and MiMo; Amazon Bedrock for GPT-6; Thinking Machines for Inkling; and Google for Gemini.
Quick lookup
Read a page and answer a question
- Rounds
- 1
- Input
- 1.5K tokens
- Cached
- 0%
- Output
- 300 tokens
Everyday task
Collect data or take a few actions
- Rounds
- 1–2
- Input
- 50K tokens
- Cached
- 50%
- Output
- 1.5K tokens
Long agent run
Complete a task across several pages
- Rounds
- 5+
- Input
- 250K tokens
- Cached
- 70%
- Output
- 6.5K tokens
Try a model on your task.
Choose a model in Cloud. Run the same task with another model to compare results.