rtrvr.ai

AI MODELS / BROWSER AUTOMATION

AI models for
browser automation.

Research sites, collect data, and fill out forms with rtrvr. Compare model speed, cost, and results to choose one for your task.

01COMPARE MODELS

Compare cost and features

Explore the chart, then compare every model in the table. Choose a view or filter by the features you need.

Model comparison

See how models compare on benchmark score, speed, cost, and context. See the full data table.

Farther out means a higher benchmark score, faster output, lower cost, or more context. Each axis has its own scale; the shape is not an overall score.
Select up to 5 models

3 of 5 selected. Check a model to add it to the chart.

Artificial Analysis supplies the index and API speed, recorded 23 September 2026. Index: Artificial Analysis Intelligence Index v4.3.2, maximum reasoning effort. Cost uses the everyday sample task. Models without an AA score remain in the table below.

Model data

Choose a table view

Every model and measurement, grouped by source. AA benchmarks: 23 September 2026. Historical rtrvr data: 31 August 2026. How we measure.

Filter models
12 of 12 models

All 16 columns are shown. Scroll across or jump to a group below.

Jump to columns
All model prices, sample task costs, independent benchmarks, historical measurements, and input features
ModelEstimated model creditsPrice per million tokens · USDArtificial AnalysisHistorical rtrvr sampleSupported inputs
QuickEverydayLong runInputCached inputOutputCache writeIndexmax effortSpeedtokens/secPlan, mediansecondsPlan, 90th %ilesecondsCache hitOutput speedtokens/secContexttokensReads images
DeepSeek FlashFree Mode

Collecting data across pages; reads images

0.080.903.0$0.285$0.006$1.14—39.523212.985.372%771MYes
DeepSeek Pro

Tasks with longer reasoning

0.253.211$1.056$0.035$3.168—36.067——60%116128KNo
Gemini Flash

Pages with images and PDFs

0.222.69.4$0.75$0.075$3.75—Not rated315——34%—1MYes
Gemini Flash LiteFree Mode

Quick lookups and summaries

0.121.24.4$0.30$0.03$2.50—22.2367——42%2091MYes
GPT-6 LunaFree Mode

Quick questions, tool calls, and images

0.030.351.3$0.10$0.01$0.50$0.12537.3162————1MYes
GPT-6 Sol

Complex sites and extensive tool use

0.607.025$2.00$0.20$10.00$2.5047.5131————1MYes
Inkling

Text tasks with longer reasoning

0.273.513$1.00$0.17$4.05—25.0114————256KNo
Inkling SmallFree Mode

Low-cost text-based tasks

0.081.14.1$0.30$0.06$1.20—27.8199————256KNo
MiMo V2.5Free Mode

Long text with repeated context

0.020.341.1$0.119$0.003$0.238—Not rated63——70%341MNo
MiMo V2.5 Pro

Long documents and deeper reasoning

0.060.862.7$0.305$0.003$0.609—26.048————1MNo
GLM 5.3

Complex text-based tasks

0.243.413$0.98$0.182$3.08—44.861————1MNo
GLM 5.3 FlashFree Mode

Code, tool calls, and screenshots

0.020.321.2$0.09$0.018$0.30—41.864————1MYes

Scores use Artificial Analysis Intelligence Index v4.3.2. AA tests at maximum reasoning effort; rtrvr uses a lower effort by default. Models without a score are marked “Not rated”.

02LIVE PERFORMANCE

How models perform on rtrvr

Speed and usage from real browser tasks in the extension, Cloud, and API. Updated hourly.

Time period
Updated hourly ·

Share of agent calls

Every hour, across all browser-agent traffic
  • GPT-6 Luna 66%
  • DeepSeek Flash 15%
  • GLM 5.3 Flash 9%
  • GPT-6 Sol 2%
  • Other 8%

Model and provider data

Table rows
Data sources

All 12 columns are shown, including task outcomes. Scroll across or jump to a group below.

Jump to columns
Production model performance over the last 24 hours
Model · providerUsageSpeed & wait timesErrors & cacheTask outcomes
Shareof callsFirst tokenmedian waitFirst token90th percentileSpeedtokens/secResponse timemedianErrorsgenerationCache hitAnsweredact runsDoneplan runsReviewed doneplan runsStuckact runs
GPT-6 Luna
66%1.3 s3.0 s1033.2 s0.2%85%97%64%44%3%
OpenAIuser key
72%1.2 s1.8 s852.4 s0.2%85%97%60%45%3%
Amazon Bedrock
21%2.7 s3.9 s1363.7 s0.0%92%————
ChatGPTuser key
7%1.6 s2.3 s985.0 s0.0%63%—74%39%—
DeepSeek Flash
15%3.6 s6.5 s1075.6 s0.2%85%92%86%41%1%
GMI Cloud
100%3.6 s6.5 s1075.6 s0.1%85%92%85%40%1%
GLM 5.3 Flash
9%3.8 s7.7 s727.1 s2.9%67%82%78%33%2%
GMI Cloud
63%5.1 s9.0 s287.5 s4.1%66%81%——2%
Azure
37%2.2 s5.1 s876.2 s0.9%68%—78%34%—
GPT-6 Soluser key
2%1.8 s5.2 s4712 s0.0%53%—84%40%—
ChatGPT
75%2.6 s5.3 s4710 s0.0%53%—83%41%—
Gemini Flashvia Google
<1%——584.7 s0.0%2%————
Gemini Flash Litevia Googleuser key
<1%——1301.4 s0.0%42%————

Speed and wait times are medians from browser tasks. Errors count failed model responses; they exclude invalid API keys and exhausted quotas. We show a row after at least five people have used it, so provider shares may total less than 100%. A dash means too little data. “Answered” and “Done” reflect what the agent reported. We do not independently verify each result.

Compare providers over time

Choose a model below to compare its providers, hour by hour.

Choose a model for these charts

Output speed

Output tokens per second, median
  • OpenAIuser key85 tok/s
  • Amazon Bedrock136 tok/s
  • ChatGPTuser key98 tok/s

First token

Time to the first output token, median
  • OpenAIuser key1.2 s
  • Amazon Bedrock2.7 s
  • ChatGPTuser key1.6 s

Round time

Median time for one model response
  • OpenAIuser key2.4 s
  • Amazon Bedrock3.7 s
  • ChatGPTuser key5.0 s

Failed responses

Calls that failed to generate
  • OpenAIuser key0.2%
  • Amazon Bedrock0.0%
  • ChatGPTuser key0.0%
03CHOOSE BY TASK

Which model should you try?

Start with one of these Free Mode models in the extension or Cloud. Free Mode includes sponsored cards and daily fair-use limits. See plans and limits

Answer questions about a page

Use GPT-6 Luna for questions and tool calls. Try Gemini Flash Lite when output speed matters most. Compare current response times in the live results above.

Collect data across many pages

Try DeepSeek Flash for collecting a dataset, filling a spreadsheet, or working through a list. It supports screenshots. Check live round times if you need a quick response at every step.

Read screenshots, PDFs, and charts

These models accept images. Choose one when the answer depends on a chart, screenshot, or scanned document.

Read long pages and documents

Both support a one-million-token context window. GLM 5.3 Flash also reads images. MiMo V2.5 offers a larger discount on cached input.

Spend fewer credits on model tokens

GLM 5.3 Flash has the lowest estimated model cost for the everyday sample workload. It also accepts images. Browser and proxy usage are billed separately on paid runs.

04DATA SOURCES

Where the numbers come from

Live results show how models run on rtrvr. Benchmarks and sample costs help you compare them.

LIVE RESULTS

Real browser tasks

We update these results hourly. Each model runs a different mix of tasks. The agent reports whether it finished; we do not independently check each result. We omit small samples.

INDEPENDENT BENCHMARK

Capability and output speed

Artificial Analysis supplies the index and first-party API speeds, recorded 23 September 2026. The index is a general benchmark, not a browser-task success rate.

ESTIMATED COST

Three sample tasks

These examples use token counts from real tasks through 31 August 2026. Estimates cover model tokens. Browser, proxy, and other usage can add to the total. Displayed rates use GMI Cloud for GLM, DeepSeek, and MiMo; Amazon Bedrock for GPT-6; Thinking Machines for Inkling; and Google for Gemini.

Quick lookup

Read a page and answer a question

Rounds
1
Input
1.5K tokens
Cached
0%
Output
300 tokens

Everyday task

Collect data or take a few actions

Rounds
1–2
Input
50K tokens
Cached
50%
Output
1.5K tokens

Long agent run

Complete a task across several pages

Rounds
5+
Input
250K tokens
Cached
70%
Output
6.5K tokens

Try a model on your task.

Choose a model in Cloud. Run the same task with another model to compare results.