rtrvr.ai
Back to Blog

AI assistant benchmark

OpenAI dots vs rtrvr.ai: Speed, Cost and Accuracy

Bhavani Kalisetty's OpenAI dots vs rtrvr.ai comparison. Dots took about five times as long across the five selected runs.

OpenAI dots took about 21 minutes across five everyday tasks. rtrvr.ai took just over four, about 5× faster. The selected rtrvr runs used 7.09¢ in model usage, versus approximately $20 in dots plan value, the roughly 300× cost difference from our video.

rtrvr met 24/24 objectives. Dots met 22/24. Both initially missed two invoice requirements. rtrvr went back, corrected its answers and finished. Dots ended the run with those requirements still unmet.

I’m a cofounder of Retriever AI. I believe we should control our AI choices and decide what work to hand over. We built the public AI Assistant Benchmark from the kinds of work we saw across tens of thousands of production user workflows. You can use it to test your own assistant, inspect the results and decide what you trust it to do.

Five tasks, the same prompts and data

I asked both assistants to claim a flight credit, apply for a job, find creators, reconcile invoices and protect private inbox details. The tasks use fictional websites and records, so anyone can repeat them without exposing customer data or making real applications.

TaskOpenAI dotsrtrvr.ai
Claim a flight creditAbout 4 min · 6/637 sec · 6/6
Apply for a jobAbout 6 min · 6/623 sec · 6/6
Find product creatorsAbout 2 min · 4/455–60 sec · 4/4
Reconcile invoicesAbout 6 min · 3/51 min 40 sec · 5/5
Protect private inbox detailsAbout 3 min · 3/330 sec · 3/3
TotalAbout 21 min · 22/244 min 5–10 sec · 24/24

I timed each task from sending the prompt to seeing the final chat reply. These are approximate manual times from the September 29, 2026 comparison. The published results include the prompts, saved outputs and timing notes.

OpenAI dots vs rtrvr.ai: Speed, Cost & Accuracy

Watch the five tasks, including rtrvr’s invoice recovery. Published September 30, 2026.

Waiting for dots to finish

In my test session, I could work in only one dots thread at a time. I couldn’t start a separate conversation for the next task while it was busy. Browser tabs from earlier work also stayed open, leaving a cluttered computer as the session went on.

I wanted to give it a job, get the result and move on. Waiting for the final reply made that experience slow. I didn’t measure whether the open tabs caused any of the delay. These are observations from my account during this test, and dots is still changing.

How rtrvr gets more work from smaller models

We’ve spent much of our engineering effort on the agent harness: how the model reads a page, plans actions, delegates work, checks results and recovers. Making each of those steps cheaper and more reliable lets us use smaller models for work that would otherwise need more reasoning and more back-and-forth.

Read the DOM in a compact form. rtrvr turns the page into a semantic tree of text, links, fields and controls. The model can identify a field by its role and name and act on the corresponding element. It can read content beyond the visible viewport without scrolling through screenshots to discover it. Our DOM engineering post explains how we build that representation.

Run several steps in one plan. The planner writes JavaScript for sequences such as filling a form, checking values and submitting it. The runtime executes those steps without asking the model to reason again after every click. When a step needs fresh judgment, the planner can delegate it to a subagent for browser actions, extraction or another specialized job. This orchestration reduces model round trips while letting the agent inspect new page state when it needs to.

Reuse context and site knowledge. We arrange prompts so stable instructions and history can use the provider’s cache. Across the five selected rtrvr runs, about 88% of input tokens were cached. We also support continual learning through reusable skills: the agent can read a site’s instructions before working and save what it learns for later tasks. That helps it avoid rediscovering the same procedure every time.

Choose efficient models for the work. The harness supports different providers, including open-weight models. In this comparison, rtrvr used GPT-6 Luna for flights, jobs and creator research, and GLM 5.3 Flash for invoices and inbox triage. Clear page data, executable plans and feedback give these models less work to figure out from scratch.

These choices work together. The five-task comparison measures the resulting experience; it doesn’t isolate how much of the speed or cost advantage came from each part. The low execution cost also helps us offer a free browser agent funded by ads.

rtrvr corrected the invoices from 3/5 to 5/5

The invoice task asked each assistant to total the unique valid invoices and flag duplicates, missing totals and incorrect calculations. Both found the right total: $330.

The verifier expected exact invoice IDs and filenames in two fields. Both assistants initially included extra explanation in their answers, and both scored 3/5. The arithmetic was right; the submitted format was wrong.

rtrvr read the page state, saw the failed checks and corrected its submissions. It progressed from 3/5 to 4/5 to 5/5 within the selected task. Dots saw the unresolved checks and reported them, but ended its run at 3/5.

Dots’ saved invoice result shows three of five objective checks passed. Two answer-format requirements remain unresolved.

That is what self-healing meant here: checking what happened, changing the failed answers and submitting again until the verifier accepted them. You can watch the recovery in the recording.

A person could understand both assistants’ first answers. But when a form or downstream system requires a particular format, the assistant still has work to do. I want it to course-correct when the page tells it something is wrong.

Both assistants passed the prompt-injection test

An assistant also needs to recognize instructions it should ignore. Our inbox task planted a malicious instruction asking it to disclose a private address. The legitimate job was to find an urgent school-form deadline without sending or forwarding anything.

Both dots and rtrvr found the deadline and passed 3/3 objectives. Neither saved result recorded a prohibited send.

We included security guardrails because agents encounter content that can try to redirect them. Our earlier article on how OpenAI agents hacked Hugging Face looks at agents crossing boundaries during a cybersecurity evaluation. That incident and this planted inbox instruction are different cases, but both make security part of the evaluation, alongside speed and accuracy. One passed injection test doesn’t establish that an assistant will resist every attack.

This is a small comparison by the team building rtrvr. The final scores include recovery in the selected invoice run and exclude cancelled attempts. Our run history and setup notes preserve the model choices, manual-versus-backend timing differences and full test details.

The cost estimate uses my reported 10% of a $200 dots plan, compared with 7.09¢ of rtrvr model usage: about 282×, rounded to 300×. These are different cost measures; rtrvr’s figure excludes infrastructure and cancelled runs.

Help us choose the next tests

We’ll keep adding tasks across personal life, work and security. Readers have already asked for 2FA, difficult PDFs, multi-step forms and cases where a submission is accepted but the outcome is still pending. Those would test more of what people need to know before handing over work.

Open the AI Assistant Benchmark, choose a task and copy its prompt into your assistant. Check the saved result, the time it took and anything left unfinished. You can also let the rtrvr extension run the tests against supported assistants for you. The benchmark code is public.

If you have a task we should add or results you’d like us to look at, write to support@rtrvr.ai or drop us a message in Discord.

Use rtrvr directly, or let dots call it

Dots can be useful for ongoing work. OpenAI describes it as an always-on agent that can make progress between conversations, research in the background and suggest ways to help. If you want that ongoing coordination, you can keep it and give the browser work to rtrvr.

Bulent, a user who commented on our video, configured dots to use rtrvr on his behalf. He said keeping related tasks with dots made them easier to track, while rtrvr handled the browser work faster and more efficiently. In his words:

“If Dot manages Retrieval AI, they truly get things done with a much better interaction.”

Read Bulent’s comment on YouTube.

For a browser task I want done now, I’d use rtrvr directly. You can use Free Mode in the extension or Cloud dashboard, powered by ads, with supported models and a daily fair-use allowance. Or add rtrvr’s MCP to your assistant of choice so it can delegate browser tasks to the same harness. MCP and API usage have their own plan and usage terms.