Data extraction & scraping
Pull structured records from listings, dashboards, and directories, usually into a sheet.
Browser agent data and evals
rtrvr has served 35,000+ users across more than 7 million live-web runs. We turn those runs into training data, evals, and browser environments for agents that need to act, recover, and verify.
Plain-text version for agents and crawlers: /data.md“apply to this job”
id=61id=66id=75id=97id=112upload_file · element 75fresh page stateReal lines from the job-application sample below.
The production signal
It records what the agent saw, what it tried, what failed, how it recovered, and whether the work was actually complete.
Why this data matters
Real runs show whether an agent can find the right control, carry state, recover, and verify the final result.
Pages can contain thousands of interactive elements, including controls outside the viewport.
Long tasks cross tabs, redirects, logins, apps, and human approval turns.
A stale element or failed tool call should repair the plan, not end the run.
The agent must check the result instead of treating its own last action as proof.
Three production samples
Choose a task, select a tool call, and load the exact page state the agent used. The viewer reads the same JSON that ships in each bundle.
Instruction “apply to this job”
Production task mix
A recent one-week window, deduplicated to distinct tasks and grouped by function. Use the full mix for a generalist agent, or start with the slice your model needs.
Pull structured records from listings, dashboards, and directories, usually into a sheet.
Applications through Ashby, Greenhouse, Lever, and Workday, including files, fields, and essay answers.
Creating and editing docs and spreadsheets, and moving results between the browser and Google Workspace.
Finding people, verifying identity, sending connection requests and messages, and confirming delivery.
Course modules, quizzes, and certification flows on SkillsBuild, SCORM players, and publisher courseware.
Government portals, invoicing systems, CRMs, and internal tools, including checkout flows that pause for approval.
Drafting and posting inside the target app's own editor, including CMS bodies and rich-text fields.
Posting, replying, reacting, and scheduling across social platforms, usually from a list of targets.
Work with rtrvr
We work with teams building browser agents, language action models (LAMs), and world models on offline policy data, online browser environments, evals, and verifiers.
Offline policy
Use a focused cut of the corpus for post-training, imitation, or offline policy learning.
Online policy
Co-design browser environments for models that learn by acting, recovering, and checking results.
Evals
Build evals around the tasks your model is close to solving but still fails in production.
For coding-agent teams
Our planner writes code against the page, executes it, receives typed errors, repairs the program, and stops only after a final verification. The same loop matters for coding agents: long-horizon plans, tool execution, recovery, and checked completion.
Research
Our public work on page understanding, agent harnesses, reliability, and cost. The commercial data and evals are scoped with each partner.
Sample data
Each ZIP includes the trajectory, every available page state, an index, summary counts, and the schema README.
Bundle contents
workflow.json full trajectory trees/<id>.json captured page states trees/INDEX.json IDs, URLs, and counts SUMMARY.json coverage and totals
@misc{rtrvr2026trajectories,
title = {Web Agent Trajectories from Production Traffic},
author = {{rtrvr.ai}},
year = {2026},
howpublished = {\url{https://rtrvr.ai/data}},
note = {Production browser-agent runs captured as enriched accessibility-tree
observations with typed actions and verified outcomes}
}Work with rtrvr
Bring us one task your model misses. We will build the data, browser environment, eval, or verifier needed to improve it.
Talk to the rtrvr team30 minutes · bring one task or failure case