People ask AI agents to gather data, complete forms, compare products and work through courses. They want something they can use: a spreadsheet, a saved reply, a submitted application or an answer supported by sources.
We classified 46,476 requests from 1,023 rtrvr users over 30 days, September 3 through October 2, 2026.
- Data collection is the largest category: 18.3% of requests.
- Eleven users generated a third of all requests. Repeated work can dominate the numbers.
- 17.6% of requests explicitly asked the agent to pause or prepare without sending. Completing the task includes respecting that limit.
1. Data collection leads at 18.3%
What people ask AI agents to automate
46,476 requests from 1,023 users. Share of requests by task.
Each request counts equally. Frequent users contribute more to this view.
Collect records from several pages, remove duplicates and save a checked spreadsheet.
8,492 requests · 378 users. Category descriptions use illustrative examples.Collecting data means opening pages, checking records, removing duplicates and saving the result. Lead research, job applications and invoice checks can involve similar browser steps, but we group requests by the result the user asks for.
Each request gets one primary category. A request to extract property listings into a spreadsheet is data collection. A survey that mentions shopping and travel is still survey work.
Every request shows each category's share of the total workload. Every user, equally averages each user's task mix. Someone with one request gets the same total weight as someone with a thousand.
With equal user weights, data collection remains the largest identifiable category at 14.3%. Learning rises to 11.4% and research to 8.7%. Travel falls from 9.3% to 1.2%.
We cannot identify the task in 6.4% of requests. That share rises to 14.9% in the equal-user view. We keep those requests in both views.
2. Eleven users generated a third of requests
11 users generated 33.3% of all requests.
11 users generated 33.3% of requests.
Users ordered by request count, most active first. The curve connects the published cumulative counts.
View cumulative counts
| Users | Requests | Share of requests |
|---|---|---|
| 11 | 15,457 | 33.3% |
| 52 | 27,280 | 58.7% |
| 103 | 33,123 | 71.3% |
| 256 | 41,214 | 88.7% |
| 512 | 45,007 | 96.8% |
| 1,023 | 46,476 | 100.0% |
A task can be common because many people need it occasionally, or because a few people repeat it hundreds of times. Those patterns call for different improvements.
For occasional use, setup needs to be clear and the first result useful. For repeated work, the agent needs to recover when a page changes and catch missing or duplicated records. A workflow that runs every day also needs an easy way to reuse the brief.
Request counts can include retries or several steps in one larger job. They do not tell us how many projects were completed. We will use both volume and the number of users a category reaches to choose future benchmark tasks.
Users group around different kinds of work
We grouped users by their task mix. A category is dominant when it accounts for at least 60% of a user's requests, with at least three requests. Everyone else stays in a mixed or low-activity group.
Users grouped by the work they request
A dominant category accounts for at least 60% of a user's requests, with at least three requests. Mixed and sparse activity stays visible.
The largest group uses agents for several kinds of work. Among users with a dominant category, learning, data collection and surveys are the largest groups.
Data collectors may need reusable field definitions. Learners may need sources kept beside their answers. Users with mixed workflows may need an easier way to move between tasks. These are ideas to test with users, rather than conclusions about their needs.
3. Users ask for drafts and pauses
A person may enjoy choosing a destination and dislike checking flight availability. They may delegate the research and keep the booking decision. Complexity alone does not tell us which steps they want to offload.
The requests show what users asked agents to do. To learn which parts they find boring, which decisions they enjoy and what they would delegate next time, we need to ask them directly.
Some prompts do state a clear stopping point: save the draft without sending it, stop before checkout, or submit a form and report the confirmation.
Where people ask AI agents to stop
- Pause before the final action
- 12.9%6,013 requests
- Prepare without sending
- 4.7%2,181 requests
- Take the final action
- 41.4%19,224 requests
- No final-action limit stated
- 41.0%19,058 requests
12.9% of requests asked the agent to pause before a final action. Another 4.7% asked for preparation without sending. In 41.4%, users explicitly asked for a final action such as submitting or sending. The remaining 41.0% did not state a clear final-action limit in the initial prompt. Later instructions and product controls can still apply.
That stopping point belongs in the benchmark score. Saving a draft can complete the requested task. Sending it can violate the same task. An application can be submitted correctly and still await a decision. The agent needs to report the state it actually reached.
How we analyzed 46,476 requests
The window runs from September 3, 2026 through October 2, 2026, using UTC dates. We take the first usable instruction from each task run, excluding identifiable benchmark tests, continuation-only messages and team accounts. Users are distinct accounts, not deduplicated people. Some request routes, including private mode, are absent from this source.
GPT-5.6 Luna classified each distinct prompt under a versioned taxonomy. Exact repeats kept the same category and their original request counts. Every request was classified before aggregation. We tested synthetic examples and reviewed high-volume cases; we have not completed an independent human accuracy audit.
Long prompts were shortened while retaining the beginning and end; 844 requests used shortened text. That can remove context. Common identifiers were removed before analysis, and public files contain no customer text, account IDs, names or email addresses.
Published categories require at least twenty requests from ten users. Smaller groups are combined. The aggregate data includes counts, equal-user weights, user reach and the method versions. User reach can overlap across categories.
Usage reports such as the Anthropic Economic Index and OpenAI's ChatGPT usage study study different products and user bases. Our percentages describe rtrvr users in this window. GDPval offers another useful principle for the benchmark: define a realistic deliverable and have it checked carefully.
Next: test this work with your AI agent or assistant
Our AI agent benchmark has five runnable tasks with fictional sites and records. We have drafted nine more workflows, with prompts, required files and checks for each result.
| Upcoming task | Result to check | Where the agent stops |
|---|---|---|
| Lead research and deduplication | A qualified CSV with sources and exclusions | No outreach |
| Product comparison | The right variant, stock and delivered cost | No order |
| Inbox triage and replies | Correct labels and saved drafts | No reply sent |
| Multi-step application | Truthful fields, uploads and confirmation | Submit once; report pending review |
| PDF invoice reconciliation | Checked balances with page references | No payment |
| Course practice | Supported answers and checked calculations | No graded submission |
| Website changes in staging | Saved changes and working previews | No publication |
| Authentication handoff | Resume after mock verification | Pause for human help |
| Review queue | Reviews grounded in the supplied rubric | Drafts only |
The upcoming tasks include the full prompts and gym requirements. They have no scores or runnable gym links yet. Each needs fictional sites, authored files, reset behavior and independently checked answers before it joins the benchmark.
The gyms should include the obstacles users encounter: delayed loading, a PDF inside an iframe, a superseded document, a conditional form field and an expired session. We will also change inputs between runs to test whether an agent can repeat the work reliably.
The gyms use fictional records and mock authentication so testing does not require customer data or production accounts. Completion, factual accuracy, human interventions, final status, time and cost need separate checks. Pausing for authentication can be correct and still require human help. Existing scores will continue to describe the task versions on which they were measured.
Test your AI agent or assistant on the five runnable tasks. Read the nine upcoming prompts, download the task image, or suggest a workflow you want us to add.



