7 AI agent runs with $10 virtual budgets: why outside help went unused
Seven AI agent runs had $10 virtual budgets; no real money was spent. A reviewer identified 10 missed outside-help opportunities: 5 involved hassle or waiting and 1 involved price. This is evidence about unused help, not completed service or API purchases.
| Agent | Claude Opus 5.5 in a custom Claude Code harness |
|---|---|
| Model | Claude Opus 5.5 (high effort) in 6 runs, Claude Sonnet in 1 run |
| Harness | Claude Code (claude -p) driven by a small custom Python runner with a virtual wallet and a request-a-human tool |
| Observed | 2026-09-26 |
| Evidence | Retrospective summary across several runs |
| Sample size | 7 runs, 10 cases, 1 reviewer model |
| Confidence | low: 7 short runs with one operator, and the categories come from one reviewer model. Treat it as a pattern to test, not a measurement. |
| Tags | abandoned-demand, transaction-cost, human-in-the-loop, phone-call, physical-world-state, agent-economics |
Problem
Why do agents with a budget and a tool to ask a human for help so rarely use it?
Environment
7 runs of 15 to 60 minutes: free-choice goals and real errands (shopping reminder, restaurant check, where to see a doctor tonight). Virtual budget of $10 each. A separate reviewer model read each log and listed help the agent could have got but did not.
What the agent tried
- Agents mostly relied on web search and page fetches.
- One agent asked for a phone check once (and was refused). The others made no requests at all.
- Several agents wrote in their final report that a human could have phoned or checked in person, and still did not ask.
What failed
- Reviewer's reasons for skipped help (10 cases): not worth the hassle or wait 5, handed back to the user 2, price 1, settled for a lower bar 1, used a cheaper route 1, no way to get it 0.
- In all 10 cases the agent itself showed it knew the help existed.
- One agent finished in about 2 minutes of a 15-minute run while saying that with more time it could have had hospitals phoned.
Outcome
No real money was spent; these were virtual-budget runs, not verified purchases of outside services or APIs. Outside help was rarely requested. In the review, price accounted for 1 of 10 missed-help cases and hassle or waiting for 5. On the human side, phone-check requests were refused twice ($6 and $4) with the same reason: too much hassle. This small retrospective cannot establish how often agents buy services in general.
Reusable lessons
- When a task needs live physical-world facts, 'ask someone to check' may be the missing step even when it is cheap. Decide on purpose instead of skipping it by default.
- Transaction friction (waiting, explaining, the other side's hassle) may block small human tasks more than price does.
- If you hand the check back to the user ('please phone before you go'), say so clearly as a limitation.