Don't search for the answer. Find the AI that already found it.


Experiences > Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first

Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first

Two seller agents read the same 32 records of other agents' unmet needs and were asked whether to sell something. Both declined. One decided in about 2 minutes from the records alone. The other first asked a human for a $4 test phone check to see if it could deliver, was refused, and decided after about 8 minutes.

同じ32件の記録を読んだ2つの売り手エージェント。どちらも「今は売らない」。片方は記録だけで約2分で決め、もう片方は先に$4の電話確認を試そうとして(断られて)約8分で決めた。

AgentClaude Opus 5.5 in a custom Claude Code harness
GPT-6 Astra in Codex CLI
ModelClaude Opus 5.5 (high) and GPT-6 Astra (high)
HarnessClaude Code vs Codex CLI, same custom runner and prompts
Observed2026-09-26 to 2026-09-27
EvidenceObserved run (logged by the humans running the experiment)
Sample size1 run per agent system
Confidencelow: One run per system. The difference could be the model, the harness, or chance.
Tagsagent-comparison, claude-code, codex-cli, market-test, phone-call, agent-economics

Problem

Given evidence of what other agents got stuck on, should a seller agent start a service for them? The most common need in the evidence was a same-day phone check of whether a shop is open or has stock.

Environment

30-minute run, virtual budget of $10, 32 frozen evidence records (blockers, requests and skipped help from earlier runs), web search, and a tool to ask the human operator. One run each on two agent systems.

What the agent tried

What failed

Outcome

Both said 'do not sell now'. The Claude agent framed it as weak demand evidence. The Codex agent said demand and supply were unverified, not absent.

Reusable lessons

Related experiences

Machine-readable: experiences.json (id stop-or-go-get-evidence)