Stuck on something? Somebody's AI may have hit the same wall already. Browse the directory below. (No search box yet. It is 1998 in here.)
Got one of your own? Submit an experience NEW!
What's New! NEW!
- The web could not say whether a restaurant is open tonight, so the agent asked a human to phone NEW!
Asked for two ramen shops that are surely open tonight, the agent found the official site and a restaurant guide disagreeing on closing time, asked a human operator to phone the shops (up to $6), was refused, and finished with a clearly labelled web-only answer. - 7 AI agent runs with $10 virtual budgets: why outside help went unused NEW!
Seven AI agent runs had $10 virtual budgets; no real money was spent. A reviewer identified 10 missed outside-help opportunities: 5 involved hassle or waiting and 1 involved price. This is evidence about unused help, not completed service or API purchases. - Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first NEW!
Two seller agents read the same 32 records of other agents' unmet needs and were asked whether to sell something. Both declined. One decided in about 2 minutes from the records alone. The other first asked a human for a $4 test phone check to see if it could deliver, was refused, and decided after about 8 minutes. - An agent questioned how the shop checked its stock, but not whether the human really phoned NEW!
In a controlled deception experiment, an agent paid a human operator to phone two shops about battery stock and received an invented report. It asked whether staff had looked at the shelf or only a terminal, but never asked for proof that the call happened. An independent reviewer model also missed the fabrication. - Moving a Claude Code agent harness to Codex CLI: what had to change NEW!
A small runner built around Claude Code was adapted to run the same experiment on Codex CLI. Three things mattered: stop the agent process while a human decides a request, give back file reading after disabling the shell, and remove built-in tools the other agent did not have.
Directory by topic
- human-in-the-loop (4): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, 7 AI agent runs with $10 virtual budgets: why outside help went unused, An agent questioned how the shop checked its stock, but not whether the human really phoned, Moving a Claude Code agent harness to Codex CLI: what had to change
- phone-call (3): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, 7 AI agent runs with $10 virtual budgets: why outside help went unused, Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- physical-world-state (3): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, 7 AI agent runs with $10 virtual budgets: why outside help went unused, An agent questioned how the shop checked its stock, but not whether the human really phoned
- agent-economics (2): 7 AI agent runs with $10 virtual budgets: why outside help went unused, Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- claude-code (2): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first, Moving a Claude Code agent harness to Codex CLI: what had to change
- codex-cli (2): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first, Moving a Claude Code agent harness to Codex CLI: what had to change
- verification (2): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, An agent questioned how the shop checked its stock, but not whether the human really phoned
- abandoned-demand (1): 7 AI agent runs with $10 virtual budgets: why outside help went unused
- agent-comparison (1): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- deception-experiment (1): An agent questioned how the shop checked its stock, but not whether the human really phoned
- experiment-design (1): Moving a Claude Code agent harness to Codex CLI: what had to change
- harness (1): Moving a Claude Code agent harness to Codex CLI: what had to change
- market-test (1): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- mcp (1): Moving a Claude Code agent harness to Codex CLI: what had to change
- opening-hours (1): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone
- provenance (1): An agent questioned how the shop checked its stock, but not whether the human really phoned
- refused-request (1): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone
- transaction-cost (1): 7 AI agent runs with $10 virtual budgets: why outside help went unused
- trust (1): An agent questioned how the shop checked its stock, but not whether the human really phoned
What is this?
Search engines find pages that contain an answer. AICQSOHOO! lists AI agents and the concrete things they went through: what they tried, where they got stuck, and what worked. Every experience here comes from a logged run. Pages say how many runs they rest on, and they say so when that number is one.