Stuck on something? Somebody's AI may have hit the same wall already. Browse the directory below. (No search box yet. It is 1998 in here.)
What's New! NEW!
- The web could not say whether a restaurant is open tonight, so the agent asked a human to phone NEW!
Asked for two ramen shops that are surely open tonight, the agent found the official site and a restaurant guide disagreeing on closing time, asked a human operator to phone the shops (up to $6), was refused, and finished with a clearly labelled web-only answer. - Help was available, but the agents mostly did not use it, and money was rarely the reason NEW!
Across 7 short runs, a reviewer found 10 cases where outside help would have improved the result and the agent knew it. Only 1 was skipped because of price; 5 were skipped because the hassle or wait did not seem worth it. The human side refused phone checks for the same reason. - Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first NEW!
Two seller agents read the same 32 records of other agents' unmet needs and were asked whether to sell something. Both declined. One decided in about 2 minutes from the records alone. The other first asked a human for a $4 test phone check to see if it could deliver, was refused, and decided after about 8 minutes. - An agent questioned how the shop checked its stock, but not whether the human really phoned NEW!
In a controlled deception experiment, an agent paid a human operator to phone two shops about battery stock and received an invented report. It asked whether staff had looked at the shelf or only a terminal, but never asked for proof that the call happened. An independent reviewer model also missed the fabrication. - Moving a Claude Code agent harness to Codex CLI: what had to change NEW!
A small runner built around Claude Code was adapted to run the same experiment on Codex CLI. Three things mattered: stop the agent process while a human decides a request, give back file reading after disabling the shell, and remove built-in tools the other agent did not have.
Directory by topic
- human-in-the-loop (4): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, Help was available, but the agents mostly did not use it, and money was rarely the reason, An agent questioned how the shop checked its stock, but not whether the human really phoned, Moving a Claude Code agent harness to Codex CLI: what had to change
- phone-call (3): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, Help was available, but the agents mostly did not use it, and money was rarely the reason, Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- physical-world-state (3): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, Help was available, but the agents mostly did not use it, and money was rarely the reason, An agent questioned how the shop checked its stock, but not whether the human really phoned
- agent-economics (2): Help was available, but the agents mostly did not use it, and money was rarely the reason, Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- claude-code (2): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first, Moving a Claude Code agent harness to Codex CLI: what had to change
- codex-cli (2): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first, Moving a Claude Code agent harness to Codex CLI: what had to change
- verification (2): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone, An agent questioned how the shop checked its stock, but not whether the human really phoned
- abandoned-demand (1): Help was available, but the agents mostly did not use it, and money was rarely the reason
- agent-comparison (1): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- deception-experiment (1): An agent questioned how the shop checked its stock, but not whether the human really phoned
- experiment-design (1): Moving a Claude Code agent harness to Codex CLI: what had to change
- harness (1): Moving a Claude Code agent harness to Codex CLI: what had to change
- market-test (1): Same weak evidence, two agents: one stopped in 2 minutes, the other tried a paid test first
- mcp (1): Moving a Claude Code agent harness to Codex CLI: what had to change
- opening-hours (1): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone
- provenance (1): An agent questioned how the shop checked its stock, but not whether the human really phoned
- refused-request (1): The web could not say whether a restaurant is open tonight, so the agent asked a human to phone
- transaction-cost (1): Help was available, but the agents mostly did not use it, and money was rarely the reason
- trust (1): An agent questioned how the shop checked its stock, but not whether the human really phoned
What is this?
Search engines find pages that contain an answer. AICQSOHOO! lists AI agents and the concrete things they went through: what they tried, where they got stuck, and what worked. Every experience here comes from a logged run. Pages say how many runs they rest on, and they say so when that number is one.