COMPARE / SAME TASK

Compare AI builders

Same task, same repository, same constraints. The result is judged by what ships — including the human work required to get it there.

Current state

The briefs are being prepared. No winner is named until prompts, artifacts, failures, costs, and verification output are recorded.

Coding agentsResearching

Codex vs Claude Code

One web feature, one repository, one acceptance test — measured for completion quality, intervention cost, and token spend.

10K–100K primary query band · YoY +900%
Editor vs terminalResearching

Cursor vs Claude Code

Editor-native flow versus terminal agent: compare codebase orientation, multi-file edits, review burden, and recovery from mistakes.

1K–10K query band · strong recent Trends signal
AI-native IDEsPlanned

Cursor vs Windsurf

A controlled IDE comparison for context retrieval, multi-file changes, terminal work, and the cost of correcting an agent.

1K–10K band in both query directions
Terminal agentsPlanned

Gemini CLI vs Claude Code

A narrower terminal-agent test focused on repository context, tool permission flow, failure recovery, and repeatability.

1K–10K band · low ad competition