TESTS / REPRODUCIBLE

Public benchmark tests

The brief comes first. Readers should be able to reproduce the task, inspect the acceptance criteria, and challenge the result.

Bench 001

Build and verify one production-ready web feature inside the same repository, using the same requirements and time boundary.

01

Same brief

Each tool receives the same repository, constraints, and acceptance criteria.

02

Visible evidence

We keep prompts, artifacts, failures, costs, and the final verification output.

03

No invented winner

A recommendation remains planned until the test has actually been run and checked.

04

Freshness built in

Every claim carries a checked date because these products change quickly.

QUEUE

Tools on the first bench

  • Codex vs Claude Code
  • Cursor vs Claude Code
  • Cursor vs Windsurf
  • Gemini CLI vs Claude Code
TESTED / MINIMUM EVIDENCE

A finished-looking result is not enough

A review can use the Tested label only after its evidence package is complete and the stated acceptance checks have run.

Required evidence package
  1. Product version and checked date
  2. Locked task, inputs, constraints, and acceptance criteria
  3. Steps another builder can execute
  4. Elapsed work, attributable cost, and verification results
  5. Failed attempts and human intervention
  6. Inspectible artifacts, logs, screenshots, or source links
  7. Limitations, unknowns, and the boundary of the conclusion
ARTICLE / TEMPLATE

Reusable review structure

Every comparison, workflow, and best-of page follows the same audit trail so another builder can apply or challenge it.

01

Problem and scope

State the decision being tested, who it is for, and what the review will not claim.

02

Test environment

Record tool version, plan, repository, hardware, date, dependencies, and starting conditions.

03

Procedure

Publish the fixed brief, inputs, steps, time boundary, and acceptance checks in execution order.

04

Measured results

Separate direct measurements from observations, estimates, and unknown values.

05

Failures and intervention

Keep failed attempts, corrections, retries, and the human work required to reach the result.

06

Cost and time

Report attributable spend and elapsed work without implying precision the evidence cannot support.

07

Limitations

Name confounders, missing evidence, version drift, and scenarios the result does not cover.

08

Fit, not hype

Conclude who the tool fits, who should avoid it, and which facts support that boundary.

09

Reproduction checklist

Give readers the inputs, commands, checks, and evidence locations needed to repeat the test.