made

Research

Model benchmarks you can audit

Most AI model rankings are vibes with a chart. Made's benchmark program works the way a study should: every prompt set, metric, and exclusion rule below was frozen on 2026-08-13 — before any output was generated — and results only publish with two blinded raters, disclosed disagreement, and the raw run data attached.

How the program stays honest

Studies

The models under test run in Made at wholesale rates — see the price sheet for exact per-generation costs, or spec-level comparisons while the benchmark results are pending.