Research
Model benchmarks you can audit
Most AI model rankings are vibes with a chart. Made's benchmark program works the way a study should: every prompt set, metric, and exclusion rule below was frozen on 2026-08-13 — before any output was generated — and results only publish with two blinded raters, disclosed disagreement, and the raw run data attached.
How the program stays honest
- Preregistered. Prompt sets are frozen at version 2026-q3-v1 and published as JSON before testing starts, so prompts can't be tuned toward a preferred winner.
- Blinded. Two raters score every output with model labels and ordering randomized per prompt; both scores and their disagreement are published.
- No cherry-picking. Provider errors, refusals, and corrupt outputs count as failures — reruns beyond the preregistered count are prohibited.
- No sponsors. No provider pays for compute, reviews scores before publication, or influences the ranking. Testing runs through Made's ordinary production routing at wholesale cost.
Studies
Text-to-Video Model Comparison
PreregisteredWAN 2.2 Turbo, Veo 3.1, Kling 3 Turbo — 20 frozen prompts, 3 runs each, results publish with raw data when complete.
Image-to-Video Motion and Identity Test
PreregisteredWAN 2.2 Turbo, Kling 2.5 Turbo, Veo 3.1 — 20 frozen prompts, 3 runs each, results publish with raw data when complete.
Face and Character Consistency Test
PreregisteredNano Banana, Nano Banana 2, Nano Banana Pro — 20 frozen prompts, 3 runs each, results publish with raw data when complete.
Image-Model Prompt Adherence Test
PreregisteredFlux 1.1 Pro, Nano Banana, Nano Banana 2, Nano Banana Pro — 20 frozen prompts, 3 runs each, results publish with raw data when complete.
Cost and Time per Usable Output
PreregisteredWAN 2.2 Turbo, Veo 3.1, Flux 1.1 Pro, Nano Banana, Nano Banana 2, Nano Banana Pro — 40 frozen prompts, 3 runs each, results publish with raw data when complete.
The models under test run in Made at wholesale rates — see the price sheet for exact per-generation costs, or spec-level comparisons while the benchmark results are pending.