Technology & AI
Model Benchmarks
Same prompt, every model. Each chart shows how one model ranks coding, image, video, and web motion leaders. Models design their own chart. Click a chart for the full write-up; compare slots against each other and real benchmarks (SWE-bench, Terminal-Bench, video Elo).
Subpage
Grok Proof Lab
Free-reign capability demo from Cursor Grok 4.5 High Fast: live physics, generative fields, typed code that runs.
Subpage
Opus Proof Lab
Free-reign capability demo from Claude Opus 4.8: WebGL raymarching, BFS pathfinding, live systems topology, typed code that runs.
Claude Opus 4.8
Jul 10
Open Proof Lab →
GPT-5.6 Sol
Pending
Chart slot
Benchmark prompt
You are evaluating leading AI models as of July 9, 2026.
Task: Rank the strongest publicly known models across four capability domains:
1. Coding — software engineering, debugging, multi-file refactors, agentic terminal work, and shipping production code.
2. Image generation — still images, illustration, photorealism, typography in images, and controllable style.
3. Video generation — motion quality, clip length, native audio, character consistency, and production-ready pipelines.
4. Web, animation, and physics — UI/UX implementation, CSS and canvas animation, Three.js/WebGL, interactive motion design, and believable physics in browser or real-time contexts.
Requirements:
- Use publicly reported benchmarks, product docs, and real-world usage where available. Do not invent benchmark numbers.
- Rank the top 5 models per domain. Different domains may have different winners.
- Design the chart however you want. Free reign on layout, style, density, and format (SVG, HTML/CSS, canvas, tables, whatever fits). Do not copy another model's chart design. Make it yours.
- State your methodology, confidence level, and what is preview-only, deprecated, or region-locked.
- Separate "best overall coding model" from "best video model" — no single model wins everything.
- End with a short "best in class by domain" summary.
Output structure:
1. Methodology (5–8 sentences)
2. Your chart (designed by you)
3. Domain rankings with 1–2 sentence rationale per model
4. Best-in-class summary