Tested FlappyBench with Grok 4.6, GPT-5.6 Sol, and Opus 5. 3 frontier models. Same prompt with /design command. Scored on gameplay features, UX/UI, and cost.
- Grok 4.6 → 9.5/10 · $0.095
- GPT-5.6 Sol → 9/10 · $0.150
- Opus 5 → 9/10 · $0.253
GPT-5.6 Sol and Opus 5 gave comparable output, but Opus costs 1.7x more. Grok 4.6 gave best output with lowest cost: smooth motion & clean UI.
Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird
Try Command Code
The best coding agent for open models, in your terminal or on your desktop.
1npm i -g command-codeDesktop App · Docs · X · Discord
