We ran a test between GPT-5.6 Sol, Kimi K3, and Gemini 3.6 Flash. One prompt. Across all three models. Using our /design command.
Results (gameplay, UX, and UI):
- Kimi K3 performed well across all three.
- GPT-5.6 Sol is good, but it gets annoying when the ball hits the wall.
- Gemini 3.6 Flash is bad (took 5–6 attempts), but it has features the others don’t, like dropping coins for power-up balls and increasing the bar length.
Ranking (DX, features, and cost):
- Kimi K3: 10/10 · $0.10
- GPT-5.6 Sol: 9/10 · $0.52
- Gemini 3.6 Flash: 6/10 · $0.45
Our engineering and design team has been testing 20+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: github.com/CommandCodeAI/slash-design-showcase/tree/main/brick-breaking
Try Command Code
The best coding agent for open models, in your terminal or on your desktop.
1npm i -g command-codeDesktop App · Docs · X · Discord
