Tested DeepSeek V4 Pro 0813, Kimi K3, and GLM 5.2 through FlappyBench. 3 models, same prompt with the /design command. Scored on gameplay features, UX/UI, and cost.
- DeepSeek V4 Pro 0813 → 8/10 · $0.0005
- Kimi K3 → 9.5/10 · $0.0740
- GLM 5.2 → 9/10 · $0.0480
DeepSeek is 148x cheaper than Kimi and 96x cheaper than GLM, for 84% of the quality. But $0.0005 for a playable one-shot game is insane.
Kimi K3 had the best output, with GLM 5.2 close behind.
Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird
Try Command Code
The best coding agent for open models, in your terminal or on your desktop.
1npm i -g command-codeDesktop App · Docs · X · Discord
