← Blog
RESEARCH

DeepSeek V4 Pro 0813 vs Kimi K3 vs GLM 5.2 on FlappyBench

Tested DeepSeek V4 Pro 0813, Kimi K3, and GLM 5.2 through FlappyBench. 3 models, same prompt with the /design command.

Team Command Code
1 min read
Aug 12, 2026

Tested DeepSeek V4 Pro 0813, Kimi K3, and GLM 5.2 through FlappyBench. 3 models, same prompt with the /design command. Scored on gameplay features, UX/UI, and cost.

  • DeepSeek V4 Pro 0813 → 8/10 · $0.0005
  • Kimi K3 → 9.5/10 · $0.0740
  • GLM 5.2 → 9/10 · $0.0480

DeepSeek is 148x cheaper than Kimi and 96x cheaper than GLM, for 84% of the quality. But $0.0005 for a playable one-shot game is insane.

Kimi K3 had the best output, with GLM 5.2 close behind.

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird

Try Command Code

The best coding agent for open models, in your terminal or on your desktop.

1npm i -g command-code

Desktop App · Docs · X · Discord

Share this article