← Blog
RESEARCH

Muse Spark 1.2 vs Kimi K3 vs GLM 5.2 on FlappyBench

Tested Muse Spark 1.2, Kimi K3, and GLM 5.2 with our FlappyBench. 3 models, same prompt with the /design command.

Team Command Code
1 min read
Aug 6, 2026

Tested Muse Spark 1.2, Kimi K3, and GLM 5.2 with our FlappyBench. 3 models, same prompt with the /design command. Reviewed gameplay features, UX/UI, and cost.

  • Kimi K3 → 9.5/10 · $0.0740
  • GLM 5.2 → 9/10 · $0.0480
  • Muse Spark 1.2 → 8/10 · $0.0187

Muse Spark 1.2 delivered a completely different game style despite being super cheap. Kimi K3 and GLM 5.2 outputs are comparable and close in cost.

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here: github.com/CommandCodeAI/slash-design-showcase/tree/main/flappy-bird

Try Command Code

The best coding agent for open models, in your terminal or on your desktop.

1npm i -g command-code

Desktop App · Docs · X · Discord

Share this article