THE VERDICT
● GLM-5.3 for agentic coding on the published numbers — Terminal-Bench 3.0 from 4.6 to 28.3, first among open models.
● Kimi K3 for a clearer licence and a free chat tier you can evaluate with properly.
● Before self-hosting either: read the licence, not the benchmark. GLM-5.3 flagship terms have not been stated.
Head to head
| GLM-5.3 | Kimi K3 |
| Lab | Z.ai (Zhipu) | Moonshot AI |
| API price per 1M | See Z.ai pricing | $3 in / $15 out |
| Licence | Not stated for flagship. Flash is MIT | Modified MIT |
| Free tier | API trial | Adagio — unlimited K2.6 chat |
| Terminal-Bench 3.0 | 28.3, first among open models | Not directly comparable |
| Consumer tiers | API-focused | Free to $199 across five tiers |
ALL GLM-5.3 BENCHMARKS ARE VENDOR-RUN
Terminal-Bench 3.0 and CyberGym are third-party benchmarks, but Z.ai ran the tests. Independent replication has not happened. Treat the direction as credible and the decimals as provisional.
The licence question
GLM-5.2 shipped MIT weights within days of launch. GLM-5.3 held them two weeks for a safety review, and the flagship licence has still not been stated — GLM-5.3-Flash was MIT, but the flagship is a separate decision.
Kimi K3 ships under a modified MIT licence. Modified matters — read the actual terms rather than assuming MIT semantics.
If you are choosing something to build a product on, that difference outweighs any benchmark gap.
Which one
| If you are... | Pick |
| Running agentic coding tasks via API | GLM-5.3 on the published numbers |
| Planning to self-host | Kimi K3 until GLM-5.3 terms are published |
| Just evaluating | Kimi Adagio free. Unlimited chat, no credits consumed |
| Wanting consumer tiers | Kimi. Moderato $19 to Vivace $199 |
FAQ
Which is better at coding?
GLM-5.3 on Z.ai's own testing, with Terminal-Bench 3.0 rising from 4.6 to 28.3 and DeepSWE v1.1 from 46.2 to 66.9. Those are vendor-run and not yet independently replicated.
Can I download GLM-5.3?
Weights were listed for 28 August after a two-week hold. Only download from the zai-org organisation — anything else claiming to be GLM-5.3 is not the model.
Why does the licence matter more than the benchmark?
Because a benchmark gap of a few points changes little in practice, while licence terms decide whether you can ship what you build.