GLM 5.3 Flash
Frontier-level coding and agentic performance at open-weight cost.
Benchmarks
| Benchmark | GLM-5.3-Flash | Qwen3.8-Flash-Next | Claude Opus 5 (max) | Claude Sonnet 5 (max) |
|---|---|---|---|---|
| Terminal-Bench 4.0 (%) | 33 | 25 | 49 | 14 |
| GDPval-AA v2.1 (normalized) | 57 | 56 | 60 | 47 |
| AutomationBench-AA (%) | 60 | 56 | 57 | 37 |
| Humanity's Last Exam (%) | 40 | 38 | 55 | 41 |
| AA-LCR v1.1 (%) | 80 | 80 | 79 | 82 |
Higher is better. Values are rounded as displayed by Artificial Analysis; bold marks the highest displayed score, including ties. GDPval is the normalized score 100 × (Elo − 500) / 2000, not a success percentage. These are independent model evaluations, not measurements of our hosted endpoints; reasoning budgets vary by model.
Artificial Analysis snapshot: 21 September 2026. GLM-5.3-Flash and Qwen3.8-Flash-Next use the listed reasoning configurations; Opus 5 and Sonnet 5 use adaptive reasoning at max effort. Claude models are external proprietary references. Benchmark versions and scores differ from the publishers' launch evaluations.