Ant Group's 560B Ling 3.1 Flash lands against GLM 5.3 Flash and Qwen 3.8 Flash Next. Specs, licences, prices, and the benchmarks Ling has not published yet.
Ant Group's InclusionAI released Ling-3.1-flash on 30 September 2026. It is the third Chinese "Flash" mixture-of-experts model in five weeks. Zhipu's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash-Next both shipped in the last week of August, and all three are pitched at the same job: running coding and tool-using agents cheaply over very long contexts.
The labels match, but the three labs made different bets. Ling 3.1 is the largest model of the three and not yet open. GLM 5.3 Flash is MIT-licensed, natively multimodal and already priced. Qwen 3.8 Flash Next activates the fewest parameters per token of the three and is a preview of Qwen4's architecture. The comparison also comes with a large gap: InclusionAI has not published a single benchmark for Ling 3.1 Flash, so any head-to-head ranking of it today is guesswork.
The spec sheet
| Ling 3.1 Flash | GLM 5.3 Flash | Qwen 3.8 Flash Next | |
|---|---|---|---|
| Lab | InclusionAI (Ant Group) | Zhipu / Z.ai | Alibaba Qwen |
| Released | 30 Sep 2026 | late Aug 2026 | 26 Aug 2026 |
| Total parameters | 560B | 320B | 125B (+51B N-gram embedding, +4B MTP) |
| Active per token | ~25B | 18B | 6B |
| Context | 1M designed; 256K during trial | 1M | 262,144 native; 1M with YaRN |
| Max output | 32,768 (Vercel, Novita listings) | 131,072 (Model Studio) | not stated on card |
| Input modalities | Text | Text, image, video, file | Multimodal |
| Weights | Not yet released | Hugging Face | Hugging Face, ModelScope |
| License | Promised open source after trial | MIT | qwen-community-1.0 |
The active-parameter column is where the most money is decided. Per-token compute scales with active parameters, not total, so Qwen 3.8 Flash Next does about a quarter of Ling 3.1's work per token. Total parameters still set the memory floor, though. Alibaba's own model card says the sparse activation "reduces compute only, not storage": the FP8 checkpoint is 172.78 GiB and BF16 is 335.28 GiB, and the minimum validated setup is two GB300s. None of the three runs on a workstation.
By simple parameter arithmetic, which is our estimate and not a vendor figure, GLM 5.3 Flash needs roughly 320 GB at FP8 and Ling 3.1 Flash about 560 GB. GLM fits on a single 8×H100 node (640 GB) with room for KV cache. Ling nearly fills that node with weights alone, before any long-context cache.
Architecture: all three bet on hybrid attention
All three models replace most of the full-attention stack with something cheaper, because full attention is what makes a 1M-token context expensive.
- Qwen 3.8 Flash Next runs 48 layers as 12 repeats of three Gated DeltaNet (linear attention) layers followed by one Qwen Sparse Attention layer. Alibaba reports up to 7.6× faster prefill and 4.9× faster decode at 1M tokens from the sparse-attention kernel. It also adds a 20-million-entry bigram and trigram lookup table that can be offloaded to host RAM with async prefetch, which is where the extra 51B parameters come from. Offload currently works on NVIDIA devices only.
- GLM 5.3 Flash combines sparse and linear attention and adds Manifold-Constrained Hyper-Connections. Z.ai says it was pretrained on a 30-trillion-token multimodal corpus, and that it beats GLM-5.2 at about a tenth of the price.
- Ling 3.1 Flash has no published architecture yet. Its predecessor, Ling-3.0-flash (124B total, 5.1B active, MIT), used a 5:1 stack of Kimi Delta Attention and gated Mamba-style linear attention across 512 routed experts. Scaling to 560B/25B suggests InclusionAI kept that family and made it much larger, but until a model card appears that is inference.
Benchmarks: only two of the three can be compared
The vendors chose different benchmark suites, and only one benchmark appears for both of the published models.
| Benchmark | Ling 3.1 Flash | GLM 5.3 Flash | Qwen 3.8 Flash Next |
|---|---|---|---|
| DeepSWE 1.1 | not published | 63.4 | 58.7 |
| Terminal-Bench 2.1 | not published | 84.3 | not reported |
| SWE-bench Pro | not published | not on card | 62.5 |
| SWE-bench Multilingual | not published | not on card | 81.0 |
| GPQA Diamond | not published | not on card | 91.7 |
| CoWorkBench | not published | not on card | 73.9 |
On DeepSWE 1.1, the one shared benchmark, GLM leads Qwen by 4.7 points. That fits the size difference: GLM activates three times as many parameters per token. Z.ai's card also marks its DeepSWE figure with an asterisk for a methodology variation, so treat that gap as indicative rather than exact.
For Ling, the only reference point is its predecessor. Ling-3.0-flash scored 56.6 on SWE-bench Pro and 72.4 on SWE-bench Multilingual with 5.1B active parameters. Qwen 3.8 Flash Next beats both scores with 6B active. If Ling 3.1's jump to 25B active does not clearly move it past Qwen and toward GLM, the larger model will be hard to justify on cost.
Price
| Input / 1M | Cached input / 1M | Output / 1M | |
|---|---|---|---|
| GLM 5.3 Flash (Z.ai) | $0.15 | $0.03 | $0.50 |
| Qwen3.8-Flash (QwenCloud) | $0.16 | — | $0.47 |
| Ling 3.1 Flash | Free for two weeks; paid rate not published |
GLM and Qwen are close to a tie on price: GLM is a cent cheaper on input and Qwen is three cents cheaper on output. QwenCloud's price is for the hosted production model, Qwen3.8-Flash, not for the open-weight Next checkpoint. Z.ai also sells a faster GLM-5.3-FlashX at $0.37 input and $1.25 output. Ling 3.1 is free on Vercel AI Gateway and Novita during the trial, capped at a 256K context.
What changes for a security team
Licence is the first filter. MIT (GLM) allows almost anything. qwen-community-1.0 is not Apache 2.0, even though some launch coverage said it was. Read it before building a commercial product on Qwen 3.8 Flash Next. Ling 3.1 has no licence yet because it has no weights yet.
Beware of impostor weights during Ling's trial window. InclusionAI has promised an open release "after the trial", with no date. Popular models with pending releases attract lookalike uploads, some with pickle-format checkpoints that can execute code on load. Pull only from the inclusionAI organisation, and prefer safetensors.
Hosted APIs mean data residency questions. Free trials are an easy way for staff to paste internal code into a third-party endpoint. For regulated data, the self-hostable models (GLM and Qwen today, Ling later) keep prompts inside your own boundary, but only if you can supply the 170–560 GB of accelerator memory they need.
The verdict, for now
- Need it in production today, with a clean licence: GLM 5.3 Flash. It is priced and documented, multimodal and MIT-licensed.
- Lowest compute per token, if you own the hardware: Qwen 3.8 Flash Next, but check the community licence first.
- Ling 3.1 Flash: try it for free while the trial runs, on non-sensitive work only. Reassess when InclusionAI publishes benchmarks, a price and the weights.
Sources
- TechNode, "Ant Group launches Ling-3.1-flash with 560 billion parameters", 30 Sep 2026
- Vercel AI Gateway and Novita model listings for Ling 3.1 Flash
- Hugging Face model cards:
inclusionAI/Ling-3.0-flash,zai-org/GLM-5.3-Flash,Qwen/Qwen3.8-Flash-Next - Z.ai pricing documentation; Alibaba Cloud Model Studio GLM-5.3-Flash page
- MarkTechPost and The Decoder coverage of Qwen3.8-Flash-Next, 26 Aug 2026