Ling 3 1 Flash · Research

Ling 3.1 Flash vs GLM 5.3 Flash vs Qwen 3.8 Flash Next: Three Bets on Cheap Agent Models

Data graphic: comparison of three Chinese Flash MoE models. Ling 3.1 Flash is the largest at 560B total and ~25B active parameters with no DeepSWE 1.1 score published; GLM 5.3 Flash 320B total, 18B active, MIT, $0.15/$0.50 per 1M tokens scores 63.4; Qwen 3.8 Flash Next 125B total, 6B active, $0.16/$0.47 scores 58.7.
AK

Threat intelligence editor · Updated Sep 30, 2026, 1:45 PM EDT

Ant Group's 560B Ling 3.1 Flash lands against GLM 5.3 Flash and Qwen 3.8 Flash Next. Specs, licences, prices, and the benchmarks Ling has not published yet.

Ant Group's InclusionAI released Ling-3.1-flash on 30 September 2026. It is the third Chinese "Flash" mixture-of-experts model in five weeks. Zhipu's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash-Next both shipped in the last week of August, and all three are pitched at the same job: running coding and tool-using agents cheaply over very long contexts.

The labels match, but the three labs made different bets. Ling 3.1 is the largest model of the three and not yet open. GLM 5.3 Flash is MIT-licensed, natively multimodal and already priced. Qwen 3.8 Flash Next activates the fewest parameters per token of the three and is a preview of Qwen4's architecture. The comparison also comes with a large gap: InclusionAI has not published a single benchmark for Ling 3.1 Flash, so any head-to-head ranking of it today is guesswork.

The spec sheet

Ling 3.1 FlashGLM 5.3 FlashQwen 3.8 Flash Next
LabInclusionAI (Ant Group)Zhipu / Z.aiAlibaba Qwen
Released30 Sep 2026late Aug 202626 Aug 2026
Total parameters560B320B125B (+51B N-gram embedding, +4B MTP)
Active per token~25B18B6B
Context1M designed; 256K during trial1M262,144 native; 1M with YaRN
Max output32,768 (Vercel, Novita listings)131,072 (Model Studio)not stated on card
Input modalitiesTextText, image, video, fileMultimodal
WeightsNot yet releasedHugging FaceHugging Face, ModelScope
LicensePromised open source after trialMITqwen-community-1.0

The active-parameter column is where the most money is decided. Per-token compute scales with active parameters, not total, so Qwen 3.8 Flash Next does about a quarter of Ling 3.1's work per token. Total parameters still set the memory floor, though. Alibaba's own model card says the sparse activation "reduces compute only, not storage": the FP8 checkpoint is 172.78 GiB and BF16 is 335.28 GiB, and the minimum validated setup is two GB300s. None of the three runs on a workstation.

By simple parameter arithmetic, which is our estimate and not a vendor figure, GLM 5.3 Flash needs roughly 320 GB at FP8 and Ling 3.1 Flash about 560 GB. GLM fits on a single 8×H100 node (640 GB) with room for KV cache. Ling nearly fills that node with weights alone, before any long-context cache.

Architecture: all three bet on hybrid attention

All three models replace most of the full-attention stack with something cheaper, because full attention is what makes a 1M-token context expensive.

  • Qwen 3.8 Flash Next runs 48 layers as 12 repeats of three Gated DeltaNet (linear attention) layers followed by one Qwen Sparse Attention layer. Alibaba reports up to 7.6× faster prefill and 4.9× faster decode at 1M tokens from the sparse-attention kernel. It also adds a 20-million-entry bigram and trigram lookup table that can be offloaded to host RAM with async prefetch, which is where the extra 51B parameters come from. Offload currently works on NVIDIA devices only.
  • GLM 5.3 Flash combines sparse and linear attention and adds Manifold-Constrained Hyper-Connections. Z.ai says it was pretrained on a 30-trillion-token multimodal corpus, and that it beats GLM-5.2 at about a tenth of the price.
  • Ling 3.1 Flash has no published architecture yet. Its predecessor, Ling-3.0-flash (124B total, 5.1B active, MIT), used a 5:1 stack of Kimi Delta Attention and gated Mamba-style linear attention across 512 routed experts. Scaling to 560B/25B suggests InclusionAI kept that family and made it much larger, but until a model card appears that is inference.

Benchmarks: only two of the three can be compared

The vendors chose different benchmark suites, and only one benchmark appears for both of the published models.

BenchmarkLing 3.1 FlashGLM 5.3 FlashQwen 3.8 Flash Next
DeepSWE 1.1not published63.458.7
Terminal-Bench 2.1not published84.3not reported
SWE-bench Pronot publishednot on card62.5
SWE-bench Multilingualnot publishednot on card81.0
GPQA Diamondnot publishednot on card91.7
CoWorkBenchnot publishednot on card73.9

On DeepSWE 1.1, the one shared benchmark, GLM leads Qwen by 4.7 points. That fits the size difference: GLM activates three times as many parameters per token. Z.ai's card also marks its DeepSWE figure with an asterisk for a methodology variation, so treat that gap as indicative rather than exact.

For Ling, the only reference point is its predecessor. Ling-3.0-flash scored 56.6 on SWE-bench Pro and 72.4 on SWE-bench Multilingual with 5.1B active parameters. Qwen 3.8 Flash Next beats both scores with 6B active. If Ling 3.1's jump to 25B active does not clearly move it past Qwen and toward GLM, the larger model will be hard to justify on cost.

Price

Input / 1MCached input / 1MOutput / 1M
GLM 5.3 Flash (Z.ai)$0.15$0.03$0.50
Qwen3.8-Flash (QwenCloud)$0.16—$0.47
Ling 3.1 FlashFree for two weeks; paid rate not published

GLM and Qwen are close to a tie on price: GLM is a cent cheaper on input and Qwen is three cents cheaper on output. QwenCloud's price is for the hosted production model, Qwen3.8-Flash, not for the open-weight Next checkpoint. Z.ai also sells a faster GLM-5.3-FlashX at $0.37 input and $1.25 output. Ling 3.1 is free on Vercel AI Gateway and Novita during the trial, capped at a 256K context.

What changes for a security team

Licence is the first filter. MIT (GLM) allows almost anything. qwen-community-1.0 is not Apache 2.0, even though some launch coverage said it was. Read it before building a commercial product on Qwen 3.8 Flash Next. Ling 3.1 has no licence yet because it has no weights yet.

Beware of impostor weights during Ling's trial window. InclusionAI has promised an open release "after the trial", with no date. Popular models with pending releases attract lookalike uploads, some with pickle-format checkpoints that can execute code on load. Pull only from the inclusionAI organisation, and prefer safetensors.

Hosted APIs mean data residency questions. Free trials are an easy way for staff to paste internal code into a third-party endpoint. For regulated data, the self-hostable models (GLM and Qwen today, Ling later) keep prompts inside your own boundary, but only if you can supply the 170–560 GB of accelerator memory they need.

The verdict, for now

  • Need it in production today, with a clean licence: GLM 5.3 Flash. It is priced and documented, multimodal and MIT-licensed.
  • Lowest compute per token, if you own the hardware: Qwen 3.8 Flash Next, but check the community licence first.
  • Ling 3.1 Flash: try it for free while the trial runs, on non-sensitive work only. Reassess when InclusionAI publishes benchmarks, a price and the weights.

Sources

  • TechNode, "Ant Group launches Ling-3.1-flash with 560 billion parameters", 30 Sep 2026
  • Vercel AI Gateway and Novita model listings for Ling 3.1 Flash
  • Hugging Face model cards: inclusionAI/Ling-3.0-flash, zai-org/GLM-5.3-Flash, Qwen/Qwen3.8-Flash-Next
  • Z.ai pricing documentation; Alibaba Cloud Model Studio GLM-5.3-Flash page
  • MarkTechPost and The Decoder coverage of Qwen3.8-Flash-Next, 26 Aug 2026