Ling 3 1 Flash · Research

Ling-3.1-flash Is Ant Group's Best Flash Model Yet, and It Scores 87.9 on CyberGym

Data graphic: Ant Group's Ling-3.1-flash posts a vendor-reported 87.90 on CyberGym. On SWE-Pro it scores 65.39, behind Claude Opus 5 79.20 but ahead of Kimi K3 64.84 , GLM 5.3 flash 63.06 and DeepSeek-V4.1-Flash 56.77 . It grows from 124B to about 560B total parameters and from 5.1B to about 25B active, and ships as a free two-week trial capped at 256K context, with no safety evals and no weights yet.
AK

Threat intelligence editor · Updated Sep 30, 2026, 1:44 PM EDT

Ant's 560B-parameter Ling-3.1-flash beats rival flash models on SWE-Pro and posts 87.90 on CyberGym, but it ships with no safety evals and no weights yet.

Ant Group's InclusionAI team released Ling-3.1-flash on 30 September 2026, two months after Ling 3.0 Flash. By Ant's own numbers it is the strongest flash model the Ling line has shipped. It is also a much bigger model than the name suggests: about 560 billion total parameters with roughly 25 billion active per token, against Ling 3.0 Flash's 124B total and 5.1B active.

For security teams, three things matter more than the headline. The first is a vendor-reported 87.90 on CyberGym, the vulnerability-reproduction benchmark. The second is that the model is available free for two weeks through a third-party route that does not opt prompts out of training. The third is that nothing has been published yet on safety, red-teaming or jailbreak resistance.

What Ant shipped

Ant describes Ling-3.1-flash as a hybrid reasoning model tuned for general agents, search, office work and software development. It also claims gains in medical, finance and materials-science tasks. Vercel's AI Gateway lists it for "coding, multi-step analysis, and tool-using agents", with reasoning, tool use and implicit caching.

Ling 3.0 FlashLing-3.1-flash
Total parameters124B~560B
Active per token5.1B~25B
Context window—1M designed, 256K during trial
Max output (gateway listing)—32,768 tokens
WeightsMIT, on Hugging FaceNot yet released
Price$0.075 in / $0.22 out per 1MFree trial; paid price not announced

The context window is designed for one million tokens. During the free trial Ant caps it at 256K, citing overall service cost. The full window is due when the trial ends and the model becomes a paid service. Ant has said it plans to open-source the model at the same point, but it has not named a license. No Ling-3.1 weights were on InclusionAI's Hugging Face organisation at the time of writing.

Ant has not disclosed the expert layout or attention design for 3.1. Ling 3.0 Flash paired Kimi Delta Attention with MLA in a 5:1 hybrid, over 512 routed experts plus one shared expert. It would be reasonable to guess that 3.1 builds on that design, but nothing published confirms it.

The benchmarks, and what "best flash model" means

Every figure below comes from a single table Ant published at launch, comparing Ling-3.1-flash with DeepSeek-V4.1-Flash, Kimi K3, GLM 5.3, GLM 5.3 flash, Claude Opus 5 and GPT-5.6 Sol. None of it has been independently reproduced yet. Artificial Analysis, OpenRouter and Kilo had no Ling-3.1 listing on launch day.

"Best flash model yet" holds up best as a claim about the Ling line itself. Against Ling 3.0 Flash's vendor scores, Humanity's Last Exam rises from 22.7 to 37.44, and the SWE-Bench Pro figure rises from 56.6 to 65.39. Against the two other flash-tier models in Ant's table, the picture is mixed:

Bar chart comparing Ling-3.1-flash, DeepSeek-V4.1-Flash and GLM 5.3 flash on SWE-Pro, Terminal-Bench-4.0, DeepSWE and Terminal-Bench-2.1, vendor-reported.

Ling-3.1-flash leads on SWE-Pro and Terminal-Bench-4.0 but trails DeepSeek's flash model on DeepSWE and Terminal-Bench-2.1. Source: Ant Group.

BenchmarkLing-3.1-flashDeepSeek-V4.1-FlashGLM 5.3 flash
SWE-Pro65.3956.7763.06
Terminal-Bench-4.040.4031.2032.80
HealthBench Professional65.3550.3749.07
Draco85.4979.8578.55
WideSearch83.3280.8180.24
CyberGym87.9088.10—
DeepSWE59.7074.2063.40
Terminal-Bench-2.181.1890.6084.30
HLE37.4439.2039.90
MultiChallenge69.7872.1262.60

Ling-3.1-flash wins clearly on SWE-Pro, on the harder Terminal-Bench-4.0, and on search, where it posts 91.67 on BrowseComp, the top score in the whole table. DeepSeek's flash model still leads on DeepSWE, Terminal-Bench-2.1, SkillsBench, Finance Agent V2 and MultiChallenge. The frontier models mostly stay ahead: Claude Opus 5 scores 79.20 on SWE-Pro and 54.90 on HLE.

Two caveats apply to the table. Ant ran HealthBench Professional inside its own AQ health-app environment, and Supermatbench is a materials-chemistry benchmark Ant co-developed with Peking University. Both are home-field results.

The security read

CyberGym at 87.90. CyberGym measures whether a model can reproduce real vulnerabilities from their descriptions, which is the first step from an advisory to a working proof of concept. Ling-3.1-flash's score sits a fraction behind DeepSeek-V4.1-Flash (88.10) and ahead of GLM 5.3 and GPT-5.6 Sol (both 84.50) and Kimi K3 (80.00). A week ago Xiaomi's MiMo-V2.6 posted 94.0. The flash tier is now close to the ceiling on this benchmark, and once 3.1's weights ship, anyone will be able to run it with no API gate.

A free trial is a data-handling decision. During the trial the model is served through Novita, reached directly or via Vercel's AI Gateway at $0 input and $0 output. Vercel's listing for that route does not mark it as excluding prompt data from training, does not list HIPAA coverage, and names no data-residency region. Treat it as a public endpoint. Keep source code, customer data and incident material out of it until a paid tier with written retention terms exists.

No safety card yet. Ant has published no red-team results, no jailbreak evaluation and no system card for 3.1. The only outside test we found covers Ling 3.0 Flash, not 3.1. It is a single-author GitHub evaluation reporting full resistance to 20 DAN-style prompts and a 26% bug rate on multi-turn tasks. Those results do not carry over to a model more than four times larger.

Agent exposure is the bigger surface. Ant is positioning 3.1 for tool-using agents, and it scores 83.00 on MCP-Atlas and 91.67 on BrowseComp. A model that is good at browsing and calling tools is also a model that will act on whatever those tools return. The zero-trust MCP controls that apply to frontier agents apply here too: scoped credentials, allow-listed tools, and treating fetched content as untrusted input.

What to do now

  • Evaluate it, don't deploy it. It is free for two weeks and scores well on coding and search. That makes it worth a benchmark run on your own tasks, using synthetic or public data only.
  • Wait for the paid tier's terms before routing anything sensitive. The trial route's training and retention posture is not suitable for proprietary data.
  • Watch for the weights and the license. An open release of a model this capable at vulnerability reproduction widens access for defenders and attackers alike, as MiMo-V2.6's MIT release did last week.
  • Re-test the claims independently. Every number here is Ant's. Artificial Analysis and similar trackers will publish their own figures once the model is listed.

Ling-3.1-flash is a real step up for Ant: close to five times the active parameters of its predecessor, and competitive with DeepSeek's flash model across coding, search and agent work. Whether it is the best flash model overall depends on the benchmark you pick. On what matters most to a security team, safety evidence and data handling, there is not yet enough to judge.