Skip to content
AI Primer

Ling-2.6-flash

Faster Responses, Stronger Execution, Higher Token Efficiency

Ling-2.6-flash is an open-source instruct large language model release from inclusionAI/Ant Group, using a hybrid-linear sparse MoE architecture with 104B total parameters and 7.4B active parameters, optimized for fast, token-efficient agent workflows, coding, document processing, tool use, and long-context language tasks.

Pricing

Model profile · Current snapshot
Input / 1M
$0.10
Output / 1M
$0.30
Blended / 1M
$0.15
Output TPS
84.14
TTFT (s)
0.88

Model Intelligence

Context window
262,144 tokens
Arena ranking
14
Benchmarkable
Yes
Model level
release
Intelligence Index
14.2
Coding Index
25.3
GPQA
0.59
HLE
0.06
SciCode
0.27
IFBench
0.57
LCR
0.28
TerminalBench Hard
0.21
TAU2
0.86

Recent stories

1 linked story
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.