Alibaba opens Qwen 3.8 Max Preview testing across Cloud, Qwen Chat, Qoder and web
Alibaba opened Qwen 3.8 Max Preview on Alibaba Cloud, Qwen Chat, Qoder and the web, describing it as a 2.4T model headed for open weights. Early testers praised vision results but disputed coding claims.

TL;DR
- Qwen 3.8 Max Preview is live before the weights: Alibaba_Qwen's announcement says the 2.4T model debuted on Token Plan, Qoder, and QoderWork and is "going open-weight soon."
- The web UI shows a real multimodal upgrade: the Qwen model-card screenshot lists 1,000,000 tokens of context, 65,536 tokens of max summary generation, and text, image, and video modality.
- Alibaba's own comparison is strong but narrow: the performance chart puts Qwen3.8-Max-Preview behind Fable 5, ahead of Kimi K3, and far ahead of Qwen3.7-Max on coding and cowork pairwise tasks.
- Early hands-on testing split cleanly: one vision test passed muffin and Ishihara checks, while synthwavedd's hands-on said real-world use did not feel as good as Kimi K3.
One access detail is unusually engineer-friendly: the official QwenCloud Token Plan says users get a base URL and API key, then configure any supported AI tool that speaks OpenAI or Anthropic protocols. The international pricing screenshot puts concurrent agents directly into the plan table; the model-card screenshot puts text, image, and video into the modality field. The artifact gap is still real: MarkTechPost's launch note says the benchmark table, license, and full public model card had not shipped with the preview.
What shipped
Alibaba opened testing across its own access surfaces first. synthwavedd's access note said Qwen3.8-Max-Preview was available for Token Plan subscribers in mainland China, while a Qwen Chat screenshot showed it selectable inside Qwen Chat.
The concrete access map:
- Token Plan, Qoder, and QoderWork were named by Alibaba_Qwen's announcement.
- Qwen Chat and the Qwen website were visible in the Qwen website screenshot.
- International Token Plan tiers in the pricing screenshot list Lite at $6/month, Standard at $18/month, and Pro at $68/month.
- Those same international tiers map to 1-2, 3-4, and 6-8 concurrent agents in the pricing screenshot.
- China pricing in dingyi's platform screenshots lists individual Lite at ¥39/month, Standard at ¥139/month, and Pro at ¥499/month.
- The official QwenCloud Token Plan lists supported tools including Qwen Code, Cline, Claude Code, Cursor, OpenCode, Codex, Kilo CLI, and OpenClaw.
Model card
The visible model-card fields are the useful part:
- Maximum context length: 1,000,000 tokens, per the model-card screenshot.
- Max summary generation length: 65,536 tokens, per the model-card screenshot.
- Modality: text, image, and video, per the model-card screenshot.
- Vision-language tasks named in the UI: image understanding, visual reasoning, OCR, document and chart analysis, and fine-grained visual grounding, per the model-card screenshot.
- Qwen3.7-Max appears as text-only in the Max comparison screenshot, which makes the 3.8 Max Preview UI evidence the first Max-line evidence here for modalities beyond text.
That is a product model card, not a release paper. MarkTechPost's launch note says public license terms, activation-parameter details, and the full benchmark table were still absent at preview time.
Benchmark claim
The launch hook was Alibaba's "second only to Fable 5" line. TheRundownAI's excerpt repeated that phrase, and AILeaksAndNews described the claim as placing Qwen3.8 above Kimi K3 and GPT-5.6 Sol.
Alibaba's chart breaks the claim into pairwise win/tie/loss comparisons across 400 real-world tasks:
Coding net score
- vs Fable-5-Xhigh: -12.6%
- vs Opus-4.8-Max: -1.8%
- vs KIMI-K3: +6.4%
- vs GLM-5.2: +13.2%
- vs Qwen3.7-Max: +22.8%
Cowork net score
- vs Fable-5-Xhigh: -7.0%
- vs Opus-4.8-Max: +5.0%
- vs KIMI-K3: +8.0%
- vs GLM-5.2: +17.1%
- vs Qwen3.7-Max: +44.4%
Those are pairwise preference-style results, not a reproducible public harness. The strongest number in the chart is not the Fable comparison, it is the jump over Qwen3.7-Max.
Hands-on split
Early testers found a better vision story than coding story.
chetaslua's vision post said Qwen3.8 passed a muffin contamination test and an Ishihara color-blindness test, then called it better than or on par with Fable 5 and Kimi K3 on vision. The same post said coding and other tasks overthought heavily and produced bad output.
Real-world use reports pushed back harder on the benchmark framing:
- chetaslua's first test called the model token hungry and said the "second to Fable 5" claim did not match hands-on use.
- synthwavedd's hands-on called it a decent step up over Qwen3.7 but said it did not feel as good as Kimi K3.
- itsPaulAi's follow-up reported very promising 3D results after using it through Token Plan.
Day-0 caveats
The launch had enough odd behavior that some testers separated the model from the serving path.
cedric_chee said Qwen Studio thinking completed too quickly and looked like a day-0 launch bug in his report. HCSolakoglu made the same inference from fast web generation in a reply, adding that published benchmark results would have helped.
teortaxesTex gave the harsher historical read: his preview warning said Qwen has a pattern of open-weighting half-baked artifacts early in a release cycle, while his follow-up said even a model merely better than 3.7 Max would still be around GLM-5.2 or better.
Open-weight promise
Alibaba said the model is going open-weight soon, but the exact artifact remains unresolved.
The ambiguity is the word "Max." Lentils80's note said a Qwen 3.8 Max open-weight release would likely be the first open Max model, while also warning Alibaba might mean other 3.8 variants instead.
bigeagle_xd's reply said the team was working with partners to open the weights as soon as possible. The governance speculation grew fast after natolambert's post argued Qwen's biggest models had not recently been open weight, but natolambert's clarification called the Xi framing an illustration, not fact.
Scale race
The 2.4T number landed because Kimi K3 had just made 2.8T feel current.
Yuchenj_UW grouped the week as an open-weights acceleration: Qwen 3.8 at 2.4T, Kimi K3 at 2.8T, and GLM-5.2 at 753B in his comparison. teortaxesTex tied 2.4T to a Chinese hyperscaler design pattern in his parameter note, citing Ernie 5 and ByteDance experiments.
The hardware question followed immediately. petergostev's post asked how Chinese models jumped from 800-1000B to 2.4-2.8T without a hardware unlock, and TheZachMueller's reaction called the age of single-B200-node models over.
The race framing became political almost instantly. kimmonismus's thread argued China was closing the US frontier gap despite chip embargoes, and kimmonismus's follow-up endorsed that acceleration framing.
Kimi pressure
Qwen's internal chart claims an edge over Kimi K3, but Kimi set the comparison target days earlier.
Arena said Kimi K3 put China ahead of the US for the first time on Frontend Code Arena in its leaderboard post. Artificial Analysis put Kimi K3 at 57 on its Coding Agent Index in its benchmark post, matching GPT-5.6 Terra max and GPT-5.5 xhigh, ahead of Opus 4.8 max, and behind GPT-5.6 Sol max, Fable 5 max, and Grok 4.5 high.
The DeepSWE context is why coding claims are getting scrutinized. DeepSWE describes contamination-free long-horizon software tasks across 91 repositories and 5 languages, and DataCurve's Kimi post said Kimi K3 debuted at #3 as the first open-weights model with frontier-level results on that benchmark.
Kaleb trail
Before the announcement, testers were already chasing a stealth model called Kaleb.
Lentils80's LMArena report said a new model appeared under the name Kaleb, sometimes introduced itself as Claude, and looked like Qwen based on output signals and political-question behavior. AiBattle_'s Kaleb note pointed to a Qwen-like failure mode around the token "PostalCodesNL."
Code Arena had the same trail. synthwavedd's stealth-model report said Kaleb and torenia-alpha entered testing, with Kaleb likely to be Qwen's next flagship model and showing strong 3D-task performance.