Skip to content
AI Primer
release

Alibaba opens Qwen 3.8 Max Preview testing across Cloud, Qwen Chat, Qoder and web

Alibaba opened Qwen 3.8 Max Preview on Alibaba Cloud, Qwen Chat, Qoder and the web, describing it as a 2.4T model headed for open weights. Early testers praised vision results but disputed coding claims.

7 min read
Alibaba opens Qwen 3.8 Max Preview testing across Cloud, Qwen Chat, Qoder and web
Alibaba opens Qwen 3.8 Max Preview testing across Cloud, Qwen Chat, Qoder and web

TL;DR

  • Qwen 3.8 Max Preview is live before the weights: Alibaba_Qwen's announcement says the 2.4T model debuted on Token Plan, Qoder, and QoderWork and is "going open-weight soon."
  • The web UI shows a real multimodal upgrade: the Qwen model-card screenshot lists 1,000,000 tokens of context, 65,536 tokens of max summary generation, and text, image, and video modality.
  • Alibaba's own comparison is strong but narrow: the performance chart puts Qwen3.8-Max-Preview behind Fable 5, ahead of Kimi K3, and far ahead of Qwen3.7-Max on coding and cowork pairwise tasks.
  • Early hands-on testing split cleanly: one vision test passed muffin and Ishihara checks, while synthwavedd's hands-on said real-world use did not feel as good as Kimi K3.

One access detail is unusually engineer-friendly: the official QwenCloud Token Plan says users get a base URL and API key, then configure any supported AI tool that speaks OpenAI or Anthropic protocols. The international pricing screenshot puts concurrent agents directly into the plan table; the model-card screenshot puts text, image, and video into the modality field. The artifact gap is still real: MarkTechPost's launch note says the benchmark table, license, and full public model card had not shipped with the preview.

What shipped

Alibaba opened testing across its own access surfaces first. synthwavedd's access note said Qwen3.8-Max-Preview was available for Token Plan subscribers in mainland China, while a Qwen Chat screenshot showed it selectable inside Qwen Chat.

The concrete access map:

Model card

The visible model-card fields are the useful part:

That is a product model card, not a release paper. MarkTechPost's launch note says public license terms, activation-parameter details, and the full benchmark table were still absent at preview time.

Benchmark claim

The launch hook was Alibaba's "second only to Fable 5" line. TheRundownAI's excerpt repeated that phrase, and AILeaksAndNews described the claim as placing Qwen3.8 above Kimi K3 and GPT-5.6 Sol.

Alibaba's chart breaks the claim into pairwise win/tie/loss comparisons across 400 real-world tasks:

Coding net score

  • vs Fable-5-Xhigh: -12.6%
  • vs Opus-4.8-Max: -1.8%
  • vs KIMI-K3: +6.4%
  • vs GLM-5.2: +13.2%
  • vs Qwen3.7-Max: +22.8%

Cowork net score

  • vs Fable-5-Xhigh: -7.0%
  • vs Opus-4.8-Max: +5.0%
  • vs KIMI-K3: +8.0%
  • vs GLM-5.2: +17.1%
  • vs Qwen3.7-Max: +44.4%

Those are pairwise preference-style results, not a reproducible public harness. The strongest number in the chart is not the Fable comparison, it is the jump over Qwen3.7-Max.

Hands-on split

Early testers found a better vision story than coding story.

chetaslua's vision post said Qwen3.8 passed a muffin contamination test and an Ishihara color-blindness test, then called it better than or on par with Fable 5 and Kimi K3 on vision. The same post said coding and other tasks overthought heavily and produced bad output.

Real-world use reports pushed back harder on the benchmark framing:

Day-0 caveats

The launch had enough odd behavior that some testers separated the model from the serving path.

cedric_chee said Qwen Studio thinking completed too quickly and looked like a day-0 launch bug in his report. HCSolakoglu made the same inference from fast web generation in a reply, adding that published benchmark results would have helped.

teortaxesTex gave the harsher historical read: his preview warning said Qwen has a pattern of open-weighting half-baked artifacts early in a release cycle, while his follow-up said even a model merely better than 3.7 Max would still be around GLM-5.2 or better.

Open-weight promise

Alibaba said the model is going open-weight soon, but the exact artifact remains unresolved.

The ambiguity is the word "Max." Lentils80's note said a Qwen 3.8 Max open-weight release would likely be the first open Max model, while also warning Alibaba might mean other 3.8 variants instead.

bigeagle_xd's reply said the team was working with partners to open the weights as soon as possible. The governance speculation grew fast after natolambert's post argued Qwen's biggest models had not recently been open weight, but natolambert's clarification called the Xi framing an illustration, not fact.

Scale race

The 2.4T number landed because Kimi K3 had just made 2.8T feel current.

Yuchenj_UW grouped the week as an open-weights acceleration: Qwen 3.8 at 2.4T, Kimi K3 at 2.8T, and GLM-5.2 at 753B in his comparison. teortaxesTex tied 2.4T to a Chinese hyperscaler design pattern in his parameter note, citing Ernie 5 and ByteDance experiments.

The hardware question followed immediately. petergostev's post asked how Chinese models jumped from 800-1000B to 2.4-2.8T without a hardware unlock, and TheZachMueller's reaction called the age of single-B200-node models over.

The race framing became political almost instantly. kimmonismus's thread argued China was closing the US frontier gap despite chip embargoes, and kimmonismus's follow-up endorsed that acceleration framing.

Kimi pressure

Qwen's internal chart claims an edge over Kimi K3, but Kimi set the comparison target days earlier.

Arena said Kimi K3 put China ahead of the US for the first time on Frontend Code Arena in its leaderboard post. Artificial Analysis put Kimi K3 at 57 on its Coding Agent Index in its benchmark post, matching GPT-5.6 Terra max and GPT-5.5 xhigh, ahead of Opus 4.8 max, and behind GPT-5.6 Sol max, Fable 5 max, and Grok 4.5 high.

The DeepSWE context is why coding claims are getting scrutinized. DeepSWE describes contamination-free long-horizon software tasks across 91 repositories and 5 languages, and DataCurve's Kimi post said Kimi K3 debuted at #3 as the first open-weights model with frontier-level results on that benchmark.

Kaleb trail

Before the announcement, testers were already chasing a stealth model called Kaleb.

Lentils80's LMArena report said a new model appeared under the name Kaleb, sometimes introduced itself as Claude, and looked like Qwen based on output signals and political-question behavior. AiBattle_'s Kaleb note pointed to a Qwen-like failure mode around the token "PostalCodesNL."

Code Arena had the same trail. synthwavedd's stealth-model report said Kaleb and torenia-alpha entered testing, with Kaleb likely to be Qwen's next flagship model and showing strong 3D-task performance.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 10 threads
TL;DR1 post
What shipped3 posts
Model card1 post
Benchmark claim1 post
Hands-on split2 posts
Day-0 caveats2 posts
Open-weight promise2 posts
Scale race4 posts
Kimi pressure1 post
Kaleb trail1 post
Share on X