AI

Alibaba’s Qwen 3.8 Max costs $2 per million tokens — open weights arrive Aug 10

Adrian Kessler

A 2.4-trillion-parameter AI model just appeared on the market at $2 per million tokens for input and $6 for output. It runs on a sparse mixture-of-experts architecture that keeps per-token costs down by activating only 95 billion parameters at a time, with the rest sitting idle. The model is Qwen 3.8 Max, and it comes from Alibaba.

The pricing gap with top US models is stark. Qwen 3.8 Max posts competitive results on several standard benchmarks: 93.0 on PaperBench and 86.1 on OSWorld-Verified, placing it above Fable 5 and GPT-5.6 Sol on those specific measures. Its context window extends to one million tokens — roughly one million words of input in a single prompt — which covers long-document analysis, large codebases, and extended reasoning chains. The model is also multimodal, accepting images and video alongside text, at a price tier that was proprietary-only a year ago.

The more consequential development may be what Alibaba has promised for next week. The company says it will release the model weights publicly, meaning any team with sufficient computing infrastructure will be able to run Qwen 3.8 Max locally without API fees, without sending data to an Alibaba server, and without a usage agreement. That commitment, if delivered on schedule, would put a 2.4-trillion-parameter model in the hands of self-hosters for the first time. Alibaba is also releasing a smaller 27-billion-parameter checkpoint for organizations with more limited compute.

The caveats are real. On the text-focused Chatbot Arena leaderboard, Qwen 3.8 Max ranks fifth — behind four models in Anthropic‘s current lineup. On SWE-bench Pro, the coding-agent evaluation that matters most to enterprise AI teams, it scores 67.7 against Fable 5’s 80.0. The open-weights release is promised for the week of August 10, but as of publication no weights have appeared on Hugging Face or ModelScope. For organizations dealing with EU data-residency requirements or US export-control rules, Qwen 3.8 Max carries the same jurisdiction questions that have attached to other Chinese lab models.

The August 10 window for the public weights release is the first hard checkpoint. If Alibaba delivers, the cost structure of frontier AI inference drops substantially for any team willing to run their own infrastructure. If the release slips or arrives with restrictions, the $2 API price still makes it one of the most aggressively priced frontier models currently available.

Tags: , , , , ,

Discussion

There are 0 comments.