Qwen3.8 Max Benchmarks: Every Real Score That Exists
Spoiler: almost none do. This page tracks verified results only — vendor claims are labeled as claims, and every number carries its date and conditions.
As of 2026-07-21, there is no official benchmark table for qwen3.8-max-preview. No model card, no scores, no license. Alibaba's "second only to Fable 5" line is currently an unverified claim.
The only independent, verifiable result so far: Trilogy AI's single blind StackPerf run — Qwen3.8 Max 80 vs Kimi K3 83 (one repository task, one run: a data point, not a verdict).
The verified-results table
| Result | Source & conditions | Type | Date |
|---|---|---|---|
| StackPerf blind head-to-head: Qwen3.8-max-preview 80 vs Kimi K3 83. Kimi finished faster with fewer tokens; Qwen produced cleaner system boundaries | Trilogy AI — one matched 269-file repository task, single run | Independent | 2026-07 (verified 2026-07-21) |
| "One of the most powerful models available today… second only to Fable 5" | Alibaba announcement — no supporting table published | Vendor claim | 2026-07-19 |
| 0 of 321 tracked benchmarks have published scores for this model | BenchLM tracker (placeholder page) | Absence-of-data, itself informative | 2026-07-21 |
Why the vacuum exists
- Preview-only access. Benchmark suites need API access; Token Plan is built for interactive tools, which limits standardized runs.
- The model can drift. Preview models may be revised without notice, so today's score may not describe next week's model. Any serious result must be date-stamped.
- Unknown architecture. Total parameters (claimed 2.4T) tell you little without the unpublished active-parameter count — efficiency comparisons are premature.
Our own evaluation (in progress)
We are running a real-workload coding-agent evaluation of qwen3.8-max-preview against GPT-5.6 Sol, Claude Fable 5, and Kimi K3 — six repository-level tasks (multi-file bug fix, cross-module feature, constrained refactor, screenshot-to-code, tool-failure recovery, 1M-context codebase navigation), two independent runs each, hidden acceptance tests, blind review, and per-success cost accounting (time, tokens, credits, human interventions). It also includes a Qwen3.7-Max upgrade comparison and weekly anchor re-runs to detect preview drift.
Results publish here with full conditions and dates as tracks complete. What this page will never do: merge tracks into a single leaderboard score, or report a number without its test conditions.
How to read benchmark claims this month
- If a score has no date, discard it — preview drift makes undated numbers meaningless.
- If a "Qwen3.8 benchmark" cites MMLU-style knowledge scores, check the model ID: several circulating results actually belong to the unrelated old
Qwen3-8Bsmall model. - Single-run agent results (including Trilogy's, and early X demos) are directional at best; agent variance across two runs is large.
Last verified 2026-07-21. This page updates when official scores, the model card, or our own evaluation results land.
Sources: Trilogy AI StackPerf run · BenchLM tracker · Announcement coverage