Qwen3.8-27B
Alibaba·
Overview
Dense 27B natively multimodal model under Apache 2.0 — the companion release to the Qwen3.8 open weights and the one that actually fits a single consumer GPU. Takes text, images and video through a 27-layer vision encoder, with a native 262,144-token context extensible to ~1M via YaRN and switchable thinking. Qwen positions it for agentic work: the launch table shows the largest gains over Qwen3.6-27B in agentic coding, computer use and vision-language tasks. At NVFP4 the weights are about 15.7 GB, a Q4_K_M GGUF roughly 16.4 GB, so 24 GB of VRAM is the comfortable target.
Capabilities and innovations
Benchmarks
| Benchmark | Score | Source |
|---|---|---|
| GPQA Diamond | 89.2% | vendor GPQA Diamond 89.2 from the Qwen3.8-27B model card / launch table (up from 87.8 for Qwen3.6-27B). The model card could not be fetched directly from this environment; the value was taken from consistent secondary transcriptions of Qwen's table and not independently reproduced. |
| Humanity’s Last Exam (no tools) | 30.8% | vendor HLE 30.8 no-tools, from the same launch table, recorded for cross-model comparability. Same transcription caveat as GPQA. |
| SWE-bench Pro | 61.7% | vendor SWE-bench Pro 61.7 from the same launch table. Same transcription caveat as GPQA. |
| Terminal-Bench 2.x | 73% | vendor Terminal-Bench 2.1 73.0 (Qwen3.6-27B: 63.4) from the same launch table. No entry on the official Terminal-Bench leaderboard as of 2026-08-16. Same transcription caveat as GPQA. |
- MMLU-Pro: not reported by the vendor — Not part of Qwen's launch table, which leans on agentic and coding evaluations (DeepSWE 1.1, QwenSWEBench, LiveCodeBench v6, OSWorld, AndroidWorld).
- MMMU-Pro: independent evaluation pending — The model is natively multimodal and Qwen reports vision-language gains, but no MMMU-Pro figure was published in the launch table; no independent run available yet.
- SWE-bench Verified: not reported by the vendor — Era-1 benchmark; Qwen reports SWE-bench Pro and its own coding suite instead.
Architecture and hardware
- Parameters
- 27B total (Dense)
- Estimated VRAM at Q4
- ~16 GB, Edge / Consumer class
- Quantization formats
- NVFP4, GGUF, FP8, BF16
- Recommended runtime
- Ollama
- License
- Apache 2.0
Links
More from Alibaba
Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.