AI Model Timeline

Tracking the accelerating release frequency of frontier AI models.

Qwen3.8-27B

Alibaba·

LLMsopen-weightEdge / Consumer

Overview

Dense 27B natively multimodal model under Apache 2.0 — the companion release to the Qwen3.8 open weights and the one that actually fits a single consumer GPU. Takes text, images and video through a 27-layer vision encoder, with a native 262,144-token context extensible to ~1M via YaRN and switchable thinking. Qwen positions it for agentic work: the launch table shows the largest gains over Qwen3.6-27B in agentic coding, computer use and vision-language tasks. At NVFP4 the weights are about 15.7 GB, a Q4_K_M GGUF roughly 16.4 GB, so 24 GB of VRAM is the comfortable target.

Capabilities and innovations

27B dense parametersNatively multimodal (text, image, video)262K native context, ~1M via YaRNSwitchable thinking modeAgentic coding and computer useApache 2.0Vision-language agent at 27B denseIntegrated 27-layer vision encoder with image and video preprocessingRuns at NVFP4 in under 16 GB

Benchmarks

BenchmarkScoreSource
GPQA Diamond89.2%vendor
GPQA Diamond 89.2 from the Qwen3.8-27B model card / launch table (up from 87.8 for Qwen3.6-27B). The model card could not be fetched directly from this environment; the value was taken from consistent secondary transcriptions of Qwen's table and not independently reproduced.
Humanity’s Last Exam (no tools)30.8%vendor
HLE 30.8 no-tools, from the same launch table, recorded for cross-model comparability. Same transcription caveat as GPQA.
SWE-bench Pro61.7%vendor
SWE-bench Pro 61.7 from the same launch table. Same transcription caveat as GPQA.
Terminal-Bench 2.x73%vendor
Terminal-Bench 2.1 73.0 (Qwen3.6-27B: 63.4) from the same launch table. No entry on the official Terminal-Bench leaderboard as of 2026-08-16. Same transcription caveat as GPQA.
  • MMLU-Pro: not reported by the vendor — Not part of Qwen's launch table, which leans on agentic and coding evaluations (DeepSWE 1.1, QwenSWEBench, LiveCodeBench v6, OSWorld, AndroidWorld).
  • MMMU-Pro: independent evaluation pending — The model is natively multimodal and Qwen reports vision-language gains, but no MMMU-Pro figure was published in the launch table; no independent run available yet.
  • SWE-bench Verified: not reported by the vendor — Era-1 benchmark; Qwen reports SWE-bench Pro and its own coding suite instead.

Architecture and hardware

Parameters
27B total (Dense)
Estimated VRAM at Q4
~16 GB, Edge / Consumer class
Quantization formats
NVFP4, GGUF, FP8, BF16
Recommended runtime
Ollama
License
Apache 2.0

Links

More from Alibaba

Data curated by AI Model Timeline. See the methodology for admission criteria, benchmark eras and source priority.