Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

Alibaba Open-Sources Qwen3.8-27B: Frontier-Class AI on 17GB of RAM

Alibaba's Qwen3.8-27B matches frontier models on coding and agent benchmarks while running on consumer hardware at 17GB after quantization, under Apache 2.

Alibaba's Qwen team released open weights for Qwen3.8-27B on August 14, a 27-billion-parameter dense multimodal model that matches frontier-class systems on coding and agent benchmarks while running on consumer hardware after quantization2,8,11. The model passed one million downloads within two days and topped Hugging Face's global trending chart3,7. The Unsloth-quantized version has approached three million downloads.

Qwen3.8-27B ships under an Apache 2.0 license for commercial use. It is a natively multimodal model processing text, images, and video, with a 262,144-token native context window that extends toward one million tokens via the YaRN method14. Thinking mode is on by default and can be toggled per request, with reasoning depth configurable across settings including xhigh, medium, and low4.

At full 16-bit precision the model requires roughly 56GB of GPU memory; an FP8 version needs about 28GB. Four-bit quantization cuts the footprint to approximately 17GB, within reach of a 24GB GPU or a large-memory Apple Silicon device. Unsloth confirmed the model runs on 17GB RAM via its Dynamic GGUF format and that fine-tuning is supported1.

Alibaba's self-published benchmarks place Qwen3.8-27B at 61.7 on SWE-bench Pro, 90.3 on LiveCodeBench v6, 70.7 on CoWorkBench, and 84.3 on OSWorld-Verified. In Alibaba's comparison table, the model beats the listed Claude Opus 4.6 Max result on SWE-bench Pro (53.4) and CoWorkBench (68.2), and scores 11.6 points above Opus 4.6 Max on OSWorld-Verified. Benchmark firm Artificial Analysis scored it 52 on its Intelligence Index, the same interval occupied by OpenAI's GPT-5.6 Luna and DeepSeek V4 Flash6. SQ Magazine noted that Qwen's SWE-bench Pro figure is self-published and has not been replicated by outside labs.

A Reddit analysis found that Qwen3.8-27B shares exactly the same architecture as its predecessor Qwen3.6-27B, with all capability gains attributed to training improvements rather than architectural changes12. The model also includes Multi-Token Prediction; Simon Willison reported roughly 72% performance improvement on his DGX Spark after enabling MTP through llama.cpp. Willison also found the default xhigh reasoning setting generated more than 22,000 reasoning tokens — taking 21 minutes — for a single SVG generation task, and recommended starting from low or no reasoning. A separate Reddit user reported that Qwen 3.8 on the low reasoning setting loops substantially less than Qwen 3.6, and flagged the preserve_thinking parameter as important for avoiding redundant reasoning across turns.

Among developers of the coding agent Cline, Qwen3.8-27B became the most-chosen local model within four days. Tomasz Tunguz found that with reasoning enabled, Qwen edged ahead on quality in his agent stack but was roughly 30 times slower and 4.5 times more expensive in a small nine-task test against DeepSeek V4 Flash.

Alongside the 27B model, Alibaba released open weights for Qwen3.8-2.4T-A95B, a Mixture of Experts flagship with 2.4 trillion total parameters activating 95 billion per request across 512 experts5. This is the first Qwen-Max-class model to receive an open-weight release. It ships under a separate Qwen3.8-Max License rather than Apache 2.0, requiring firms with more than $50 million in annual revenue that offer commercial AI services to obtain a separate agreement. Alibaba reported scores of 67.7 on SWE-bench Pro, 93.0 on PaperBench, and 92.6 on GPQA Diamond for the 2.4T model. Paid API access is available at $2 per million input tokens and $6 per million output tokens internationally. A hosted version with one-million-token context is planned for Alibaba's Qwen Cloud.

ANALYSIS The dual release establishes a tiered open-weight strategy: permissive Apache 2.0 for the consumer-scale model, revenue-gated licensing for the datacenter flagship — a structure that parallels Moonshot AI's approach with Kimi K3.