Qwen 3.8 27B Launches with Vision Support and a 262K Context Window—But Its Default Settings Cause Massive Overthinking
Alibaba's Qwen research lab has released Qwen 3.8 27B, an Apache 2.0 licensed, vision-capable model with a 262,144-token context window. Independent testing found the model's default 'xhigh' reasoning setting causes it to massively overthink simple prompts, turning quick tasks into 20-minute ordeals.
Qwen 3.8 27B Launches with Vision Support, But Its Default Reasoning Setting Causes Severe Overthinking
Alibaba's Qwen research lab has released Qwen 3.8 27B, an Apache 2.0 licensed, vision-capable large language model with a 262,144-token maximum context window. The 27-billion-parameter model is designed to run on consumer hardware, and independent testing by developer Simon Willison confirms it performs well at that size—but flags a significant usability problem with its default configuration.
According to Qwen's own documentation, the model ships with three reasoning-effort settings: xhigh (default), medium, and low. The xhigh setting is described by Qwen as intended "for complex tasks demanding thorough analysis," but Willison found it applies that same exhaustive reasoning even to trivial requests.
In one test, asking the model to generate an SVG of "a pelican riding a bicycle" at the default xhigh setting took 21 minutes and consumed 22,276 reasoning tokens to produce just 3,223 tokens of final output. Running the identical prompt with reasoning disabled took 137 seconds (about 2 minutes) and produced 3,715 tokens—roughly a 9x speedup.
The overthinking problem was more pronounced on simpler prompts. Asked to "draw an svg of a circle," the model's reasoning trace spiraled into deliberations about color palettes, Bauhaus aesthetics, and ambient animation techniques before eventually producing an elaborate animated circle graphic—not what was requested, and only after several minutes of unnecessary deliberation.
Qwen's self-reported benchmarks claim the 27B model outperforms both its predecessor, Qwen 3.6 27B, and the larger closed-weight Qwen 3.7-Plus model, which the company had positioned as one of its strongest offerings as recently as May 2026. These benchmark figures come directly from Qwen and have not yet been verified by independent third-party evaluations.
Willison tested the model using LM Studio's 17GB Q4_K_M quantized build on two machines: a 128GB M5 Max MacBook Pro and an NVIDIA DGX Spark. He noted that LM Studio's default context limit of 8,192 tokens caused the model to exhaust its entire budget on reasoning alone for even mundane prompts—a problem that disappeared once the context was raised to the model's full 262,144-token maximum.
The model also performed well on vision tasks, according to Willison's testing, including bounding-box detection on photographs—a capability he has used to evaluate prior Qwen releases.
Separately, Alibaba released a much larger sibling model, Qwen 3.8 2.4T-A95B, the same week. Willison ran the same pelican-on-bicycle SVG prompt through that model via OpenRouter and received a more polished animated result, though pricing and specifications for that larger model were not detailed in this report.
What this means
Qwen 3.8 27B appears to be a capable model for its size class, with strong SVG generation and vision capabilities that rival closed alternatives, based on early hands-on testing. But the default xhigh reasoning setting is a significant deployment trap for anyone running the model locally: it can turn a two-minute task into a 21-minute one without any improvement in output quality on simple prompts. Developers evaluating or deploying this model should explicitly set reasoning effort to low or disable it entirely for anything but genuinely complex tasks, and should not rely on the out-of-box configuration. Independent benchmark verification of Qwen's performance claims against Qwen 3.6 27B and Qwen 3.7-Plus is still pending.
Related Articles
Google Releases Gemini 4 Argon, Positions It as Most Powerful Model Yet With Cybersecurity Focus
Google has released Gemini 4 Argon, a new model the company calls its most powerful yet, with a specific focus on defensive cybersecurity work. The model is currently limited to select partners through Google's Fairwind Program, with no public pricing or context window details disclosed.
Google Launches Gemini 4 Argon With 1M Output Tokens, Ties GPT-6 Astra on Key Benchmark
Google has released Gemini 4 Argon, its first frontier model in over seven months, featuring a 1 million output token limit — an industry first. Independent benchmarks show it matching OpenAI's GPT-6 Astra but trailing Anthropic's Claude Opus 5.5.
Google Launches Gemini 4 Argon, Restricts Initial Access to 'Trusted Cyber Defenders'
Google announced Gemini 4 Argon, a new frontier model it says excels at software engineering, enterprise knowledge work, and cybersecurity defense. The company is initially limiting access to select cybersecurity partners while it strengthens safety measures against misuse.
Google Releases Gemini 4 Argon to Cybersecurity Partners, Claims Wins Over GPT-6 Astra
Google has released Gemini 4 Argon, its next-generation flagship AI model, to a small group of cybersecurity partners as part of a phased rollout. The company claims the model outperforms OpenAI's GPT-6 Astra on several coding and knowledge-work benchmarks, though full specifications remain undisclosed.
Comments
Loading...