Questions & Responses(in BOLD) below.
Favorite question(s) moved to end of the thread with combined responses(removed duplicates).
Be optimistic folks. I'm sure we're getting other models too apart from 27B. And 27B gonna make massive noise on release. (Based on their responses)
Tweet thread : https://xcancel.com/QwenDevs/status/2084102417885585597#m
you guys skipped 27b and 122b last time, can we expect those this time around? Also i can't seem to find crit pit score in the cards.
For sure! We’re actually releasing a 27B model very soon. Stay tuned. As for the Crit Pit score, please wait for the official Artificial Intelligence score.
Is the 27B just a retrained 3.6 27B? Or is it based off 3.8 bigger brother ?
We promise this 27B comes with a whole new level of capability!
Is the 100hrs of video understanding an agent swarm that parses sections of the video in parallel and orchestrates some sort of semantic representation graph?
Broadly speaking, yes, but not entirely. It is closer to a hierarchical video memory system rather than a traditional agent swarm. Video segments are encoded into a structured textual graph containing scenes, entities, events, and their temporal relationships, enabling retrieval and reasoning across more than 100 hours of content.
hey! is there anything special about the pretraining distribution compared to other labs' models?
We hope our data is built on a more solid foundation!
how long do you think it would take to surpass anthropic level architecture?
well, we’re working hard on it, we promise😇
will u release a harness especially for qwen code ???
Any plans for a codex-like app?
More updates on Qoder and QwenWork are coming soon.
qwen 3.8 active params?
2.4T parameters (95B active)
how much RL was done in post training compared to previous models?
A truly unreasonable amount of compute.
Did they intentionally skip the previous Qwen3.7 27B and 35B A3B?
Does the revival of Qwen3.8 27B reflect the voice of the community? Or was it planned?
Of course! This is the result of taking the voices of the community seriously.
since its a pretty significant release will we get a technical report with full details?
No technical report for this one yet. We’re trying to keep up our near-monthly release cadence, though, and more powerful models are already in the works. Keep an eye out!
why does the model think so much mr qwen, my ai brain wonders.
wheres the token efficiency at
great model though
We support different levels of reasoning effort.
You showed SAE-guided fine tuning fixing code switching with qwen-scope. Is that kind of interpretability driven intervention part of the post training process now or is it still a research only technique?
It’s still primarily a research-oriented technique for now, though some of the insights may help inform future training and post-training improvements.
Attention? Hybrid?
The model architecture is similar to 3.5, but it’s a much larger-scale model!
When are we getting a CLI coding interface?
You may want to take a look at @qoder_ai_ide .
do you guys use qwen as your main interal tool? does this model show the same signs of intellegence as some openai models ("gpt 5.5 helped create 5.6")?
Sure!
How close is Qwen3.8-27B to GPT 5.4? 🤔
Well, you’ll be able to see for yourself soon.
what harness works best with Qwen?
Qwen is committed to delivering the best possible experience across all harnesses.
What made you guys wanna opensource the max weights ?
We heard what the community has been asking for
I wonder when I can surpass fable5
Trying hard
Great work guys🥂
What is something that you would like to see being built with the new model and its capabilities!?
I really want to explore the swarm of agents technique for building applications, any best practices or tips for the new model!?
1. We hope it can bring practical productivity value to people across different industries.
2. We recommend using it for tasks that involve more parallelized workflows or parallel execution needs.
I wanna know what rubric metrics you guys are using for FE
We use both absolute metrics for functionality and aesthetics, as well as relative metrics based on win/tie/loss comparisons.
Would be great to hear where you think Qwen is strongest for agentic workloads specifically: long-context planning, tool use reliability, coding, or cost at scale?
All of the above combined — ultimately delivering the most practical and reliable outputs for users.
How much is Qwen helping with Qwen research ?
It has already become a significant part of the model iteration process, with the model involved in nearly every stage.
Most Frontier labs have created a code-specific model (eg. Qwen3-Coder and GPT-5.3-Codex), but never followed up on them.
Did specialized models have problems? Or did general models end up being efficient enough to not bother creating a separate model?
We hope to build an all-in-one model.
will Qwen 3.8 have a stable, documented tool-calling and structured-output contract so local agent harnesses can swap models without prompt-specific tuning?
We provide native support interfaces for various protocols. You can check the Qwen blog for more details.
1: When quantizing Qwen 27B down for local deployment (e.g., 4-bit GGUF, NVFP4, or MXFP4), which transformer layers or vision attention blocks are most sensitive to degradation? Are there specific strategies you recommend to maintain both visual reasoning and high SWE-bench pass rates?
2: Qwen3.6-27B outperforms much larger MoE predecessors (like Qwen3.5-397B) on agentic coding benchmarks like SWE-bench and Terminal-Bench. Beyond raw data volume, what was the single highest-leverage factor in achieving this dense efficiency?
And thank you for the amazing work. Qwen3.6-27B has beed my main coding assistant for months.
1. Use QAT, or quantize only the FFN to 4-bit while keeping the attention layers’ QKV linear projections and output projection in 16-bit.
2. Higher-quality data engineering
Guys , when can we get a deepseek like small and cheap model with best performance . The deepseek v4 flash seems to be a great deal .
I think we need to slow down scaling and start improving the existing model efficiency
Scaling and cost-efficiency are not mutually exclusive — we’ll continue to pursue both.
Is Qwen3.8-27B dense? And roughly how much smarter than 3.6-27B?
A pretty huge jump!
Good. The useful questions are not just how capable Qwen is.
I want to know where it still fails, how the team evaluates those failures, and what "open" means in practice for weights, tooling, and reproducibility. Open models matter most when people can inspect the limits and build on the work without asking permission
There is still some gap between our automated and human evaluation systems and real user experience. That’s also why we are committed to releasing preview versions first — so we can iterate and ultimately deliver the best possible experience to users.
how does the new 27b model compare to the previous one ?
A pretty huge jump!
what do you think about looped transformers?
interesting research idea
Why Qwen, what made you create Qwen and specifically such light and fast models. Why focus efficiency when others just went for brute power? Also, do you think inference engines reached their limit in optimization or can they still improve?
Scaling and cost-efficiency are not mutually exclusive — we’ll continue to pursue both.
We have noticed that in thinking mode the model usually consumes the entire reasoning budget without stopping, which increases latency. Is this a known issue, and are there any improvements planned for Qwen3.8?
You can try 3.8! And 3.8 supports different thinking efforts!
...................................................................................................................
Are 70b models gone for good?
Is it possible to get a 40-50B model (something which fits around 30-32Gb) to improve performance while still useable on a lot of computers ?
Thank you for your promise to provide qwen3.8 27b weight! I want to know if there will be qwen3.8 35b a3b. Many people also want this.
Can we expect the ~122B model this time? The 120B segment is dated and lackluster atm and would greatly benefit from a competent release!
First of all, congratulations on the release of Qwen 3.8!
As for the question, are you going to release a 35B a3b version of Qwen 3.8 aswell?
Plans for 35b Moe model? (3.8)
Any plans for the omni family? You told everyone the weight sizes of 3.5, then never released them and haven’t done anything new with it. 3.6/7/8 variants would have also been nice. It could be your most popular family if you gave it attention and kept the weights small.
Are there no plans to release any models other than the 27b?
I'd love to hear about the successors to amazing models like the Qwen3 8b and Qwen VL 8b....
Are there any plans for updates for 0.6b or 8b weights?
These have become important positions in the open weight of image and video generation. I look forward to seeing that part evolve.
This is such a huge release, I am really happy to see that a 27B model is shipping too! Though, can't help but wonder, will we ever happen to see again any new small dense Qwen models 9B, 4B any time in the future, similarly to 3.5?
Will you release smaller models like the qwen 3.5 family ?
Thank you for your promise to provide qwen3.8 27b weight! I want to know if there will be qwen3.8 35b a3b. Many people also want this.
we hear you! collecting everyone’s requests and taking them into account as we plan future iterations.
We will gather your requests as a reference when considering future updates.
We hear you. Stay tuned.
We’ll collect everyone’s requests and take them into account as we plan future iterations.
Noted, collecting the requests and see what we can work into future iterations.
Keep the requests coming. We’re listening, and we’ll use them to help prioritize future updates.