Qwen offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.

346 points•jjcm•13 days ago•138 comments•

138 comments

mavamaarten13 days ago
I'm wondering, is there a tool or something out there that helps me pick a model, in the vast sea of models out there these days? Every time I need a model for something I see the list on openrouter and I'm completely overwhelmed.

I'd love to be able to explain my use case, my cost preferences and have a tool select a few good models to try.

E.g. I wrote a tool that cleans out my email spam box. It classifies emails that are already flagged as spam, and if it's very obviously spam it removes it permanently (keeps a copy on disk though). And after x emails, it goes through the list of deleted spam mails and suggests email rules. What model would be best suited? I'd love to be able to explain this use case and get this info served to me. The list of models and the information about what they're good at is just too splintered and spread out. I landed on google/gemma-4-31b for now, because it's cheap and good enough and also supports Dutch and French a bit. But I can't realistically try them all.

Systemerror7A6913 days ago
My approach to this problem is to just...not try them all.

As long as the model you're using solves the problems you have to your satisfaction, there is no need to try any other models, except for financial reasons maybe.

So I start with a relatively cheap model (GLM 5.3 flash for me) and as long as it accomplishes the task (it did so far) I don't have to change. And even if it can't do something, the first thing I change is see if I can give it more tools or better context (useful even if I switch models later) or trying a different approach to the problem.

If google/gemma-4-31b works, you don't need to overthink it.

epolanski13 days ago
> If google/gemma-4-31b works, you don't need to overthink it.

Up until recently I had a gemini flash 2.0 api deployed that did summarization and translation of news articles/corporate statements fast and cheap and had no reason to update it.

If it works fine, this chase of the latest LLM is bit pointless.

insightfulornot13 days ago
Starting with GLM-5.3 Flash was a pretty decent first try! I started with other models, and ended up settling on this exact one because all the others were either too slow or unreliable for my tasks. Qwen 3.8 didn't do it for me, whatever tweaks I added to my harness. Where I'm getting at is you did start with an incredible model in the first place, which greatly helps sticking to it.
barrenko13 days ago
I use Gemini(s) because I can send pdfs as files to their API and not worry too much. I've started to diverge and consacrate a part of my pipeline to sending image based pdf pages to glm flash 5.3, not sure how to address / test it properly.

Long term I have fears I can't depend of the Google's AI api.

bmordue13 days ago
satisficing instead of optimising
testycool12 days ago
I use models.dev's CLI tool, which I think gets data from OpenRouter, and ArtificialAnalysis so your coding agent can help you narrow it down.

<sidenote>

Similarly, HuggingFace has a CLI + a few skills, and they are very useful.

I had a production image processing using Gemini 2.5 Flash Lite (which is getting discontinued in October), and in 20 minutes Claude Code + HF Cli recommended the best replacement small model (Qwen VL 3B something) and proceeded to fine tune it on my datataset. All this while I was in a rush to get dressed and go to the store.

It cost ~$3 I think, and results were excellent. Not perfect, but not far from perfect either.

We didn't replace Gemini in prod at the time, because we didn't have time to do all the math on how to end up with a smaller bill/mo.

</sidenote>

sisve13 days ago
Have you tried openrouters auto model? Tries to give you the best model based on prompt and price

https://openrouter.ai/docs/cookbook/coding-agents/openclaw-i...

ComputerGuru12 days ago
That’s if you have disparate prompts and don’t want to actively select a model. If you’re developing a pipeline, it’s a terrible idea. You want to choose a model, validate it, then stick to it.
mavamaarten12 days ago
I have, but honestly that was exactly what I am not looking for. Sometimes it picked a model for a Dutch email that totally does not support Dutch. Other times it would work fine. It's just a layer of indeterminism I wasn't looking for.
rudicjd274712 days ago
If you use the Chinese ones at least the energy comes from solar - aside from that, pick one and see if it solves your problems
paimapi12 days ago
I doubted this but it does look like there's an actual government initiative that mandates 80% clean energy for all new data-center builds: https://www.fastcompany.com/91578780/how-china-is-powering-n...

that said, China is also rapidly scaling up coal-fired plants: https://apnews.com/article/china-coal-power-plant-carbon-cli...

those presumably support all of the surrounding infrastructure + people + manufacturing so it's not as if it's truly solar-powered. but it's still handily better than the state-by-state abandonment of clean energy goals here in the US - I lay this out a bit here: https://news.ycombinator.com/item?id=49700743

yorwba12 days ago
If you use a model hosted in China while it's night there, the energy obviously doesn't come from solar, as there's not enough storage capacity. Even if you use it during daytime, most of the energy still won't come from solar, because the ideal solar power locations in the sparsely-populated west are far from the ideal data center locations in the densely-populated east and there's not enough transmission capacity between them.

Additionally, using a Chinese provider doesn't mean the model will be hosted in China, e.g. for Qwen Omni here, the supported regions are: China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt), and US (Virginia). https://www.alibabacloud.com/help/en/model-studio/qwen-omni#...

idiotsecant12 days ago
if you mine 4 tons of coal, anywhere on earth, 1 ton of it will be going to china to be burned to make electricity. China's energy production is extraordinarily dirty, the cleanest part of it is the PR. More than half of the energy they produce is from burning coal. They certainly want to integrate more renewables but they have the same problems with that everyone else does - storage and transmission are expensive and essential for a renewable heavy grid.
ezst12 days ago
> at least the energy comes from solar

If by that you misspelled coal, sure

https://ourworldindata.org/grapher/share-elec-by-source?coun...

bonoboTP13 days ago
It's not like new releases come with fully mapped out capability scores for exactly the aspects that you're interested in. There are benchmarks, but reality is often different. It's simply unknown to humanity how well each model will perform in your own bespoke context unless you just try them. You can read experiences and vibes by others but often they will use them in different ways or have different preferences etc.

They generally all try to make them good at everything, it's not like they'd declare "this model is not made for task X".

_ache_13 days ago
If the performances are comparable, and there is no evidence it's not.

in/out ($) Gemini : 1.5 / 9.0 | Qwen 3.8: 0.15 / 0.47

That is a massive cost reduction.

Refs: https://www.alibabacloud.com/help/en/model-studio/model-pric... https://runware.ai/gemini-omni

killingtime7413 days ago
You can't just look at the per token cost, but how many tokens it takes on average to do a task. The difference can be massive.
vntok13 days ago
True, but it would have to be more than massive (order(s) of magnitude) to offset that gap.
tidbeck13 days ago
Also cache write/read cost + cache efficiency.
_ache_12 days ago
I know, but it's a good enough proxy.
syntaxing13 days ago
> audio-visual performance close to Gemini 3.8 Flash and overall audio performance that exceeds Gemini 3.8 Flash

Wow crazy if true. I think Gemini's audio capability and multi language was the "selling point" for a lot of people. Other capability also matches or exceeds 3.8 Flash.

They also made a new harness but github link seems to 404.

testaburger13 days ago
probably their distill target
conception13 days ago
3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
spijdar13 days ago
They seem to be doing something different with the "Qwen4" architecture as demoed in Flash-Next. I've noticed the reasoning behaves ... weirdly. Like, really weirdly compared to any model I've ever seen before.

I've noticed between tool calls, it'll sometimes say things like:

  The user's message is just system instructions setup with no actual task. There's no question to answer yet. I should acknowledge briefly and wait for the actual request.

  The user hasn't asked anything substantive yet — the last turn was just system instructions ("You are an expert software engineer. Helps user to solve problems."). My previous response was a brief acknowledgment. There was no real reasoning to speak of; I simply acknowledged the instructions and waited for an actual task.

  【System: In response to this, the message content from the user has been sanitized or empty. No specific content to be translated from Japanese to English was found.】
These don't clearly reflect ... anything, and it keeps performing tool calls correctly anyway. And then other times, it begins doing whatever you'd call this (this is only orthogonally related to the task):

  A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.
anon37383913 days ago
This is a serving bug or quantization issue. I had all kinds of issues that were like this on DGX Spark until I found a single-GB10 vLLM recipe [1] that uses Nvidia's NVFP4 quant. The community quants did not work well.

Another failure mode you may see is inordinately long CoT. Properly served, the model is good at calibrating its CoT length to the difficulty of the immediate task.

[1] https://github.com/blazux/qwen3.8-Flash-DGX

kouteiheika13 days ago
> I've noticed the reasoning behaves... weirdly

Is this with the full unquantized weights? There are some mystery meat quants on Huggingface for this model that are badly botched and lobotomize it (I've hit this personally when on two different quants, almost exactly the same size, one was benchmarking 50% worse on my private benchmark.).

saghm13 days ago
Earlier today I was playing around with the "Union Alpha" stealth model (which I guess exited stealth later in the evening), and I noticed it had a habit of trying to respond to the subagents it spawned while giving me an answer. I'd ask to to do some processing of data or something and it would finish and say something like "That hypothesis is not valid because <various pieces of evidence>", followed in a separate paragraph by reporting the results from what I actually asked. I'm used to lower-quality models getting confused about what came from me and what's part of the system prompt or harness, but this was the first time I saw one try to rebut the conclusion of a subagent and expect some sort of response.
denom13 days ago
> ... a person who cannot grow, and never will , vs. a person who changes completely every seven years

Wow, that is unexpected. But honest?

Morizero13 days ago
I saw some corrupting when using https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-S... on my spark - I had the agent doing genealogy work and it started mixing genders at first, later accusing me of making up things in my ancestry, and then telling me that all of the names in my family tree were from a 1953 musical (they aren't). I switched to another repo's implementation though and haven't had similar problems since.
rubslopes13 days ago
> RL it to oblivion.

What would that mean in this context?

pennomi13 days ago
Tuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness.

I swear I spend more time telling Claude not to do things than telling it what to do.

khafra13 days ago
Others have given examples, but here's the theory: https://www.lesswrong.com/posts/fuSaKr6t6Zuh6GKaQ/when-is-go...

Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric.

But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.

cleaning13 days ago
See 5.6, Astra, and Opus 4.8 for examples
conception13 days ago
In this context, benchmaxing, if you will, so hard towards agentic coding benchmarks that everything else suffers.
podocarp13 days ago
Please what is flash pro ultra and all these, can they just use semver or something
xutopia12 days ago
The different names have different meanings and help make decisions on which to use.

Omni means you can use multiple types of input and have multiple types of outputs like audio, video, images and text. Flash means that it is built for speed and smaller than the more complete ones.

spacebanana713 days ago
I believe in general flash models prioritise speed, ultra/pro/omni do more slow reasoning to the effect of sometimes better intelligence, and lite models prioritise cost.
drbscl13 days ago
Omni usually means multimodality (in terms of input and/or output type, text, images, audio, etc)

Read the full thread on Hacker News →

Related stories