Qwen offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.

739 points•jmillikin•10 days ago•199 comments•

199 comments

vunderba10 days ago
So thoughts

Positives

• It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc.

• It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported.

• It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090.

Negatives

• The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again.

Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them.

https://genai-showdown.specr.net

vunderba10 days ago
Well, the results are in, at least for text-to-image (the editing bench will come later).

Qwen-Image 2.1 is definitely a pretty big leap over the last open-weight version, Qwen-Image 1.0, released back in August of last year and managed to score 7 out of 15 as opposed to its predecessor which scored 4 out of 15.

Even though it's significantly smaller, 7b vs 20b, it's multimodal (so you don't need a separate image-to-image model like you did with Qwen-Edit), more coherent, and significantly faster even when outputting at higher 2K resolutions. However, in my testing, I found that I had to play with dialing up the CFG depending on the complexity of the prompt.

I've also added a progress dropdown under Model Performance so you can see how cloud vs. local models have been trending since 2024. Spoiler: June of this year released some of the biggest bangers (Krea 2, Ideogram 4, and the kind of slept-on Boogu-Image 0.1).

Downsides:

- It was clearly trained on at least some level of synthetic training data, and it shows in some of the subpar outputs in terms of fidelity. Some of this you might be able to iron out with a refiner model downstream or a custom LoRA but time will tell.

- They've moved away from the permissive Apache license. Commercial usage is only allowed by request.

Comparisons:

https://genai-showdown.specr.net

If you just want to compare local models only:

http://genai-showdown.specr.net/?models=local

chr15m10 days ago
What's to stop somebody using the output at scale on local hardware to distill their own model and then making that available open weights?
SV_BubbleTime10 days ago
> and the kind of slept-on Boogu-Image 0.1

Not slept on at all. It was absolute trash, and I’m super curious why people pretend otherwise. There isn’t a single thing that model did better than any temporal peer.

Epitaque10 days ago
That benchmark might have some issues. You prompted the models to generate an image of striking a ring against a crucible. Then you (presumably, manually?) scored the images that depicted an anvil higher than the ones striking something resembling a crucible.
vunderba10 days ago
That’s a good catch. Yes, all scoring is done through manual review since relying on a VL model for these kinds of meta-metrics is a sort of loose equivalent of gödel's second incompleteness theorem.

I’ll have to think about this one. When I crafted the prompt, I wasn’t really thinking about the differences between a crucible and an anvil. It was more the visual of an archangel smelting halos for newly arrived heavenly beings.

zargon10 days ago
It would make more sense to compare with Qwen Image 2, since that was the last open weights Qwen model.

Edit: This is wrong.

vunderba10 days ago
Wait... is that true? I don't think the original Queen Image 2.0 was ever released beyond an API. At least, I don't remember a public weights release.
Chance-Device10 days ago
Native transparency isn’t so hard to do by the way, I made an image AE (I don’t say VAE deliberately as none of these are VAEs, I don’t know why they keep being called that since the variational part is completely absent) that supported this about two years ago as a hobby project. I haven’t really been following the space recently, I’m surprised it’s taken so long for this to come out if it’s a first.
mattnewton10 days ago
It’s not hard architecturally, but it is hard to find or create good datasets of images on the magnitude you want. I suspect the qwen team heavily used synthetic data for this.
tempesttempo8610 days ago
Could you add new OAI 2.5 image models?
vunderba10 days ago
Can do! GPT-Image-2 already scored unsurprisingly very high: 12 out of 15 on text-to-image, and 10 out of 12 on image-to-image.

The three benchmarks it failed on (D20, Flat Earth, and Banded Snake) are pretty difficult, so I'd be surprised if 2.5 manages to pass them, but I’ll add it for completeness’ sake later this week.

jfoster10 days ago
A lot of the previous Qwen models seem to have used Apache licenses, among others:

https://en.wikipedia.org/wiki/Qwen#List_of_models

Unfortunately, it looks like this model is using a much more restrictive license:

https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE

kloud10 days ago
Calling open-weights as open-source in marketing materials is the usual misrepresentation. But now with the restriction on commercial use (which is against opensource definition) it is not even open-weights, technically it would be more accurate to call it weights-available.
NewJazz10 days ago
And the only reason source available had any significance is that you could look at something and understand it. Weights are much much more opaque.

It's just freeware.

unrented797710 days ago
I'm willing to bet a nonzero amount of its training material is GPL, so I'll treat it as GPL licensed instead and use it however the fuck I want.

If AI labs get to ignore licenses, so do we.

urbnspacecowboy10 days ago
> I'm willing to bet a nonzero amount of its training material is GPL

Image data?

> GPL licensed instead and use it however the fuck I want

GPL is not a "use it however the fuck I want" license. Maybe you're thinking of the WTFPL?

soulofmischief10 days ago
You are not wrong, but will a judge and jury be competent enough to understand the difference after you've spent several hundred thousand in litigation?
user4392810 days ago
It's not going to matter unless you plan to commercially deploy the model, as far as I see.

If you were to generate outputs for commercial use, I think it would still violate this research license, but it's not like they are going to know, are they?

That said, I am disappointed that the model is not actually open-weights as I expected based on the headline.

JaggerJo10 days ago
Agreed.
gregoriol10 days ago
Was going to post about this: the last image models with Apache 2.0 license seem to be from 2025, recent Qwen models are "non-commercial use".
Luker8810 days ago
Companies can use llm to license-wash open source code regardless of license.

How difficult would it be to use this model to create a second model without licensing issues?

bloaf10 days ago
I love the non-commercial clauses because of how many people are using these for deceptive ads and “virtual staging” and fake social media accounts. Anything that makes those guys lives harder while still letting me make silly pictures for my kids and tapestries for my D&D campaign feel fine by me.
user4392810 days ago
I have seen comments on X that they will consider a revenue cap for the non-commercial restriction.

A researcher commented among the lines that they have no interest in restricting creators from using it for monetized YT content.

The way he put it, it suggested they are at this time looking to understand use cases rather than necessarily limit or charge for commercial use, and he recommended reaching out to the commercial team.

So it sounds like one could likely receive a free commercial license if needed.

plufz10 days ago
What are the top image models that still use a less restrictive license today?
vunderba10 days ago
Boogu-Image has the Apache 2.0 License [1] (good coherence, but outputs can look synthetic).

And Krea 2 has a community license [2] that is fairly permissive - I think commercial usage is allowed under $1 million.

Boogu-Image scored 6/15 and Krea 2 scored 7/15 on my GenAI Showdown benchmark [3] - only Ideogram4 eclipses them in terms of local models, but its got a far more restrictive license and the JSON structured inputs can be a pain to work with.

[1] - https://github.com/Boogu-Project/Boogu-Image

[2] - https://www.krea.ai/krea-2-licensing

[3] - https://genai-showdown.specr.net/?models=fd,hd,kd,qi,f2d,zt,...

jjcm10 days ago
I run a prompt-to-ui design site that uses image models for the design process[1]. The text rendering especially makes this model deeply interesting to me, despite the license. Here are some tests using my harness comparing the outputs of gpt-image-2 and qwen 2.1:

https://html.non.io/qwen-comparison/

The text rendering definitely is much, much better than anything else on the open weights market right now. Small text fidelity is quite good. It seems like the text encoder however gets a little bit overloaded with larger prompts - note the presence of hex codes in the design output, those were inputs from the expanded prompt.

I'll be trying a post-training run on this for web design, it has some serious potential.

[1] diffui.ai

cloudking10 days ago
Those simple prompts produce nearly the exact same layout in the 2 different models?
jjcm10 days ago
My harness expands the prompt into a json representation that specifies layout much more rigorously, which is why you see such that amount of alignment between the two.

That internal json backing helps significantly when you want to maintain consistent design system components/patterns across multiple pages. The aligned layout is it working as intended.

orbital-decay10 days ago
Totally normal for modern models due to training on the same datasets supplied by third parties, dataset contamination, and mode collapse, especially for simple prompts that don't have enough semantic capacity. -isms are often very similar even without distillation, and tend to come and go in waves along with model generations.
supermatt10 days ago
Equally confused with this. They must be using a lot more guidance than just the provided prompt.
howdareme10 days ago
Qwen is trained off of gpt’s outputs. This is both a positive and negative
BoorishBears10 days ago
Qwen's latest image models have a ton of distillation from gpt-image, same with Grok Imagine.

Even the artifacts are getting picked up.

xienze10 days ago
> The text rendering definitely is much, much better than anything else on the open weights market right now

Really? Because basically everything in those screenshots is completely garbled. I didn't follow it super closely but I thought Ideogram or whatever was really good for this particular use, with actual clear text.

vunderba10 days ago
This is my experience as well. Ideogram4 (assuming you are willing to put in the work to use the proper structured JSON input) is very accurate when it comes to text rendering in an image.
fishfasell10 days ago
The capabilities of local LLM text-to-image is honestly pretty damn impressive. IMO, I think local image generation is currently ahead of local code generation. I can get an image in seconds locally with the quality being way higher than what I'd expect from a local model. However with coding it's much slower and much less impressive. I'm sure there's a reason for this and I'm not an AI expert so I'll let the smarter folks tell me why, but that's just been my observation thus far.
mft_10 days ago
I've played with diffusion models on and off since the first release of Stable Diffusion - just for amusement, without a particular goal.

Recently, I've been helping a friend's wife with some basic vector images for her sewing hobby (she has what is essentially a CNC sewing machine) and have been super-impressed with FLUX.1-Kontext, which I've been running on my Macbook Pro with mflux. Its ability to (for example) take a photo of a human or an animal and return a line drawing which is recognisably them (rather than just a generic similarish image as I've experienced with other models) is excellent.

It's an older model now, but (AIUI) has the text-handling features baked in, and in my various testing is very reliable at giving me the outputs that I want, without the randomness I've experienced previously. It's big and relatively slow (~3 mins per 512x512 image edit on my M1 Max Mac) but excellent to work with. It's also very straightforward to set up, without the harness complexity of e.g. comfyui.

jLaForest10 days ago
is the cnc sewing machine an off the shelf model or something DIY? I'd love to hear more
rahimnathwani10 days ago
How are you converting the bitmaps into vector images?
gavmor10 days ago
Remember that quality output is a necessary but insufficient property of a generative model.

Prompt-adherence is really hit-or-miss—especially if one lacks the visual vocabulary. Likewise with coding, I find junior devs don't think to prompt re: respecting this-or-that interface, or refactoring to point-free style, etc.

So, as others have said, the artist knows better.

tarcon10 days ago
I think there was a lot more brainpower invested in the media generation side of things. The noise-based diffusion technique is further developed. It had a discovery of applying a physics-based understanding of Brownian motion to guide it. Image generation has comparatively simple training process - this is an image with dog, and without dog (contrastive learning).

Might be worth to watch the diffusion based LLMs.

victorbjorklund10 days ago
I mean I’m sure it’s the reverse for an artist. They would be less impressed with the image and more impressed with the code quality
gedy10 days ago
To generalize, LLMs are great at what you are not skilled at.
26d010 days ago
The point I think is interesting is that this is just 7B. The current SOTA 7B LLMs are barely usable for quite simple coding.
fishfasell10 days ago
That's a fair statement, I agree. I'm quite an abysmal artist so I could be a victim of my own bias here
hgufj10 days ago
I am really grateful to the Chinese Labs for open sourcing their best models. If it was left to the Americans, we would be forced to pay obscene API fees to use them.
jfoster10 days ago
Note that the license on this has this in it:

> You shall not use the Materials for any commercial purpose without obtaining a separate commercial license from us.

It probably will be much cheaper to use than other image models, but it seems that will be up to the whims of Qwen/Alibaba rather than just being the cost of putting it in a cloud provider.

https://github.com/QwenLM/Qwen-Image-2.1/blob/main/LICENSE

RIMR10 days ago
Honestly, that's fine. The commercial license isn't that bad, and cloud providers selling API access to this can afford it.

I am just happy I can run these models on my own hardware. Hopefully in 10 years, self-hosted models far exceeding what's currently available will run comfortable on commodity hardware.

tenuousemphasis10 days ago
Good luck to them enforcing that license.
Gasp0de10 days ago
Which lab is open sourcing their model?

Read the full thread on Hacker News →

Related stories