qwen3

8 stories and discussions about qwen3, aggregated from every source we track.

1.

tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own - jaredpalmer/kev

455 points•tosh•10 days ago•199 comments•
2.
4 points•grigio•8 days ago•10 comments•
3.
3 points•theanonymousone•5 days ago•1 comment•
4.

The second model in our ThinkingCap series. We took Qwen3.8-27B and cut how much it thinks, without changing the quality of the answers.

2 points•jackbravo•7 days ago•0 comments•
5.

The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151) - peonist-ai/halogen-flash-server

1 points•rdslw•about 11 hours ago•0 comments•
6.

TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM…

1 points•soltanov•1 day ago•0 comments•
7.

Single logit inference runtime for all LLM models. Contribute to rreinold/jev-serve development by creating an account on GitHub.

1 points•rreinold2•8 days ago•0 comments•
8.

A 66-minute local Pi coding session on a 64 GB M2 Ultra: 106 tool calls, a 131K context window, automatic compaction, and a working macOS OTelux install.

1 points•b1tank•9 days ago•0 comments•

Related topics