qwen3
8 stories and discussions about qwen3, aggregated from every source we track.
tiny Jev-like family of decision models built on top of Qwen3.5 you can train and run on your own - jaredpalmer/kev
The second model in our ThinkingCap series. We took Qwen3.8-27B and cut how much it thinks, without changing the quality of the answers.
The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151) - peonist-ai/halogen-flash-server
TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM…
Single logit inference runtime for all LLM models. Contribute to rreinold/jev-serve development by creating an account on GitHub.
A 66-minute local Pi coding session on a 64 GB M2 Ultra: 106 tool calls, a 131K context window, automatic compaction, and a working macOS OTelux install.