Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.
249 comments
A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.
I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.
That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.
For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!
Or are the subagents generating your training data using a closed/paid model?
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
I ask cause would this be a kind of model distillation?
I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
When I do something "heavy", my 9800x3D will kick on the fans and start making lot of heat. This is fine when I'm intentionally doing say a transcode, but if I'm just querying syntax it could get pretty annoying. Do you "feel" it? I know that will be pretty machine dependent.
Not my project
this is our preferred open weight token vendor
this work may explain why recent models like qwen-3.8-flash and MiMo-2.6-* have not made it into their offering, which has given me reason to pause my excitement for Fireworks
…trust me bro.
It’s obviously easier to believe when they’re not training models.
Eh, anyway this whole thing is just an ad:
> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.
Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.
So this is a bit of a pivot to “use our training infrastructure too…!” imo.
Pivot? Sure. Go them. Not what I signed up for though. /shrug
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
[1] https://philippdubach.com/posts/jev-model-router-for-pi/
Kimi K3 itself isn't FOSS. Speaking of reciprocity: Fireworks is presumably paying Moonshot serious money for the right to do what they are doing here, since Kimi's license[0] excludes commercial inference providers (such as Fireworks) from gratis use. It requires them to: "...enter into a separate agreement with Moonshot AI before using the Software or its derivative works..."
[0] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE#...
I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.
So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.
The lock-in is less pronounced as it is with AWS or MS.
Read the full thread on Hacker News →
Related stories
- Hacker News · 44 points · 11 days ago
- DEV Community · 1 points · 10 days ago
- Vibes vs. Evidence: What delivers AI code review qualitydsifry.github.ioHacker News · 3 points · 9 days ago
- Lobsters · 18 points · 3 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 3 points · 11 days ago