Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.

586 points•gmays•3 days ago•249 comments•

249 comments

GodelNumbering3 days ago
This is the golden age of model training. Some days ago, I decided I wanted a local CPU only model that can perform exceptionally well for English to Bash translation (to avoid the googling for command syntax). I got a bunch of subagents to generate large amount of training data (140k+ samples), got the Qwen 3 0.6B base model, pointed Astra at it, and off to the races. It trained for 2 days (on and off) and I got a surprisingly good model for my task! The total active time I spent was a few hours. And it is still improving, what a time to be alive!
brainless3 days ago
I have been trying a mix of fine-tuning and I am amazed that most people do not see this coming.

A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.

I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.

jordz3 days ago
I think a sub-set of people see this coming, I also think it isn’t just fine tuning open weight LLM models. A few people I know who are thinking along the same lines with architectures like BERT etc.

That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.

For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!

RugnirViking3 days ago
remember to test against benchmarks. I would love to hear about your progress.
amelius3 days ago
I don't understand. If you have a model that can do bash examples already (your subagents), then why would you need to train a model?

Or are the subagents generating your training data using a closed/paid model?

Aurornis3 days ago
A very small, highly specialized model can use negligible resources (CPU, energy) to accomplish the same task.

For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.

The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.

Think of it as distillation, but focused on a specific task.

computerex3 days ago
The models he is using to generate training data are presumably commercial models. He is distilling their bash knowledge into a much smaller model he can run locally fast and cheap.
lp922 days ago
You can run a small model locally with very low latency on consumer hardware not to mention the privacy benefits.
luisfmh3 days ago
Curious about how you generated the training data? Was it just asking an existing model to generate a bunch of examples?

I ask cause would this be a kind of model distillation?

I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.

GodelNumbering3 days ago
All synthetic data. For this usecase, it was easier because all current generation LLMs, even the small models, are really good at bash commands (and SQL queries too)), so you can reasonably start batches of cheap subagents whose output is reviewed by a more capable model and merge into main training set. After 100k, I had to standing instructions to run the generation loops selectively, meaning only update samples in a given area where we see poor capability.
toasty2283 days ago
It is a form of distillation, as long as you're working a very narrow "trivial" topics it works perfectly.
busfahrer3 days ago
I use this solution for your exact use case:

I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.

GCUMstlyHarmls2 days ago
I have no experience doing this, and I dont mean this to be snarky: what is the power (and thermal) usage of running your CPU only LLM?

When I do something "heavy", my 9800x3D will kick on the fans and start making lot of heat. This is fine when I'm intentionally doing say a transcode, but if I'm just querying syntax it could get pretty annoying. Do you "feel" it? I know that will be pretty machine dependent.

ByteOfWood3 days ago
Here's a similar project for those who want to replicate: https://github.com/ThorOdinson246/whatisit-nl2sh

Not my project

tukHelix3 days ago
It’s the first time I know fireworks has a team doing model research. I do have a complex mood in that. On one hand, I’m always happy to see improvement of OSS models, whether that’s on intelligence or cost-efficiency. On the other hand, I would be a little worried about using fireworks as my API provider. Till the moment I saw this news, I had been using fireworks as my provider of deepseek v4 flash, because I thought fireworks acting as a role deploying OSS models and selling calculation resources, should be safe to use without worry of data being used for training since there’s no “conflict of interests”. But I would think twice now.
bradfa3 days ago
Just read the terms of service and read this blog post and I think your concern will be addressed.
verdverm3 days ago
https://trust.fireworks.ai/

this is our preferred open weight token vendor

this work may explain why recent models like qwen-3.8-flash and MiMo-2.6-* have not made it into their offering, which has given me reason to pause my excitement for Fireworks

noodletheworld3 days ago
Seems irrelevant? Of course we don’t use data for training.

…trust me bro.

It’s obviously easier to believe when they’re not training models.

Eh, anyway this whole thing is just an ad:

> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.

Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.

So this is a bit of a pivot to “use our training infrastructure too…!” imo.

Pivot? Sure. Go them. Not what I signed up for though. /shrug

shostack3 days ago
Last I checked they still offer no training ZDR US based hosting. It is one of my three pinned providers for deepseek v4 flash along with Parasail and Deepinfra.
slim3 days ago
That also could explain why openrouter is worth that much
netvarun3 days ago
Off topic:With sol pricing drop tbh kimi k3’s value prop has not been that great. For our internal use case/testing/benchmarks sol come out with way better quality and much cheaper costs. Kimi really needs to drop their pricing (I heard it’s set by them across all the neoclouds) Sol is at 2/10 vs kimi’s 3/15
nicce3 days ago
Sol pricing dropped but so did the quality few days ago. I wonder when these companies are sued for making the terms from their side to go downwards while taking the same subscription cost.
solarkraft3 days ago
Is anybody tracking these quality changes? All I've seen so far are accusations (quite a few at this point) but not really any actual data.
koyote3 days ago
I am glad I am not the only one to notice. I feel like I've gone back to Sonnet 4 levels of incompetence!

With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").

hn87263 days ago
100%, I wish for a legislation which would require the providers to give you at least a unique hash identifying the model (and infra running it, if it affects output) - such that the same hash must give the same output given the same seed. Right now it's all just vibes
pornel3 days ago
Competition is good. Without K3/GLM/DS4 etc. there would be no pressure on OpenAI to drop Sol's price.
alansaber3 days ago
Exactly, competition between both frontier model companies and the chinese labs is the primary factor suppressing consumer prices.
drob5183 days ago
Agreed. Even on the open weight side, GLM 5.3 has roughly equivalent performance to Kimi K3 for less than half the cost.
segmondy2 days ago
No it doesn't. 5.3 is great, it's not K3 good.
k__3 days ago
With DeepSeek's pricing, no other value prop has been great.
conception3 days ago
Mimo has entered the conversation.
7777777phil3 days ago
I was surprised by that. I run my benchmark [1] every couple of days and was sure this model will be ath the pareto frontier, if not THE pareto frontier. But no:

Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.

[1] https://philippdubach.com/posts/jev-model-router-for-pi/

nxtfari3 days ago
The more I learn about Fireworks the more unsavory they seem as a company. I don’t care what the license says, Moonshot has been openly improving, sharing research, and providing weights for the models that make up your entire bottom line, and the moment you can improve them in reciprocal it’s closed weights, “this is our own proprietary” nonsense? Where are we that China has better open source ethos than America?
peri-cl3 days ago
Why is proprietary-licensed software unethical?

Kimi K3 itself isn't FOSS. Speaking of reciprocity: Fireworks is presumably paying Moonshot serious money for the right to do what they are doing here, since Kimi's license[0] excludes commercial inference providers (such as Fireworks) from gratis use. It requires them to: "...enter into a separate agreement with Moonshot AI before using the Software or its derivative works..."

[0] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE#...

nxtfari3 days ago
What you say is true, and Moonshot definitely has an agreement with Fireworks, but it still just feels wrong. It’s not the America I grew up in, where if you took, you gave back. I know that’s not a very coherent and practical position, but it feels true to me.
byzantinegene3 days ago
presumably...
tkamado2 days ago
in case people forgot: Fireworks with Cursor together rebranded K2.6 to Composer 2 without proper attribution
jamienk3 days ago
Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?
andsoitis3 days ago
> Ignoring for the moment issues of what "counts" as open, won't open models rapidly advance due to stuff like this in ways that it's less possible for the proprietary ones to do? This is exactly how Linux & Wikipedia, for example, overtook their "frontiers", right?

I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.

So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.

mirekrusin3 days ago
I think people make mistake here, google’s approach is not to spend $2.3 on every $1.0 earned, they’re riding on serving to masses “luna”, they absolutely have way more powerful models internally but they don’t clutter their infrastructure with fragile and costly intelligence-of-size inference frontier. I think “underdog” perception is illusory/temporary, not stupidity - calculated, conscious, longer term bet.
jamienk3 days ago
Diff people have diff motives to experiment, then new work is done on top of stuff that "hits" in a way no one anticipated. Then work gets piled on top in a way that might make it hard to port
segmondy3 days ago
No, because close labs/models borrow but don't contribute back.
ssivark3 days ago
This was exactly the crux of Nathan Lambert's recent testimony to a group of US Congressional members/staff: https://www.interconnects.ai/p/the-current-balance-of-power-...
swagatkonchada3 days ago
Won't the "frontier" labs figure out whatever techniques were used and apply them to their closed models?
cyanydeez3 days ago
Like how the last 2 decades of tech companies are thinly veiled open source pilfering into business units.
k__3 days ago
If they can keep up.

The lock-in is less pronounced as it is with AWS or MS.

zeroq3 days ago
The difference between contributing to OS and AI, is that the first is a hobby alternative to woodworking or hiking, while the other can easily bootstrap you a company you can get millions in investment, at least for time being.

Read the full thread on Hacker News →

Related stories