Run Q4 small language models entirely in the browser with WebGPU. Chat, benchmark, and compare PetitGPT, SmolLM2, and friends on-device.

280 points•logicallee•2 days ago•113 comments•

113 comments

demibabs2 days ago
Cool project, but I'd really suggest looking at the UI.

The text is too small and it's way too dense with information in general. Considering how simple this product is to use, it's kinda crazy that I have to scroll through over a page length of (mostly useless, AI-generated) information before getting to the actual interface.

Also what is going on with the footer (why does it link back to the site itself, why is it telling me to "serve over HTTP").

post-it2 days ago
> The text is too small and it's way too dense with information in general.

The Claude special.

fl0id2 days ago
For real. all these dashboards look quite similar
amelius2 days ago
Also, they should give some examples of what you can use these Small Language Models for, and what their limits are.
phist_mcgee2 days ago
This is the future of software, sloppy ui.
azan_2 days ago
Yes, before LLMs we never had sloppy UIs.
logicallee1 day ago
I've thought some more about your feedback, since other replies gave the same feedback. At the same time, I didn't want to remove the "mostly useless" information since clearly it was useful for a lot of people! The submission was on the front page of HN (often in the #3 or #4 spot) for 12 hours and last I checked generated more than 300,000 requests from 36,000 unique IP's (on an ordinary 2-day period there are around 1,600 unique IP's) and people downloaded 600 GB of models, including 8.5k downloads initiated for the first model (3,000 of which were completed and 5.5k of which did not wait for it to complete, since download was a little slow due to the high number of concurrent users.)

So, clearly, the presentation format resonated with a lot of people. Therefore, I kept the presentation but put a tl;dr in enormous can't-miss-it shimmering font at the top. I marked all the "dense information" you mentioned as "Optional reading" (it's a lab after all, there should be a reading for it) and increased the font size.

I removed the parts of the footer that you said didn't make sense. The page isn't on the front page anymore so I don't know if people would like the changes or not, but I've changed the page layout in response to your feedback.

janalsncm1 day ago
I should preface this by saying it is a cool demo. I love SLMs and I think they will become even more popular in the future as capability per byte improves and hardware improves to support my bytes.

So I think this post performed well in spite of the annoying parts, not because of it.

> I didn't want to remove the "mostly useless" information since clearly it was useful for a lot of people

The information was not “clearly” useful. It had obviously incorrect information that no one even noticed. In my opinion that is strong evidence for the opposite conclusion, that people ignored it because it was noise.

Why would people ignore this information? Demos are a show, don’t tell thing. For example, you don’t have to tell people that the latency is low. They should be able to see it from the demo.

langurmonkey2 days ago
It's not loading for me on Firefox (v156.0.1, Arch Linux), works fine on Chromium.

Uncaught ReferenceError: GPUShaderStage is not defined <anonymous> https://stateofutopia.com/experiments/microllmlab/engine/web... webgpu-metal.js:24:17 <anonymous> https://stateofutopia.com/experiments/microllmlab/engine/web... [MicroLLM lab] App loader failsafe triggered after 6s microllmlab:55:21

Doohickey-d2 days ago
Firefox doesn't support WebGPU on Linux, that's probably why.
butz1 day ago
There is "Prefer WebGPU" toggle (on by default), but it does nothing. How hard would it be to add simple support test to check if browser has all required capabilities to run your application? I think at least allowing to browse models does not require WebGPU, right?
logicallee1 day ago
I've tried to fix the Firefox on Linux issue, can you check again? Firefox doesn't have WebGPU enabled by default on Linux so you will have to use the wasm fallback which is slower.
tolugenius2 days ago
I did the default arithmetic with PetitGPT research-v1

>What is 2+2?

Answer

> To find 2 + 2, we need to add 2 to both sides of the equation.

> 2 + 2 = 4

> So, 2 + 2 = 4 + 2.

Brilliant

tecleandor2 days ago
GPT-2 124M is terrible...

  > what is 2+2?
Answer:

  > 3+3? 4+4? 5+6? 7+8?
  > Reply ~18000 0 ~10 min 2 By : 1-1: I'm a beginner. 3x2 is my best option, but if you're not sure about the other options then just go for it and try again
krackers2 days ago
It's not instruct tuned looks like? It's closer to a base model rather than a chatbot.
logicallee2 days ago
GPT-2 is an interesting one because it is a February 2019 model. (You can see some information about it below the card if you click on the card.)

That was 2-3 years before the big "ChatGPT moment" (the highly coherent ChatGPT research preview was released in November 2022, I think it was ChatGPT 3.5). Back in 2019 the models really were not producing very coherent output. Now you can see it for yourself right in your browser :) Everything has come a really long way since then!

kasumispencer22 days ago
Sounds about right about something that's 124M. I trained one myself a few weeks ago and it's about the same level of being terrible.
logicallee2 days ago
I got the correct output for PetitGPT research-v1: https://ibb.co/0pP9DS2T
dotancohen2 days ago
That's not incorrect.

LLMs produce semantically correct sentences, not factually correct statements. Have we forgotten this so soon?

anyfoo2 days ago
It is, in every sense, incorrect. Which statement in this short snippet is "semantically correct"? (Better LLMs get this right, of course.)
NicuCalcea2 days ago
It's not very good semantically either.

> Give me a recipe for soup.

> Here is a recipe for soup:

> Saffa-Cake-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-

mgaunard2 days ago
> how would you compare your capabilities to that of claude fable 5.1 by anthropic

> Comparing your capabilities to that of claude fable 5.1 would be very similar. Both are stories about a clown who is a cunning and manipulative animal. They both have problems and challenges, and they face challenges to overcome them. However, they have different goals and approaches. The clown

I think that tells you all you need to know.

Reviving15142 days ago
That's hilarious. At least the smaller model is more self aware!
atobe2 days ago
Similar, simpler: https://tiny.tobelabs.com/

Small LM trained on Tiny Stories. My first goal was to get people to see tokens from the model ASAP, the second was to demo Tiny Stories for those who may be unfamiliar. Oh and play with ternary models. https://arxiv.org/abs/2305.07759

There's larger QA model on the same page too.

logicallee2 days ago
Interesting project, thanks for sharing.

Read the full thread on Hacker News →

Related stories