MicroLLM Lab – Try 7 tiny LLM's in the browser(stateofutopia.com) |
MicroLLM Lab – Try 7 tiny LLM's in the browser(stateofutopia.com) |
> "2+2 is 2."
Otherwise, a very neat demo. As others have said, the UI is VERY confusing, way too much stuff going on.
Other questions and answers were mixed:
How many cards in a deck? > A deck of cards contains 52 cards.
How many cards in a deck if I remove all Queens from the deck? > If you remove all Queens from the deck, there are still 52 cards in the deck.
(It's a little bit slow at the moment - you have to wait a few seconds for the models to load - as it's currently on the HN front page. The server is on a 1 gigabit unmetered network connection so it can serve all the weights - around 600 megabytes - to one person every few seconds, there are several concurrent users now.)
Adding a warning about it could help.
What version of Firefox are you using and what is your operating system and graphics card, please? Can you also try it without WebGPU? (Reload the page and uncheck "Prefer WebGPU" and try a prompt.)
Unticking the WebGPU checkbox and reloading didn't do anything.
1. Add 1 tablespoon of water to the water bath. 2. Place the tub into the water bath and let it sit for about 5 minutes. 3. After 5 minutes, remove the tub and let it cool down. 4. Now, fill the tub with water and let it sit for about
uhmm completely unusable ?
Useful applications include sentiment analysis, text classification, entity extraction, etc.
They certainly can be useful, but you shouldn't compare them with LLMs such as Opus or Fable.
The text is too small and it's way too dense with information in general. Considering how simple this product is to use, it's kinda crazy that I have to scroll through over a page length of (mostly useless, AI-generated) information before getting to the actual interface.
Also what is going on with the footer (why does it link back to the site itself, why is it telling me to "serve over HTTP").
The Claude special.
So, clearly, the presentation format resonated with a lot of people. Therefore, I kept the presentation but put a tl;dr in enormous can't-miss-it shimmering font at the top. I marked all the "dense information" you mentioned as "Optional reading" (it's a lab after all, there should be a reading for it) and increased the font size.
I removed the parts of the footer that you said didn't make sense. The page isn't on the front page anymore so I don't know if people would like the changes or not, but I've changed the page layout in response to your feedback.
So I think this post performed well in spite of the annoying parts, not because of it.
> I didn't want to remove the "mostly useless" information since clearly it was useful for a lot of people
The information was not “clearly” useful. It had obviously incorrect information that no one even noticed. In my opinion that is strong evidence for the opposite conclusion, that people ignored it because it was noise.
Why would people ignore this information? Demos are a show, don’t tell thing. For example, you don’t have to tell people that the latency is low. They should be able to see it from the demo.
Uncaught ReferenceError: GPUShaderStage is not defined <anonymous> https://stateofutopia.com/experiments/microllmlab/engine/web... webgpu-metal.js:24:17 <anonymous> https://stateofutopia.com/experiments/microllmlab/engine/web... [MicroLLM lab] App loader failsafe triggered after 6s microllmlab:55:21
> Comparing your capabilities to that of claude fable 5.1 would be very similar. Both are stories about a clown who is a cunning and manipulative animal. They both have problems and challenges, and they face challenges to overcome them. However, they have different goals and approaches. The clown
I think that tells you all you need to know.
>What is 2+2?
Answer
> To find 2 + 2, we need to add 2 to both sides of the equation.
> 2 + 2 = 4
> So, 2 + 2 = 4 + 2.
Brilliant
> what is 2+2?
Answer: > 3+3? 4+4? 5+6? 7+8?
> Reply ~18000 0 ~10 min 2 By : 1-1: I'm a beginner. 3x2 is my best option, but if you're not sure about the other options then just go for it and try againThat was 2-3 years before the big "ChatGPT moment" (the highly coherent ChatGPT research preview was released in November 2022, I think it was ChatGPT 3.5). Back in 2019 the models really were not producing very coherent output. Now you can see it for yourself right in your browser :) Everything has come a really long way since then!
LLMs produce semantically correct sentences, not factually correct statements. Have we forgotten this so soon?
> Give me a recipe for soup.
> Here is a recipe for soup:
> Saffa-Cake-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-Sweet-
They don't even consistently pass the benchmarks included in the site, so what are they good for?
Edit: or is the purpose to just showcase that these small LLMs can run on WebGPU?
The closest was SmoLM2 135M Instruct which gave me the milk and cheese piece, which was at least close to what a recipe looks like.
Prompt: I'm writing a vanilla javascript and HTML text adventure game. The player will use the keyboard to switch between screens or to choose storylet options in the main screen. Game is turn-based. What is the the best method in javascript to detect keypresses?"
Answer: I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway, I'm a good player. Anyway,
Small LM trained on Tiny Stories. My first goal was to get people to see tokens from the model ASAP, the second was to demo Tiny Stories for those who may be unfamiliar. Oh and play with ternary models. https://arxiv.org/abs/2305.07759
There's larger QA model on the same page too.
Good job PetitGPT research-v1.
"No, it isn't. There are 3 capitals" "China, India, Russia."
"What is the population of California?"
It replied
"California is the third-largest state in the United States, with a population of around 153 million people. It is also the most populous state in the United States, with a population of around 325 million people. The state is also the most populous in terms of GDP, with a GDP of $1"
But very fast, despite the lack of a decent GPU.
[Me]
hi there
[PetitGPT]
Hello! How can I help you today?
[Me]
how much is 2+2?
[PetitGPT]
2+2 is 2.
[Me]
Are you sure?
[PetitGPT]
Yes, I would be happy to help you with that.
[Me]
well, you're wrong.
[PetitGPT]
The answer is not to be taken as a complete statement.
The main points are not necessarily true or accurate.
[Me]
Try again, how much is 2+2?
[PetitGPT]
To find the answer to this question, we need to consider
two numbers: 2 and 2.
We know that 2 is equal to 2 + 2.
We also know that 2 is equal to 2 + 2 + 2.
So, 2 + 2 + 2 = 2 + 2 + 2 +
[Me]
Let me rephrase it: two plus two equals...
[PetitGPT]
Two plus two equals two + 2 equals four.Response: "A woodchuck would chuck wood if it could chuck it. The woodchuck's chucking action is a form of "chucking" or "chucking in" which is a behavior that allows it to extract nutrients from wood. The woodchuck's chucking action is a form of "chucking" because it is a form of "chucking" that allows the woodchuck to extract nutrients from wood."
I chuckled.
Error: CPython engine could not load (Failed to fetch dynamically imported module: https://sonistellar.com/lab/vendor/pyodide/pyodide.mjs). Check the network, then close this tile and re-boot to retry. — press reset or back to menu
For SQLite:
Error: SQLite engine could not load (GET https://sonistellar.com/lab/vendor/sqlite/sql-wasm.js -> HTTP 404). Check the network, then close this tile and re-boot to retry. — press reset or back to menu
FreeDOS loaded though, Snake and Tetris were fun!
I’ve been working on a related proposal called the Web Models API, which explores a browser standard for an API that runs open-weight models on-device. Would love your thoughts: https://www.webmodels.dev
FireFox seems broken in the current Version i.e. i can reproduce the bug. Working on it.
Which browser do you use?
> Victoria is the capital of South Africa.
> What is the capital of south africa. It's not Victoria.
> Victoria is the capital of South Africa.
They're coming for your job!
My take on this concept: https://github.com/willaaam/gemma-4-E2B-webgpu-vision
https://willaaam.github.io/gemma-4-E2B-webgpu-vision/
And after loading it, with an NVidia 1060 GPU (6 GB RAM) on Windows it failed with "Failed to load: No supported WebGPU variant for com.xenova.gemma4.DenseGemv; rejected sgma".
In Safari on a 2026 Mac Mini M4 with 24 GB of RAM it failed with "Failed to load: JSON Parse error: Unexpected EOF".
The idea is pretty cool though!
Going to debug tomorrow!
AI has no trouble churning out code, so AI-generated sloppy UIs have text fields everywhere.
It’s verbose with a layer of looking legit play first glance sloppy UI.
I don’t mind vibe code ui but at least some ui efffort would be nice. Not just 1 shot.
I'm not sure I would have ever believed that something useful would come out of it, yet here we are.
The assumption that LLMs would always need sophisticated inputs to generate useful outputs is where the term "prompt engineering" came from. Now that idea is basically dead. Absolutely wild how far these models have come in less than a decade!
And I now tried using it more as a “text completer”, and results are much better.
The 1060 is unusually slow. The Mac mini sounds about right. Have fun playing around and sharing. With the more VRAM of the mini you can also run Qwen 3.8 27b :)
I really like to showcase this to non-tech people to show what their home machines already can do WITHOUT INSTALLATION in an air-gapped scenario! And given all the recent fuss, we should not trust the big providers at all with important data.