Why we write our own C and C++ inference engines(localai.io) |
Why we write our own C and C++ inference engines(localai.io) |
But the cpp port of vllm looks great, that'd be great if you'll maintain that. I hit the same limitations with vllm.
While the page looks vibe-coded[1], the content itself does not have any AI tells. What are the tells you are seeing?
[1] Too many sites I find on HN frontpage these days slow my PC to a crawl. I assume they are all using the same autogenerated HTML, Javascrip and CSS to make animated backgrounds :-( On this specific site scrolling is laggy.
> The method, the measurements, and what it costs us.
> That is the general shape of these wins.
> Parity is the gate, speed is the follow-up
I could go on and on, but you probably get the point. If you don't find anything funny with the above, you might have not been enough-exposed to slop.
Wasm size from 30Mb to 300kb and 1.5x speedup. It's definitely worth it for performance or distribution size.