42x faster prompt lookup drafting in llama.cpp(jadidbourbaki.github.io) |
42x faster prompt lookup drafting in llama.cpp(jadidbourbaki.github.io) |
https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/comment...
Any advice for what I can do? Due to this, I cannot create a PR or issue in the llama.cpp repository. However, I am worried about bothering the maintainers on other channels in case it aggravates them further. Thank you for your help!
I’ll benchmark his change and add it to the article, crediting him for this improvement.