Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.
12 comments
https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/comment...
Any advice for what I can do? Due to this, I cannot create a PR or issue in the llama.cpp repository. However, I am worried about bothering the maintainers on other channels in case it aggravates them further. Thank you for your help!
Email that author (email can usually be found via GitHub, or contact via other private way they've shared somewhere) and explain the situation, don't lambast them publicly on social media or similar ways. If that doesn't work, do the same but for another maintainer. Don't spam all of them straight up, wait a week or something before contacting someone else.
> Emailing me about a temporary ban will result in a permanent one instead.
As far as I can tell the ban was given based on a single accidental ping to the maintainer, who then reviewed the draft PR before realizing it was not for the mainline llama.cpp repo.
I’ll benchmark his change and add it to the article, crediting him for this improvement.
https://github.com/jadidbourbaki/llama.cpp/pull/12#issuecomm...
https://news.ycombinator.com/item?id=49859982#49863097
but commented on the post.
Read the full thread on Hacker News →
Related stories
- Hacker News · 2 points · 2 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 4 points · 3 days ago
- Llama.cpp Under the Hoodcppdepend.comHacker News · 2 points · 8 days ago
- Hacker News · 2 points · 5 days ago
- Using Llama-cpp-Python grammars to generate JSONtil.simonwillison.netHacker News · 1 points · 8 days ago