drafting

3 stories and discussions about drafting, aggregated from every source we track.

1.

Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.

70 points•pptadversary•4 days ago•11 comments•
2.
2 points•Exoristos•9 days ago•0 comments•
3.
1 points•andsoitis•10 days ago•0 comments•

Related topics