drafting
3 stories and discussions about drafting, aggregated from every source we track.
1.
Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.
2.
3.