llama
15 stories and discussions about llama, aggregated from every source we track.
Four changes to the n-gram caches of llama.cpp make drafting up to 41.6x faster, load the static cache up to 23.5x faster, and lower peak memory up to 2.65x.
(Edit: apologies, I should have clarified initially I'm running on Linux OS. I didn't realize it might not be obvious from the screenshot alone for a non-Linux users.All tests are done on Ubuntu ba...
I spent a weekend benchmarking llama.cpp on my Xiaomi Book Pro 14. The machine has Intel's Core Ultra X7 358H and Arc B390 integrated graphics, with 30 GiB of unified memory. That last part matters. It means the GPU…
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
jeva.cpp - a llama.cpp fork with a JEV-compatible decision API for all LLMs supported by llama.cpp. - PragmaTwice/jeva.cpp
How to point Oh My Pi at vLLM, llama.cpp, SGLang, and other local servers or a gateway, plus two config fixes for omp 18.2.7 and later.
Static analysis of llama.cpp with CppDepend: Code City hotspots, GRASP-based module decoupling, C-style POD design, minimal STL footprint, and why pragmatism beats dogma in high-performance C++.
turn any llama-server into a jev system one endpoint - khimaros/verdict
A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex
Agentic gateway providing persistent LLM memory via a markdown wiki. Open-source core of Kortexio. - Kortexio/ContextMemory
Fork off LLama cpp that newly scales GPU 2x 4x etc also on big models not fully fitting in vram - neurall/llama.cpp
llama.cpp recently added the ability to control the output of any model using a grammar.
We’re on a journey to advance and democratize artificial intelligence through open source and open science.