cuda
6 stories and discussions about cuda, aggregated from every source we track.
1.
A langgraph based workflow with a C++ CUDA suite to optimize CUDA kernels - bertaye/agentic-cuda-optimizer
2.
Gemma 4 E2B q4_0 served by llama.cpp on one laptop, CPU-only and on a GTX 1650 Ti, rebuilt on CUDA 13.4 and re-measured in CPU, GPU, GPU, CPU order with a temperature gate. The card takes decode by 4.14x, and run order moves the answer by about 2%.
3.
A GPU-accelerated Pedersen vector commitment engine for Ethereum Verkle trees (EIP-6800) - Dyslex7c/cuda-verkle
4.
5.
6.
A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex