gpu
63 stories and discussions about gpu, aggregated from every source we track.
[Experimental] A virtio device for near-native NVIDIA GPU access in KVM virtual machines. - nestrilabs/virtio-nvgpu
The real difference between texture-atlas, SDF, MSDF, and Slug glyph rendering on the GPU: how each works, where each breaks, and when to use which.
There's a new kind of technical debt, and it doesn't come from cutting corners. It comes from...
Serving Gemma 4 E2B q4_0 through llama.cpp on one laptop, twice: CPU-only and on a 2021-era 4 GB GTX 1650 Ti. Same GGUF, same binary, same prompts, one flag apart. The card takes decode by 4.3x, and needs only 1598 MiB to do it.
Gemma 4 E2B q4_0 served by llama.cpp on one laptop, CPU-only and on a GTX 1650 Ti, rebuilt on CUDA 13.4 and re-measured in CPU, GPU, GPU, CPU order with a temperature gate. The card takes decode by 4.14x, and run order moves the answer by about 2%.
Step-by-step: running Google's quantization-aware-trained Gemma 4 E2B on a 10th-gen Core i7 laptop with a 4 GB GTX 1650 Ti — why bf16 and int8 cannot fit, why the QAT GGUF does with room to spare, and managing it with an MCP server.
Apple's container CLI 1.4.1 cannot boot the official debian:13 image as a machine, because it has no /sbin/init. A small Dockerfile fixes that and gives you a persistent Debian 13 VM with systemd, your Mac user and your home folder. The VM has no GPU, so Ollama runs on macOS and the VM calls it at 192.168.64.1:8000. From inside Debian, gemma4:e2b answered at 45.2 tok/s, loaded 100% on the M3 GPU.
How our Kimi K3 megakernel on TPU v7 reaches over 700 tokens/s with speculative decoding and nearly 2× GB200's batch-one decode throughput.
Maximizing perf on AI-SQL queries with the KV-optimal left-deep join
Nobody ever thinks of RISC-V when they think of Nvidia, but it's actually incredibly important to Nvidia's GPUs.
Bulk Synchronous Parallelism (BSP) in distributed deep learning clusters induces severe rate-of-change current transients (dI/dt) across data center power delivery networks. When thousands of accelerators synchronously…
[Experimental] A virtio device for near-native NVIDIA GPU access in KVM virtual machines. - nestrilabs/virtio-nvgpu
A GPU-accelerated Pedersen vector commitment engine for Ethereum Verkle trees (EIP-6800) - Dyslex7c/cuda-verkle
GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can…
Crop a line of text and get the closest fonts from Google Fonts, DaFont, Adobe Fonts, GitHub font repositories, Debian and more, with their weight and italic.
Glyd stores open models' bf16 weights in about 11 bits instead of 16 and decodes them inside the GPU's matrix multiply: bit for bit, often faster.
VRAM-aware single-node GPU job scheduler. Contribute to ceruleane/gpusched development by creating an account on GitHub.
NEO Emacs (WIP): GPU powered Emacs written in Rust with a modern display engine. Aiming for modern design & multi-threaded Elisp, 10x performance, zero-pause concurrent GC and 100% Emacs compat...
A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex
We won 2nd overall at HackMIT for building a GPU in 24 hours.
A modern LLM can spend most of its time doing something that looks almost embarrassingly simple: C...
GPU-Rendered UI library built from scratch in Rust [WIP] - artemtsitronov/glacex
Five years of GPU infrastructure at Beam — from ECS and Knative cold starts to a custom container runtime, FUSE lazy-loading, and a trustless binary.
Stop wasting hours waiting for video to be built from the web page in a browser. Vibe code your video on GPU using Rust & SVG framework.
Institutional-grade real-time AI cloud GPU spot rates, cluster arbitrage, and inventory intelligence across 31 global cloud providers.
Maximizing perf on AI-SQL queries with the KV-optimal left-deep join
In gory detail: reliability, performance, support, pricing—and, of course, security—in our most thorough analysis of GPU cloud providers globally.
Transcript of Onur Satici's QCon London 2026 talk on streaming Vortex files from S3 into GPU memory.
tmux for GPUs: share one GPU in space and time. Runs anywhere Docker does. - numinous-technology/gmux
Explore real-time GPU cloud capacity across providers and regions. Track availability scores, compare instance types, and find GPU resources — powered by Datadog.
Explore real-time GPU cloud capacity across providers and regions. Track availability scores, compare instance types, and find GPU resources — powered by Datadog.
Price a whole AI project against buying the cards: published rental rates per GPU-hour, your purchase quote, and electricity at the vendor's rated board power. Shows the price a card would have to hit for buying to win.
Fast, lightweight micro virtual machine for cloud streaming — GPU included - nestrilabs/nesbox
Lifeboat — downloads for macOS, Windows and Linux, plus Docker and Kubernetes install instructions. Run language models on your own hardware behind an OpenAI-compatible API. - IterateAI/lifeboat-re...
I built Vitest and Jest environments that give your tests real, GPU-backed WebGL and WebGPU contexts directly in Node.js. They use the same ANGLE and Dawn implementations as Chrome, so you get the same pixels without…
An interactive NVIDIA-GPU process viewer and beyond, the one-stop solution for GPU process management. - XuehaiPan/nvitop
Run PyTorch from your laptop on a remote GPU. Setup, training, CUDA remoting, and practical performance guidance.
See the wind that sends climbers home: the next 72 hours at the summit of 32 mountaineering objectives, read at summit height across several weather models and checked against summit stations.
Native Hardware Serving & Continuous Batching Acceleration for Dense LLMs (CIPO CA 3,322,620) - cortexLab011/floria-serving
If you've ever had a training job ready to go and then sat there refreshing your cloud console...
Contribute to Themaister/pyrowave development by creating an account on GitHub.
The ASSAY-1 specification, the graded GPU capacity registry, and the weekly Delivered Compute Report.
Contribute to tudormunteanu/gpu-cloud-instance-boostraps development by creating an account on GitHub.