gpu

63 stories and discussions about gpu, aggregated from every source we track.

1.

[Experimental] A virtio device for near-native NVIDIA GPU access in KVM virtual machines. - nestrilabs/virtio-nvgpu

144 points•WanjohiRyan•7 days ago•59 comments•
2.

The real difference between texture-atlas, SDF, MSDF, and Slug glyph rendering on the GPU: how each works, where each breaks, and when to use which.

72 points•ibobev•about 11 hours ago•35 comments•
3.

There's a new kind of technical debt, and it doesn't come from cutting corners. It comes from...

25 points•cyclopt_dimitrisk•3 days ago•21 comments
4.

Serving Gemma 4 E2B q4_0 through llama.cpp on one laptop, twice: CPU-only and on a 2021-era 4 GB GTX 1650 Ti. Same GGUF, same binary, same prompts, one flag apart. The card takes decode by 4.3x, and needs only 1598 MiB to do it.

15 points•xbill•14 days ago•6 comments
6.

Gemma 4 E2B q4_0 served by llama.cpp on one laptop, CPU-only and on a GTX 1650 Ti, rebuilt on CUDA 13.4 and re-measured in CPU, GPU, GPU, CPU order with a temperature gate. The card takes decode by 4.14x, and run order moves the answer by about 2%.

10 points•xbill•7 days ago•2 comments
7.

Step-by-step: running Google's quantization-aware-trained Gemma 4 E2B on a 10th-gen Core i7 laptop with a 4 GB GTX 1650 Ti — why bf16 and int8 cannot fit, why the QAT GGUF does with room to spare, and managing it with an MCP server.

10 points•xbill•20 days ago•3 comments
8.

Apple's container CLI 1.4.1 cannot boot the official debian:13 image as a machine, because it has no /sbin/init. A small Dockerfile fixes that and gives you a persistent Debian 13 VM with systemd, your Mac user and your home folder. The VM has no GPU, so Ollama runs on macOS and the VM calls it at 192.168.64.1:8000. From inside Debian, gemma4:e2b answered at 45.2 tok/s, loaded 100% on the M3 GPU.

9 points•xbill•17 days ago•1 comment
9.

How our Kimi K3 megakernel on TPU v7 reaches over 700 tokens/s with speculative decoding and nearly 2× GB200's batch-one decode throughput.

5 points•xutingl•7 days ago•1 comment•
11.

Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

4 points•charles_irl•4 days ago•1 comment•
12.
4 points•winwang•5 days ago•0 comments•
14.

Nobody ever thinks of RISC-V when they think of Nvidia, but it's actually incredibly important to Nvidia's GPUs.

4 points•giuliomagnifico•11 days ago•0 comments•
15.

Bulk Synchronous Parallelism (BSP) in distributed deep learning clusters induces severe rate-of-change current transients (dI/dt) across data center power delivery networks. When thousands of accelerators synchronously…

4 points•samganyx•11 days ago•0 comments•
17.
3 points•jinhongyii•1 day ago•0 comments•
18.
3 points•sonabinu•3 days ago•0 comments•
19.

[Experimental] A virtio device for near-native NVIDIA GPU access in KVM virtual machines. - nestrilabs/virtio-nvgpu

3 points•jeanthomas•5 days ago•0 comments
20.

A GPU-accelerated Pedersen vector commitment engine for Ethereum Verkle trees (EIP-6800) - Dyslex7c/cuda-verkle

3 points•grobat79•7 days ago•0 comments•
21.
3 points•thinkevolve•11 days ago•0 comments•
22.
2 points•ksec•about 19 hours ago•0 comments•
23.

GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can…

2 points•matt_d•about 20 hours ago•0 comments•
25.
2 points•ibobev•2 days ago•0 comments•
26.
2 points•mattyw•3 days ago•0 comments
27.

Crop a line of text and get the closest fonts from Google Fonts, DaFont, Adobe Fonts, GitHub font repositories, Debian and more, with their weight and italic.

2 points•made_by_dy•3 days ago•0 comments•
28.
2 points•gtsnexp•3 days ago•0 comments•
29.

Glyd stores open models' bf16 weights in about 11 bits instead of 16 and decodes them inside the GPU's matrix multiply: bit for bit, often faster.

2 points•surya-koritala•3 days ago•0 comments•
30.

VRAM-aware single-node GPU job scheduler. Contribute to ceruleane/gpusched development by creating an account on GitHub.

2 points•mothproof•4 days ago•0 comments•
31.

NEO Emacs (WIP): GPU powered Emacs written in Rust with a modern display engine. Aiming for modern design & multi-threaded Elisp, 10x performance, zero-pause concurrent GC and 100% Emacs compat...

2 points•tagfowufe•4 days ago•2 comments•
32.

A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time 'System 1' agent decision loops. - lateos-ai/reflex

2 points•leochong•5 days ago•0 comments•
33.
2 points•AbuAssar•6 days ago•0 comments•
34.

We won 2nd overall at HackMIT for building a GPU in 24 hours.

2 points•arghunter•6 days ago•0 comments•
35.

A modern LLM can spend most of its time doing something that looks almost embarrassingly simple: C...

2 points•shrsv•8 days ago•0 comments
36.

GPU-Rendered UI library built from scratch in Rust [WIP] - artemtsitronov/glacex

1 points•adamnemecek•about 8 hours ago•0 comments•
37.

Five years of GPU infrastructure at Beam — from ECS and Knative cold starts to a custom container runtime, FUSE lazy-loading, and a trustless binary.

1 points•Mernit•1 day ago•0 comments•
38.

Stop wasting hours waiting for video to be built from the web page in a browser. Vibe code your video on GPU using Rust & SVG framework.

1 points•neogoose•3 days ago•0 comments•
39.

Institutional-grade real-time AI cloud GPU spot rates, cluster arbitrage, and inventory intelligence across 31 global cloud providers.

1 points•nurturely•3 days ago•0 comments•
40.
1 points•signa11•3 days ago•0 comments•
41.

Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

1 points•birdculture•3 days ago•0 comments•
42.

In gory detail: reliability, performance, support, pricing—and, of course, security—in our most thorough analysis of GPU cloud providers globally.

1 points•nathanscully•4 days ago•0 comments•
43.
1 points•cromka•4 days ago•0 comments•
45.

Transcript of Onur Satici's QCon London 2026 talk on streaming Vortex files from S3 into GPU memory.

1 points•surprisetalk•6 days ago•0 comments•
46.

tmux for GPUs: share one GPU in space and time. Runs anywhere Docker does. - numinous-technology/gmux

1 points•joshkolo•6 days ago•0 comments•
47.

Explore real-time GPU cloud capacity across providers and regions. Track availability scores, compare instance types, and find GPU resources — powered by Datadog.

1 points•hwayne•6 days ago•0 comments•
48.

Explore real-time GPU cloud capacity across providers and regions. Track availability scores, compare instance types, and find GPU resources — powered by Datadog.

1 points•hwayne•6 days ago•1 comment
49.

Price a whole AI project against buying the cards: published rental rates per GPU-hour, your purchase quote, and electricity at the vendor's rated board power. Shows the price a card would have to hit for buying to win.

1 points•the_app_guy_1•7 days ago•0 comments•
50.

Fast, lightweight micro virtual machine for cloud streaming — GPU included - nestrilabs/nesbox

1 points•WanjohiRyan•7 days ago•0 comments•
51.

Lifeboat — downloads for macOS, Windows and Linux, plus Docker and Kubernetes install instructions. Run language models on your own hardware behind an OpenAI-compatible API. - IterateAI/lifeboat-re...

1 points•chanuiterate•7 days ago•0 comments•
52.

I built Vitest and Jest environments that give your tests real, GPU-backed WebGL and WebGPU contexts directly in Node.js. They use the same ANGLE and Dawn implementations as Chrome, so you get the same pixels without…

1 points•bhouston•7 days ago•0 comments•
53.

An interactive NVIDIA-GPU process viewer and beyond, the one-stop solution for GPU process management. - XuehaiPan/nvitop

1 points•Nina_antalpha•8 days ago•1 comment•
54.

Run PyTorch from your laptop on a remote GPU. Setup, training, CUDA remoting, and practical performance guidance.

1 points•boxstream•8 days ago•0 comments•
55.

See the wind that sends climbers home: the next 72 hours at the summit of 32 mountaineering objectives, read at summit height across several weather models and checked against summit stations.

1 points•duranduran•8 days ago•0 comments•
56.

Native Hardware Serving & Continuous Batching Acceleration for Dense LLMs (CIPO CA 3,322,620) - cortexLab011/floria-serving

1 points•cortexlab1•8 days ago•0 comments•
57.

If you've ever had a training job ready to go and then sat there refreshing your cloud console...

1 points•sumukh_dev•9 days ago•1 comment
58.

Contribute to Themaister/pyrowave development by creating an account on GitHub.

1 points•open-paren•9 days ago•1 comment•
59.

The ASSAY-1 specification, the graded GPU capacity registry, and the weekly Delivered Compute Report.

1 points•RentAnAgent•9 days ago•0 comments•
60.

Contribute to tudormunteanu/gpu-cloud-instance-boostraps development by creating an account on GitHub.

1 points•tudorizer•9 days ago•0 comments•

Related topics