A langgraph based workflow with a C++ CUDA harness to optimize CUDA kernels - bertaye/agentic-cuda-optimizer
Hello; I was working on optimizing some CUDA kernels and I thought may be it is a good oppurtunity learn langgraph as well. I created a simple C++ CUDA Test Harness and handed that to AI agents. They can run kernels, get benchmarks, and even can profile via nsight
11 comments
fooblaster6 days ago
Can someone explain why this isn't better accomplished through a single prompt to Claude code or codex? I don't think I understand.
bertaye5 days ago
Hello, indeed you can just use that.
The basic idea here is just automating and limiting the steps that AI can take. These are described as ‘nodes’ and their actions are limited/more descriptive from developer perspective.
The langgraph simply allows you to set some fences around the AI agent for a goal, instead of raw terminal flow. Is it better? Arguable.
osti6 days ago
Yup, this thing is basically useless. I just did /goal optimize the cuda kernel, and ai agent just proposed and tested a bunch of ideas by itself, and it profiled them using nsight ncu etc. by itself, which lead to one order of magnitude faster kernel.
generalizations6 days ago
Very cool. Did you also try using the karpathy autoresearch? How do you think this compares?
bertaye6 days ago
honestly I know it exists but I never used it so can't compare
aidiveyt5 days ago
in mine, only an agent's final report reaches the orchestrator, never its transcript, so a wrong turn inside a node stays invisible. that's the fence i'd want first.
Read the full thread on Hacker News →
Related stories
- Hacker News · 4 points · 8 days ago
- Hacker News · 1 points · 5 days ago
- Hacker News · 1 points · 6 days ago
- Hacker News · 7 points · 8 days ago
- Hacker News · 1 points · 10 days ago
- Hacker News · 1 points · about 13 hours ago