Deployment runtime for VLA policies. Contribute to Agents2AgentsAI/vla-edge development by creating an account on GitHub.
On MolmoAct2, a 30-action chunk went from 611 ms to 113 ms. Most of that comes from CUDA graphs and TensorRT. The last 1.5x was from the agentic search. No distillation or pruning was used.
On the recently published ABC-VLA model, the optimized solution is 32.4 ms compared to the TensorRT conversion at 63.4 ms. A significant part of this gain came from a lossless weight decoder, similar to DFloat11 and ZipServ. This was done automatically by the agents without any human guidance. Again, this is without any distillation or pruning. The plans are bf16/fp16, and no 8-bit or 4-bit quantization.
More information is in the README and blog posts. Please let us know if you have any problems reproducing on your hardware.
0 comments
No comments yet.
Related stories
- Hacker News · 1 points · 10 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 11 points · 10 days ago
- Hacker News · 1 points · 1 day ago
- Rho – A Foundation for Efficiently Adaptable VLA Modelsmicrosoft.github.ioHacker News · 1 points · about 1 hour ago
- Designing a Programming Language for the AI Erazena-lang.devHacker News · 1 points · 1 day ago