The Next 3× in Inference Won't Come From Faster Kernels

2 points•floathub•1 day ago•1 comment•

1 comment

floathub1 day ago
One unique idea for solving the chip shortage: pack more models into each existing GPU by having them share the resources. :-)

Read the full thread on Hacker News →

Related stories