Peek at an image, get booleans back: typed visual questions (choice, score, yes/no) with calibrated probabilities from a 500M VLM, ~400 ms on a MacBook - bykof/peekaboolean
I trained a vision-language model that answers typed questions about an image. It uses also choice, score and noul, like Jev.
I thought, why Jev only processes text? There should be the possibility to process an image too. I will experiment with some pictures and questions in the next days to see, how well it performs in real life.
On my M1 Pro a request with six questionut 400 ms p95, about 60 ms on a desktop GPU. The server encodes the image once and scores each option as a short suffix against the KV
Looking forward to answer your questions! :)
0 comments
No comments yet.
Related stories
- The Verge · 0 points · 5 days ago
- The Verge · 0 points · 3 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Lobsters · 86 points · about 1 year ago
- The Verge · 0 points · 12 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago