Felt extra autistic today and made jev an autoregressive chatbot by classifier small token vocabularies. https://github.com/giga-james/jevgpt had some trouble getting it to explore enough token…
https://github.com/giga-james/jevgpt
had some trouble getting it to explore enough token options so i did a small spec decoding pass to expand token vocab (jev as drafter and verifier lol)
it aint much and it aint honest work either
4 comments
Where do you see the biggest potential improvement gains?
the underlying mechanism has to sample tokens (1000/90000 for openai's tokenizer) by some shallow heuristic to build the decision pool (more details are in https://github.com/giga-james/jevgpt/blob/main/docs/experime... !), while standard autoregressive models have access to the entire token vocabulary to predict the next word
jev allows 255 decisions (token vocabulary here), so if we want to support a token vocabulary of 1000, we have to parallelize into multiple jev calls to allow that expansion. the verifier pass simply chooses the best out of the parallel decisions so we can exceed the 255 decision limit. in theory this should allow for up to 255^2 token limit, but if you increase the verifier depth you can increase that exponential factor as much as you want (it'll be slow though)
Read the full thread on Hacker News →
Related stories
- Hacker News · 3 points · 3 days ago
- Hacker News · 1 points · 6 days ago
- Hacker News · 1 points · 2 days ago
- Hacker News · 421 points · 6 days ago
- Hacker News · 1 points · 4 days ago
- Hacker News · 2 points · about 11 hours ago