78 comments
It had this to say in the linked post:
This led to the natural question: can gzip do language modeling? (...). Here’s some real, unedited output after priming it on tiny Shakespeare:
gzipt --corpus data/tinyshakespeare.txt --prompt $'MENENIUS:\n' --length 200
MENENIUS:
'Though all at once canq
MARCIUS:
Pray now, nocamest thou to a morsel.
LARTIUS:
Hence, and
I' the end admire, where G
again; and after it ag .
Now thinking back, what's missing so that gzip could unwind the correct body of work from Shakespeare is just a correct sequence of bytes. One way to arrive at this is by just getting the body of work and doing the inverse, compressing it to get that golden sequence of bytes.The other is what thinking does, it tries to predict the missing sequence of tokens from a high entropy source, the prompt, in order to increase the likelihood of correctly decompressing the desired results from its weights.
Like for text, what would that involve? How do you compress a string or multi-line string without losing information and hopefully structure (paragraphs, would it be like replacing periods and the following space with just sticking the starting capitalized letter of the following word to the previous sentence's last letter and when it decompresses theres some kind of note that converts that back into the. First letter of the next sentence
How can things be compressed without losing information or structure?
Because the initial content is rarely the most efficient representation, so it's possible to store fewer bytes that can deterministically be converted into the original. Like for text, what would that involve?
Most compression algos don't care what information you're compressing. All they see (all they need to see) is bytes. It ends up being way more sophisticated than removing repeated periods and whitespace.Like if you had eight boxes of loose lego, simply shuffling around the boxes wouldn't give you much in the way of reducing the space the legos take up. but if you took the legos (bytes) themselves out of the boxes, you end up saving a lot more space.
In many ways, we are communicating using a compressed channel (words) since we both have pre-agreed on the meaning of these words.
``` th #1 - numbers are just comments ng #2, this has a space at the end information #3 compress #4 letter #5 this has a space at the start and #6 this has a space at the start and end sentence #7 this has a space at the start
How can 1ings be 4ed without losi23 or structure?
Like for text, what would 1at involve? How do you 4 a stri2or multi-line stri2wi1out losi236hopefully structure (paragraphs, would it be like replaci2periods61e followi2space wi1 just sticki21e starti2capitalized5 of 1e followi2word to 1e previous7's last56when it de4es 1eres some kind of note 1at converts 1at back into 1e. First5of the next7 ```
I'm on a phone, so I may have mistakes here, but I'm pretty sure that's shorter than your original text, in bytes, by about (9+18+20+21+12+12+16=108), minus the dictionary size of 51 -- so, 57 bytes shorter, but still containing your full text. With predistributed compressor binaries and a lot of analysis, you can even predistribute a global dictionary for common sequences, and simply specify "xyz0", x, y, and z being 24-bit numbers, or whatever bit size can index into your full reference dictionary, and 0 meaning end-of-file-dictonary. then, assuming the byte sequences in your text above are common enough to be in the 24-bit indexed dictionary, that initial dictionary could be just 22 bytes (21 and a terminator) -- so, 86 bytes shorter than the original, but still containing your original message unaltered. ..assuming i didn't make mistakes in my hand-compression.
I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.
I know a large organization who's built their AI framework completely around this concept, and I feel that it's not really meaningful concept with the capabilities of current models.
Do we? I just learned from a speaker[1] that we literally need words to recognize emotions. People who have a poor vocabulary have lower emotional intelligence because without being able to attach a word to an emotion, the brain is unable to recognize & process it.
[1] Dude seemed to be knowledgeable about the subject. He's a specialized trainer, should be educated in this exact field. So hopefully I'm not lying to anyone here :)
'slow' means making one or several action-dependent forecasts, evaluating the expected value of the outcomes, and making a decision based on that.
Neither map exactly to the situation with LLMs, but very roughly, the first is analogous to trained classifiers and the second to reasoning models.
The analogy breaks down, since each instance of token being produced is an example of a policy execution (system 1), and reasoning is just stringing lots of these together. But there are those who argued, before LLMs, that system 2 is just "policy composition" anyway...
That seems to fit the fast vs slow model of human thought reasonably well.
That’s still several orders of magnitudes too slow to fit fast vs slow. Think of 30ms vs 3-4 seconds to get an idea of what we’re talking about here
Of course the shapes of what an AI can do in fast vs slow are quite different.
So the entire debate is fubar.
Meta-cognition makes sense in a dynamic and updatable and modular system, for example I can monitor thoughts coming from my amygdala with my prefrontal cortex and then adjust how I process these thoughts.
In LLMs it makes zero sense, even if you feed the output of one model into another, there is no way they can update the heuristics behind how those were computed.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 3 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 5 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Hacker News · 73 points · 9 days ago
- The Verge · 0 points · 12 days ago