122 comments
So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.
What you will have is a potential denial-of-service attack - although this one isn't particularly great because there's zero amplification (they might as well just send garbage into your firewall)
Additionally, OOM inside a low level routine can be a troublesome attack, since OOM handling in many applications does questionable (nee vulnerable) things when crashes occur in not-known-to-be-memory-intensive code. Sure, that’s sloppy engineering, but it’s common.
Not until they decide to expand the emoji range, allocate space for all past and future fictional languages, as well as birdsong and dog barks.
Someone at the consortium is rubbing their hands with glee with all the newfound space.
But honestly, cool hack! If you invent a method to encode large numbers into bytes, why limit yourself to 24-bit numbers?
For compatibility with UTF-16:
o Restricted the range of characters to 0000-10FFFF (the UTF-16
accessible range).
* https://datatracker.ietf.org/doc/html/rfc3629#section-12* https://en.wikipedia.org/wiki/UTF-16
The original spec had 31 bits (the UTF-32/UCS-4 range):
And also, for personal aesthetic reasons I hate that it limits the Unicode codepoint range to an awkward non-power-of-two number (now there are 0x110000 codepoints in total). UTF-8 and UTF-32's 2^31 feels much more natural.
Technically current UTF-8 only goes up to 21 bits (that's the current UNICODE range), for the encoding itself that is an arbitrary limit though, with the 'single lead byte' method of traditional UTF-8 it could go up to 36 bits "payload".
For emojies they already make heavy use of the Zero-Width-Joiner. So a woman firefighter is the woman emoji + ZWJ + fire engine. Sure the UTF-8000 approach is much better encoding size wise.
I mean, if someone's seriously going to try encoding birdsong and dog barks, at this point they're basically reinventing tokens for multi-modal language models.
Nobody needs more than 4.47 trillion characters. (famous last words)
There cannot be 4 trillion characters because humans would need to know all of them and humans cannot know that many things.
I don’t actually know if this is LLM-generated, but phrasing like this is weirdly triggering to me now
Read the full thread on Hacker News →
Related stories
- Hacker News · 6 points · 8 days ago
- Hacker News · 3 points · 8 days ago
- Hacker News · 6 points · 1 day ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 2 points · 10 days ago
- Hacker News · 2 points · 4 days ago