134 points•vismit2000•11 days ago•122 comments•

122 comments

2shortplanks11 days ago
On a practical matter, it seems like a bad idea to have codepoints that can take up to an arbitrary number of bytes - this just screams buffer overflow problems.

So in practicality, you’re going to want an arbitrary limit on this (the article suggests as much). But if you place a limit on it then you’ve got one implementation of the standard that can decode certain characters and another that can’t. Better to have one standard that puts a hard limit on the number of bytes and another standard that uses more bytes and so on.

Pannoniae10 days ago
You don't have a buffer overflow problem if you read it in a memory-safe way i.e. read it in chunks and realloc when you reach the size of your allocation.

What you will have is a potential denial-of-service attack - although this one isn't particularly great because there's zero amplification (they might as well just send garbage into your firewall)

zbentley9 days ago
I think the DoS has pretty common amplification vectors in the form of APIs that split or otherwise copy (e.g. materializing code points for Unicode regex searching).

Additionally, OOM inside a low level routine can be a troublesome attack, since OOM handling in many applications does questionable (nee vulnerable) things when crashes occur in not-known-to-be-memory-intensive code. Sure, that’s sloppy engineering, but it’s common.

LorenPechtel9 days ago
Memory allocation?? Why? Real world, you'll bump into limits based on fonts long before you'll get a buffer that's too large for the stack and worth using the memory allocator for. I can't see any reason to support more than 2^64 characters and lots of headaches from trying to beyond that. You check your buffer writes and reject the character if it's too long.
torgoguys10 days ago
DOS in what way? Can you clarify? Thx.
Retr0id10 days ago
In regular unicode, a grapheme can be made up of an arbitrary number of codepoints (and thus an arbitrary number of bytes), which does cause issues at times.
flohofwoe11 days ago
OTH UTF-8 is just one variable-length stream encoding among many others (RLE, LBE128, etc...).
DmitryOlshansky10 days ago
The bonus is synchonizing at arbitrary point in stream and that ASCII is UTF-8
saghm10 days ago
Unless I'm misremembering, even UTF-16 is variable. You need to bump up to UTF-32 to get fixed-width.
explodes10 days ago
Limit the codepoint to the number of atoms in the universe (less than 32 bytes).
__david__10 days ago
I don’t think there’s harm is speccing out the arbitrary encoding and then having a different spec that references that spec but puts hard limits on it. Many rfcs are like that.
sph11 days ago
> UTF-8000 is in no way endorsed by or representative of the Unicode Consortium.

Not until they decide to expand the emoji range, allocate space for all past and future fictional languages, as well as birdsong and dog barks.

Someone at the consortium is rubbing their hands with glee with all the newfound space.

But honestly, cool hack! If you invent a method to encode large numbers into bytes, why limit yourself to 24-bit numbers?

throw0101a11 days ago
> […] why limit yourself to 24-bit numbers?

For compatibility with UTF-16:

    o  Restricted the range of characters to 0000-10FFFF (the UTF-16
       accessible range).
* https://datatracker.ietf.org/doc/html/rfc3629#section-12

* https://en.wikipedia.org/wiki/UTF-16

The original spec had 31 bits (the UTF-32/UCS-4 range):

* https://datatracker.ietf.org/doc/html/rfc2279

* https://en.wikipedia.org/wiki/UTF-32

orangeboats9 days ago
We really ought to deprecate UTF-16 someday. The fact that it pretends to be a fixed-length encoding has caused all sorts of bugs over the years, with many people assuming n(UTF-16 codepoints) == n(characters) which breaks when the string contains non-BMP characters.

And also, for personal aesthetic reasons I hate that it limits the Unicode codepoint range to an awkward non-power-of-two number (now there are 0x110000 codepoints in total). UTF-8 and UTF-32's 2^31 feels much more natural.

flohofwoe11 days ago
> ...24-bit numbers?

Technically current UTF-8 only goes up to 21 bits (that's the current UNICODE range), for the encoding itself that is an arbitrary limit though, with the 'single lead byte' method of traditional UTF-8 it could go up to 36 bits "payload".

sharktheone11 days ago
I think that wouldn't change much. They would just make use of more grapheme clusters.

For emojies they already make heavy use of the Zero-Width-Joiner. So a woman firefighter is the woman emoji + ZWJ + fire engine. Sure the UTF-8000 approach is much better encoding size wise.

mitxela11 days ago
I wonder how they're going to encode a female fire engine in the future.
TeMPOraL10 days ago
I predict eventual convergence between UTF-whatever and most popular tokenizer for whatever LLM escapes to become world-ruling AGI.

I mean, if someone's seriously going to try encoding birdsong and dog barks, at this point they're basically reinventing tokens for multi-modal language models.

Sharlin11 days ago
UTF-8 originally supported up to six-byte encodings (see eg. RFC 2279), but it was restricted to four bytes in 2003 in order to match UTF-16 constraints :(
delamon11 days ago
We still have about 85% of codepoint space unused. Hopefully, by the time it becomes a problem, UTF-16 will be long dead
nasso_dev11 days ago
i hope so too, but UTF-16 being used by languages such as java and javascript makes me fear it might be here to stay.... i hope im wrong
colejohnson6610 days ago
But by then, the 4-byte limit of UTF-8 will itself have ossified. Even today, reverting back to the 6-byte limit is nigh impossible.
yyyk11 days ago
Just limit it to 8 bytes at which point you always do 'know the number of follow on bytes' from the first byte.

Nobody needs more than 4.47 trillion characters. (famous last words)

mitxela11 days ago
Important to recognize that characters have individuality, that's why there can only be a limited number of them. Unicode is enumerating a finite set of things, not encoding an infinite set. Aenything without this property - any generic form of encoding - is not characters, it's something else like images. If it's not in any alphabet it shouldn't be in unicode, you should use an escape tag for image data instead. (Emojis probably shouldn't, but they do behave like an alphabet)

There cannot be 4 trillion characters because humans would need to know all of them and humans cannot know that many things.

stbenjam11 days ago
> No special cases introduced. All properties preserved.

I don’t actually know if this is LLM-generated, but phrasing like this is weirdly triggering to me now

Neywiny11 days ago
Yeah that kind of line is what I see all the time in my chats. Even worse worse is when they put it in code comments.
account4210 days ago
It's not even true as being able to tell the character length from the first byte is not a property that extending UTF-8 past 36 bit payloads preserves.

Read the full thread on Hacker News →

Related stories