I built my own binary serializer and a custom lang. heres how I did it

40 points•BurnerBurner•3 days ago•17 comments•

17 comments

trashb3 days ago
I'm missing some important parts that I would argue any (binary)format needs.

  - a magic header
  - a version number
  - a crc or data corruption check
Additionally I would argue one would be better off writing a custom text parser instead of parsing a binary format in this case, this is similar to the debate about unixlike config vs windows regedit. I would prefer something other then json but sill readable as txt.

Even the json example provided at the bottom can be minified from 418 characters to 166 by replacing the field names with single characters and removing the spaces. Almost all of the savings in this format come from not including the the field names and having a position dependent layout. You can choose to have delimiters or arrange for a byte to indicate the type(+size) for example the following string encodes the example data almost (85bytes vs 80bytes) as efficient but is still readable and supports utf-8 interpretation.

  123456789;LeroyJenkins;60;alliance;p,100,200,300;i,999,1,1;i,45,100,0;a,s,120;a,a,45;
You may optimize it further by allowing recurring entries and allowing assumed values from a defined default and only sending delta's can be dependent on the type of data you are expecting.

  { "itemId": 999, "quantity": 1, "isSoulbound": false }
  i,999,1,0
could become:

  default = { "itemId": 0, "quantity": 1, "isSoulbound": false }
  { "itemId": 999}
  i,999
orf2 days ago
> You may optimize it further

You should just use protobuf, far before you get to this step

ErikHuisman3 days ago
I also want to be a real developer so i just GZIP the JSON to make it binary.
nacozarina2 days ago
Yeah, I’m trying to imagine how this is superior to simply adding ZIP’d JSON support to your app.
entrope2 days ago
It depends. My personal experience is with RINEX, a text file format for GNSS observation data. gzip compresses large sets of files by about 4x. Someone named Yuki HATANAKA came up with a compact text format (CRX) that shrinks it by about a factor of 4.1x. You can gzip CRX for a total ratio of almost 12x, or bzip3 it for a total ratio of 17x. But if you combine Hatanaka's technique (essentially higher-order delta coding) with some others, you get a binary format with a compression ratio of 19x that reads faster than the CRX text.

When one's uncompressed RINEX files are 200 GB/day, it's worth some effort to shrink the files, especially if it means they are faster to read.

masklinn2 days ago
If you want to be a real developer you should use raw deflate (or atleast zlib streams), they're way harder to identify so they look more really developpery.

Also gzip adds something like a dozen unnecessary bytes.

jareklupinski2 days ago
just Poob it
kstenerud3 days ago
Schemas save you space, but then you lose the ability to understand the data without the schema. That's where JSON has always been handy despite its inefficiencies.

Of course you can get the same kind of thing in binary. I wrote a drop-in binary JSON replacement because it's easy to write a binary one that's twice as fast as simjson and yyjson. The important thing is to never give the drop-in replacement any extras that break roundtrip compatibility.

Was a fun little project to write, and quite useful for me: https://github.com/kstenerud/bonjson

masklinn3 days ago
> Schemas save you space, but then you lose the ability to understand the data without the schema. That's where JSON has always been handy despite its inefficiencies.

Of course that’s orthogonal to using a text v binary representation, you can have a schema’d text format, and a self-describing binary format.

BSON, UBJSON, MessagePack and CBOR are binary but self-describing.

deepsun3 days ago
Cool. And the next enlightening step would be to generate custom parser code from message schemas. So that when de-serializer reads a next field, it does not read what name and type it will be, it expects and fails if it's not. Safer and more compiler friendly.
Skwid3 days ago
One of my favourite yak shaving adventures in a previous job was writing a parse in place UBJSON decoder for ~1MB of data on a device with about as much free memory. Fast (enough) access by key, binary search with a few shortcuts for the mostly numeric payload. It built an index of maybe 30 bytes to speed things along, but also to let it keep working on the tail of the old message whilst the new one was overwriting it's start.

Should the system design have required all this of a device with 4MB of memory? Probably not. But it worked, and I had a great time

Read the full thread on Hacker News →

Related stories