112 points•ibobev•5 days ago•166 comments•

166 comments

adrian_b4 days ago
I agree that the C flexible integer sizes were still necessary at the time of its creation, when some important computers still had word sizes that were not powers of two.

Nonetheless, I started to use C for programming only in 1990, when I got access to the Microsoft C and Borland Turbo C compilers.

At that time, 36 years ago, the C flexible integer sizes were already obsolete.

Since that time until now, while using C on a great variety of computers, from servers and workstations to the smallest microcontrollers, I have seen plenty of portability problems created by the existence of the flexible integer sizes.

The only programs that had no portability problems were those that never used the flexible integer sizes, but only integers with a definite size, e.g. 8-bit, 16-bit, 32-bit or 64-bit.

While sizeof solves the problems of memory allocation or copying, it does not help in preventing unexpected integer overflows, because even the size of "char" may be unknown, and even if the size of "char" is known, writing code with multiple paths that would check or prevent overflow for different integer sizes is very cumbersome.

Flexible integer sizes would work well only on the old computers, where integer overflow generated a hardware exception, so installing an overflow handler would have been sufficient to make the C code work correctly regardless of the size of the native integers.

nickcw4 days ago
The fixed size int types int32_t and friends weren't introduced until C99. Microsoft held out until 2013 before it put <inttypes.h> into Visual Studio!

So there has been a really long time in C's evolution where we haven't had fixed size types which has been a super annoying mess of #ifdefs in portable code.

The variable size ints have allowed some super weird architectures though. I remember looking at the datasheet for the Motorola 56000 DSP and noting that the C compiler set char = short = int = 24 bits! That was because the hardware could not address anything smaller than 24 bits. I think long could be 48 bits.

ThrowawayB73 days ago
The product is named Visual C++, that it happened to do C was always kind of a sideshow as far as I could tell. Aside from that, I seem to recall that VC++ had the WORD, DWORD, and eventually QWORD macros to specify unsigned 16, 32 and 64 bit types respectively.
convolvatron3 days ago
I always just defined these types in a per-platform portability header myself.
someonebaggy3 days ago
Microsoft had [unsigned] __int{8,16,32,64}
ddlsmurf4 days ago
I would agree if overflow on those types wasn't undefined behaviour, or unpredictable
bobmcnamara3 days ago
Plus char could be signed or unsigned!
jibal4 days ago
Per the C standard `sizeof(char)` is always 1, regardless of how many bits it has.
Maxatar2 days ago
You are mixing some things up. Yes the `sizeof(char)` is defined to always be 1 by the standard... but the size of char is not 1, it's defined by CHAR_BIT which can in principle vary from platform to platform.
ddlsmurf4 days ago
because it's the unit of addressable memory that C is concerned by with sizeof, otherwise its values still depend on CHAR_BIT
dmitrygr3 days ago
What you seek is CHAR_BIT
locknitpicker4 days ago
> At that time, 36 years ago, the C flexible integer sizes were already obsolete.

This is a highly ignorant comment. You're confusing the fact that you only had to work with a single target architecture with the whole concept of multiple processor architectures being somehow obsolete, as if there was a sudden law of nature that forced every single computer, being full blown HPC stuff or small microcontrollers used in embedded applications.

Take a look at arduino. They still have 16-bit models out there. Also noteworthy, it seems some DSPs also have ints larger than 32 bits.

flohofwoe4 days ago
The parent is completely right in the sense that for actually portable C code it was always better to use fixed-width integer types which were chosen for the problem to solve instead of target hardware capabilities.

For instance if your integer arithmetic needs to happen with 32 bit precision (no matter if the code runs on a 16- or 32-bit CPU), there is no scenario where using 'int' makes sense. Instead you'd use a fixed-width 32-bit integer type and accept that math operations are compiled into two instructions on a 16-bit CPU.

And OTH if you only require 16 bits integer width, there's not much point in picking a 32 bit integer type. Since two's-complement integer encoding has been standard since at least the 70s, the CPU can do narrow operations in the native register width. Any overflow/wraparound is still correct when only looking at the lowest 16-bits of the result.

adrian_b4 days ago
As I have said, I have not worked with a single architecture.

Before 1990, I had worked with a variety of ISAs, from IBM mainframes and DEC minicomputers to many kinds of microprocessors.

After 1990, I have used C on a great variety of x86, Motorola 68xxx, IBM/Motorola PowerPC and many generations of ARM ISAs.

Even if you use explicit 32-bit integers in a program, that will not create any correctness problem when the program is run on 16-bit microcontroller. At most such a program may have a suboptimal performance. Performance problems are much easier solved during porting than obscure bugs.

There have been some popular DSPs with 24-bit integers, e.g. Motorola 56xxx. Nonetheless, nobody would want to run on such a DSP a program that was written for another kind of CPU, even for another kind of DSP, because the performance would be pathetic. Any program for such a fixed-point DSP, even when derived from an existing program, would need to be rewritten while using at every point in the program the knowledge that the size of "int" is 24 bits (because the programs for fixed-point DSPs need copious amounts of scaling operations, to avoid overflows and underflows), so such a program should not actually use "int", but it should typedef an "int24_t", to make this assumption explicit.

layer83 days ago
> Language types such as char, int, short, and long do not come with a guarantee of how many bytes they occupy in memory.

Char is actually guaranteed by C to occupy exactly 1 byte in memory. It’s just that a byte can have more than eight bits in C. “Byte” is simply the smallest unit of memory addressable by a pointer.

Further down the article acknowledges that “C requires char to have at least 8 bits (CHAR_BIT >= 8), not exactly 8” and mentions the Honeywell 6000 as an example of a C implementation with 9 bits (and 36-bit ints).

Historically in computing, the size of a byte was hardware-dependent and not standardized. The Wikipedia article on “byte” cites Knuth’s 1968 TAOCP where byte denotes a unit which “contains an unspecified amount of information […] capable of holding at least 64 distinct values […] at most 100 distinct values. On a binary computer a byte must therefore be composed of six bits”.

kvemkon3 days ago
> C requires char to have at least 8 bits (CHAR_BIT >= 8)

The famous TI C55x DSP with 16 bits byte for example: https://news.ycombinator.com/item?id=3112704.

And they are still available new for purchase.

habitue3 days ago
It was intentional, sure. It was an attempt to solve a particular kind of problem.

In hindsight though, it was a mistake.

Evidence: when the world moved to 64 bit, we didnt just let int mean 8 bytes on amd64. That's a clear acknowledgement that the design was not correct once we understood things better.

sparkie3 days ago
`int` being 32-bits on amd64 was the correct decision, even in hindsight. If you are familiar with the ISA you will understand this. Existing 32-bit code just worked on the 64-bit chip, because the instruction encodings for 32-bit are unchanged. The 64-bit instructions are basically "opt-in", by placing a REX prefix on them - that makes them more expensive to encode and uses more instruction cache. Even today, compilers will emit 32-bit versions of instructions when the upper 32-bits are not needed, because it's cheaper.

When amd64 was released, x86 was almost ubiquitous on desktops and ran the majority of servers - most of the software used by the world could continue being used. If AMD had not gone through this effort to make it backward compatible, it's likely IA64 would've won and we wouldn't have this debate. Hard to understate the importance of not breaking things.

If you were designing a greenfield 64-bit ISA, then yes, it might make sense to have `int` be 64-bits, but it was definitely not, and still is not a mistake that it's 32-bits on amd64.

On RISC-V for example, it's questionable. The RV32 ecosystem is tiny and almost irrelevant - if they decided to break things for RV64 it wouldn't be a big problem - probably better to fix any problems early rather than hold baggage to run software that never existed - though it's much easier to port software to RV64 if `int` is still 32-bits.

Veserv3 days ago
You are agreeing with them.

They called it 'int' instead of 'int32_t' because it was meant to be the natural word size. It was supposed to float as the natural word size increased and allow code that worked on one word size to transparently work on the new word size.

Not floating to 64-bits on amd64 means their rationale for not giving it a fixed size was wrong. That is a simple width-doubling to another power of 2, the easiest possible case to expect to work transparently, on a compatible instruction set and they still chose to not float the size.

You can not just transparently change the word size and expect it to work, so you should just fix the size or use a fixed size. The notion of only having floating sizes as primitives was a mistake.

InvisibleUp3 days ago
Where the flexible integer sizes break the most is when dealing with ABIs, which weren't really a concern before dynamic linking existed but are very much a concern today. We've also, for some reason, decided that the standard way of defining a library ABI is with a C header. That means that everyone has to worry about precisely defining integer sizes, as well as more esoteric types like size_t or intmax_t. Good writeup on all that here: https://thephd.dev/to-save-c-we-must-save-abi-fixing-c-funct...
codedokode3 days ago
I think it didn't work out well, because "int" being different size makes programming difficult. For example, a system must manage up to 100 000 records. Can I use int for record number? What if it is 16 bits? What if I need to send data between machines, how can I use "int" if it can be different size?

Probably someone noticed that it is inconvenient, and on 64-bit machine ints are still 32-bit and not 64.

The computers with 16-bit ints or 9-bit bytes are long gone, but the language still has to carry that legacy.

Read the full thread on Hacker News →

Related stories