Samsung Electronics is set to more than double HBM4 and HBM4E output next year, lifting glass carrier demand 2.5-fold, industry sources said.
459 comments
"Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway."
When I worked on satellites we had some electronics which used a radiation hardened FPGA which cost something like $500k per chip. Carrying that around was nerve wracking.
I had to remind myself a few times the practical value to anyone was about $22.
The truck has multiple drivers and private security.
Security is an issue, though, these are treated as a high value load and typically would be followed by unmarked, plain clothes armed security guards tailing around the truck. (That's how Apple often does warehouse deliveries of trucks full of iPhones. It wouldn't surprise me if there was going to be some rolling lead soon following some trucks full of Duos...)
https://www.wired.com/story/the-worst-ive-ever-seen-cargo-th...
Likely ending up on a container to China, export bans be damned.
What stayed with me from that talk was that all the data that was captured from those massive antennas was packaged in boxes, loaded in a cargo van and travel 1.600 kilometers to Santiago to be uploaded into the systems.
https://www.supplychain247.com/article/how_do_you_get_all_of...
[0] https://aws.amazon.com/blogs/aws/aws-snowmobile-move-exabyte...
[1] https://www.cnbc.com/2024/04/17/aws-stops-selling-snowmobile...
It’s not a step I’ve seen discussed very much, so it was kind of neat to see it discussed front and center in a lay article like this.
https://resources.pcb.cadence.com/blog/2023-the-planarizatio...
Not dissimilar to lens or mirror making, the process of grinding to a precise flatness. The key step here is bonding it to a glass carrier, presumably because otherwise it gets too thin to support itself.
See also the "reduced kerf diamond wire saw": technology critical to cheap solar panels is the ability to saw the boule into thinner and thinner pieces with less waste at each cut, like a big ham slicer.
Zero chance you could do it if a human had to touch the die/chip directly during those steps.
But clean room + vacuum pick and place machine? A matter of fine tuning, and it should work just fine.
The lack of ASML EUV machines certainly hurts, and pushing DUV so hard results in abysmal yields of good chips, but you can compensate by running more wafers or making smaller chips, and the net result is that Huawei's Ascend production volume is limited by CXMT's HBM capacity not processor dies.
The problem is that HBM manufacture requires many steps (die thinning, via drilling, plating, alignment) where the equipment used by everyone else (Samsung, SK Hynix, Micron) is also blocked by sanctions, so the Chinese are having to develop all of this themselves too, which they have, but yields are currently low, even when using shorter HBM stacks.
Edit:
And, that, my fellow keyboard high-minded thinker but powerless corp peons, was also part of the China's plan all along.
And it’s not a given that they’ll catch up. I think it’s fair to say that their development of jet engines is on track to be obsolete before they’re competitive.
They may in principle have established the technologies to be self sustained but that’s only relevant and sustainable if they develop an economy based on domestic consumption rather than exports. Which they are struggling with as well.
Prefill (input tokens) is heavily compute bound. And the ratio of input to output continues to rise, as typically in agentic sessions you have a few tokens output for a tool call and (many) thousands of input from the tool result.
Then you have cached input tokens, which is a totally different issue, system RAM or NVMe bound.
Obviously output tokens is VRAM memory bandwidth bound, but this is less and less of the bottleneck these days for overall agentic speed.
If your yield is low and you still want the volume, then run more wafers, but then you need access to more DUV machines and more wafers. If your defect density is too high, then design smaller chips with a higher chance of avoiding defects, which is what Hauwei are doing - split processor into multiple chiplets (NVidia do this too, to raise yields and reduce cost).
HBM3E usually has 5-6000 vias, but doing this with multi-pattern DUV without defects is tough, so CXMT currently drop that to 3000, at a cost of some loss of thermal and voltage stability.
HBM:processor production ratios aren't what Huawei would like, so they mitigate this at system level by putting an optical memory bus on the GPU chip and sharing memory across the system.
Not everything is a huge LLM - smaller models like recommendation systems don't have so many parameters and need so much memory, so why waste HBM on them? ByteDance use their SeedChip acelerator for this, currently made by TSMC (using an older sanctions-approved 28mm process), which instead etches dense RRAM on die beside the processor.
There is also a time dynamic to this, with China still using pre-sanctions equipment and chips as their domestic alternatives ramp up to replace them. One interesting part of this is HBM ... HBM is very demanding to make, and everyone had problems with it, with initially only SK Hynix being successful. Samsung and Micron took a year or so to catch up, and during this time Samsung had made a ton of HBM that didn't meet NVidia's specifications, so ended up, pre-sanctions, selling it to China, where it has acted as a stockpile to carry them over as CXMT's domestic capacity ramps up.
China seems to be doing fine. Of course they would like sanctions lifted, moreso for HBM than anything else, but all that sanctions have really achieved is accelerating their semiconductor independence. They are still building 1T+ SOTA models, standing up 100K clusters of domestic AI accelerators, etc, etc.
If it could be made to work, you could run a fable-grade model in tens of watts.
I gather a practical max ceiling today is a stack of 16 chips in height yielding 64GB?
These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5. That's what gives them their order-of-magnitude bandwidth speedup. But even though they technically pack in more capacity per square millimeter of motherboard, I gather they take up more space than older technologies once you account for the vias and interconnects to route all those signals.
HBM is meant to be integrated into the same package as the CPU, so no more DIMM sockets. It also has higher latency apparently.
> These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5.
The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU.
So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second!
These are not normal requirements, other than for a GPU.
Or 256-512 bits on medium to high end consumer CPUs if you're apple.
At least DDR6 is probably widening things 50%.
But in reality we also already have unified memory architecture systems, integrated graphics etc.
The advantage in going wide is transferring cache lines rapidly, not the CPU bus interface.
Buy the way you win with CPUs is with latency, and not bandwidth, which is why Apple M series actually uses DDR with lower latency because of the stacking.
Everything in their statement can be true and it be a bad thing for consumers of non-HBM RAM.
"HBM capacity to expand to 250,000 wafers a month"
So let's say their current HBM capacity is 100k wafers/month (pure speculation/random number for illustration), and their total RAM capacity (including HBM( is 300k wafers/month, then non-HBM capacity reduces from 200k to 50k.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · about 14 hours ago
- Ars Technica · 0 points · 7 days ago
- Micron to Double General-Purpose DRAM Capacity per Serverinvestors.micron.comHacker News · 1 points · 1 day ago
- Hacker News · 2 points · 9 days ago
- Hacker News · 2 points · 9 days ago
- Hacker News · 2 points · 8 days ago