Samsung Electronics is set to more than double HBM4 and HBM4E output next year, lifting glass carrier demand 2.5-fold, industry sources said.

561 points•giuliomagnifico•10 days ago•459 comments•

459 comments

ricardobayes10 days ago
The article features a truck that the chips get loaded onto. This quote from Tanenbaum immediately jumped into mind:

"Never underestimate the bandwidth of a station wagon full of tapes hurtling down the highway."

leoedin10 days ago
I was wondering how much financial value they put in one truck. Presumably the value density of this stuff is incredibly high (although not really to anyone else, so security probably isn't a huge issue). At what point do you split the shipment between multiple trucks to avoid huge losses if there's a crash?

When I worked on satellites we had some electronics which used a radiation hardened FPGA which cost something like $500k per chip. Carrying that around was nerve wracking.

briffle9 days ago
I had a cousin that worked private security, and for a few years (I think it was early 2000's) his main gig was to work as a team of armed security, and follow trucks from the two Intel campuses in Oregon to the airport in unmarked vehicles, get on the cargo plane with the cargo, and fly to the location where they did the packaging of the wafers (I want to say it was the Philippine Islands) and then fly back. Lots of overtime.
ceejayoz10 days ago
I once had to keep a $20M (technically; a million download codes for the same $20 movie for a promotion) thumbdrive in my home safe because it arrived too late to put in the safety deposit box like it was supposed to be.

I had to remind myself a few times the practical value to anyone was about $22.

roflchoppa9 days ago
My buddies company shuffles chips around the west coast, some deliveries are a few million dollars. They shove the chips into the truck, then shove a giant movable wall, leaving about an inch or two gap around the perimeter.

The truck has multiple drivers and private security.

kotaKat9 days ago
It's treated the same way as any other generic shipment of cargo, really. It's fine throwing $20 million in one truck.

Security is an issue, though, these are treated as a high value load and typically would be followed by unmarked, plain clothes armed security guards tailing around the truck. (That's how Apple often does warehouse deliveries of trucks full of iPhones. It wouldn't surprise me if there was going to be some rolling lead soon following some trucks full of Duos...)

tencentshill9 days ago
They are getting hijacked/stolen and disappearing.

https://www.wired.com/story/the-worst-ive-ever-seen-cargo-th...

Likely ending up on a container to China, export bans be damned.

sebastiansm710 days ago
Almost 10 years ago at my university in Chile. My professor of Computer Networks invited a PhD Student that was working on the ALMA Observatory [0] doing some research on signal processing.

What stayed with me from that talk was that all the data that was captured from those massive antennas was packaged in boxes, loaded in a cargo van and travel 1.600 kilometers to Santiago to be uploaded into the systems.

[0] https://www.almaobservatory.org

KellyCriterion9 days ago
Amazon did this initially as well for large data pools when launching AWS S3 etc.: Sending you a truck/container to put all the disks in:

https://www.supplychain247.com/article/how_do_you_get_all_of...

rpozarickij9 days ago
I wanted to mention AWS Snowmobile [0], but after looking it up I found that apparently it was discontinued [1]. It seems that nowadays there are more economical/efficient alternatives (I guess theoretically you could still beat these alternatives in certain metrics, e.g. by using more trucks at once to get more bandwidth).

[0] https://aws.amazon.com/blogs/aws/aws-snowmobile-move-exabyte...

[1] https://www.cnbc.com/2024/04/17/aws-stops-selling-snowmobile...

lilbigdoot9 days ago
I forget how much it is exactly now but I saw a short video on the bandwidth of moving modern cold storage down the highway and its actually nuts
hacker_homie10 days ago
Just hope nothing changes the latency is terrible.
y1n010 days ago
I’ve often wondered about die thinning. I mean just looking at sd cards and micro sd etc, it was clear that there had to be a thinning step (or many steps) and it seemed amazing to me that such a complex step could be economically viable. And yet it clearly is. All a matter of being able to operate at scale I guess.

It’s not a step I’ve seen discussed very much, so it was kind of neat to see it discussed front and center in a lay article like this.

pjc5010 days ago
Planarization has always been part of the process: https://jeez-semicon.com/blog/Planarization-in-Semiconductor...

https://resources.pcb.cadence.com/blog/2023-the-planarizatio...

Not dissimilar to lens or mirror making, the process of grinding to a precise flatness. The key step here is bonding it to a glass carrier, presumably because otherwise it gets too thin to support itself.

See also the "reduced kerf diamond wire saw": technology critical to cheap solar panels is the ability to saw the boule into thinner and thinner pieces with less waste at each cut, like a big ham slicer.

KurSix10 days ago
Yeah, the scale is kind of mind-bending. You're taking something that's already extremely fragile, grinding most of its thickness away, handling it through several more process steps, and somehow doing that cheaply enough for mass-market products.
xadhominemx10 days ago
The critical back end step for HBM is thermal compression bonding - where they actually compress 13 very thin die separated by solder balls. Cracking is a major issue.
lazide10 days ago
Automation, automation, automation.

Zero chance you could do it if a human had to touch the die/chip directly during those steps.

But clean room + vacuum pick and place machine? A matter of fine tuning, and it should work just fine.

HarHarVeryFunny10 days ago
On a related note, I was reading yesterday that apparently the real bottleneck for Chinese production of AI accelerators is HBM production, not processors or ASML equipment.

The lack of ASML EUV machines certainly hurts, and pushing DUV so hard results in abysmal yields of good chips, but you can compensate by running more wafers or making smaller chips, and the net result is that Huawei's Ascend production volume is limited by CXMT's HBM capacity not processor dies.

The problem is that HBM manufacture requires many steps (die thinning, via drilling, plating, alignment) where the equipment used by everyone else (Samsung, SK Hynix, Micron) is also blocked by sanctions, so the Chinese are having to develop all of this themselves too, which they have, but yields are currently low, even when using shorter HBM stacks.

jacquesm10 days ago
The only thing the West is achieving here is that sooner or later China will be able to compete on their own terms rather than ours. It may buy some time but the end result is very predictable.
bigcat1234567810 days ago
As other comments pointed out, that was always China's plan, the value here is really that when the Chinese overlord annillate the world's industrial competition, they'll have the moral high ground of being out of self-reliance and hostility by western monopolists.

Edit:

And, that, my fellow keyboard high-minded thinker but powerless corp peons, was also part of the China's plan all along.

audunw10 days ago
What the West is doing is forcing China into a situation which is impossible to sustain when their economic dividend runs out. Especially considering these investments are fuelled by an extremely high level of debt.

And it’s not a given that they’ll catch up. I think it’s fair to say that their development of jet engines is on track to be obsolete before they’re competitive.

They may in principle have established the technologies to be self sustained but that’s only relevant and sustainable if they develop an economy based on domestic consumption rather than exports. Which they are struggling with as well.

stingraycharles10 days ago
It baffles me how shortsighted the policymaking here is. Like, what did they expect to happen?
deepsun10 days ago
Well Soviet Union tried to be self-sufficient, but the truth is such strategy loses to global free trade. Whether we want it or not, global trade is what keeps the world from WW3 for now, not the nuclear MAD.
KurSix10 days ago
Sanctions can absolutely create the incentive to build a domestic supply chain, but incentive doesn't automatically translate into catching up on the same timeline
andy_ppp10 days ago
Tokens per second is almost entirely memory bandwidth at inference time, training obviously needs more compute but you can add more chips for that.
martinald10 days ago
Not quite, it's got quite a bit more complicated with agentic use cases.

Prefill (input tokens) is heavily compute bound. And the ratio of input to output continues to rise, as typically in agentic sessions you have a few tokens output for a tool call and (many) thousands of input from the tool result.

Then you have cached input tokens, which is a totally different issue, system RAM or NVMe bound.

Obviously output tokens is VRAM memory bandwidth bound, but this is less and less of the bottleneck these days for overall agentic speed.

cubefox10 days ago
According to SemiAnalysis, both inference and post-training (RLVR) is mostly memory bandwidth bound. Only pre-training is compute bound, but it now only takes a small share of overall data center capacity.

https://x.com/EugeneNg/status/2099315982959616369

FooBarWidget10 days ago
The "abysmal" yield using DUV is an overblown statement by western commentators who don't look closer at the development. Certainly yield is worse than with EUV, but after multiple iterations of development, yield has become pretty good, within economically acceptable bounds, still making scaling possible. Volume is still ramping up. A lot of capacity will come online in 2027. The removal of western middlemen such as Mediatek actually improved the economics. Further yield improvements are still coming. Volume is already so high that domestically made phones and NPUs are making a real impact, yet are also selling like hot cakes.
HarHarVeryFunny10 days ago
It depends on what you are trying to make, but there are workarounds for most things, and it becomes more of a cost and volume issue than a showstopper.

If your yield is low and you still want the volume, then run more wafers, but then you need access to more DUV machines and more wafers. If your defect density is too high, then design smaller chips with a higher chance of avoiding defects, which is what Hauwei are doing - split processor into multiple chiplets (NVidia do this too, to raise yields and reduce cost).

HBM3E usually has 5-6000 vias, but doing this with multi-pattern DUV without defects is tough, so CXMT currently drop that to 3000, at a cost of some loss of thermal and voltage stability.

HBM:processor production ratios aren't what Huawei would like, so they mitigate this at system level by putting an optical memory bus on the GPU chip and sharing memory across the system.

Not everything is a huge LLM - smaller models like recommendation systems don't have so many parameters and need so much memory, so why waste HBM on them? ByteDance use their SeedChip acelerator for this, currently made by TSMC (using an older sanctions-approved 28mm process), which instead etches dense RRAM on die beside the processor.

There is also a time dynamic to this, with China still using pre-sanctions equipment and chips as their domestic alternatives ramp up to replace them. One interesting part of this is HBM ... HBM is very demanding to make, and everyone had problems with it, with initially only SK Hynix being successful. Samsung and Micron took a year or so to catch up, and during this time Samsung had made a ton of HBM that didn't meet NVidia's specifications, so ended up, pre-sanctions, selling it to China, where it has acted as a stockpile to carry them over as CXMT's domestic capacity ramps up.

China seems to be doing fine. Of course they would like sanctions lifted, moreso for HBM than anything else, but all that sanctions have really achieved is accelerating their semiconductor independence. They are still building 1T+ SOTA models, standing up 100K clusters of domestic AI accelerators, etc, etc.

vatsachak10 days ago
China should invest in an analog inference chip. It's a hail mary but why not.
danielheath10 days ago
The two challenges there are firstly - that analog design has been a separate electrical engineering school for most of a century, so there are few who could design it - and secondly - that every single chip will have subtle variations in its computations, necessitating some sort of model finetuning per chip. Possibly the chip could be characterised at the factory, and ship with the characterisation data burned into a controller rom or something, but if that doesn’t pan out the whole thing is likely a non-starter.

If it could be made to work, you could run a fable-grade model in tens of watts.

geysersam10 days ago
They can probably afford to do both.
bell-cot10 days ago
I'd assume that they have - but will keep mum 'till they have a major breakthrough or large-scale operational deployment to announce.
sroussey10 days ago
That would be like skipping land line phones for mobile...
bobmcnamara10 days ago
Or on DRAM on die compute
chvid10 days ago
Huawei uses their own non-standard HBM called HiZQ probably not produced by CXMT.
fooker10 days ago
What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?
bob102910 days ago
It's not that there's a blocker. It's that it takes roughly 3x the manufacturing capacity to produce an HBM package at the same storage capacity as DRAM. We are sacrificing total bytes for bandwidth.
rkagerer10 days ago
I don't fully understand the source of the "total bytes" constraint, but a major factor may be because HBM4 / HBM4E can only make use of the footprint directly above the processor/logic die (or in direct vicinity of its interconnect), while traditional DRAM can be placed further away where there's lots of real estate on the motherboard.

I gather a practical max ceiling today is a stack of 16 chips in height yielding 64GB?

These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5. That's what gives them their order-of-magnitude bandwidth speedup. But even though they technically pack in more capacity per square millimeter of motherboard, I gather they take up more space than older technologies once you account for the vias and interconnects to route all those signals.

tliltocatl10 days ago
How so? FEOL is pretty much the same, BEOL is almost the same save TSVs, the packaging tech is different and more advanced, but not exactly 1:1 comparable. Do TSVs really occupy 3x the area of DDR IO's?
phkahler10 days ago
>> What's the main blocker (other than the current inflated cost) for using HBM instead of DRAM as the primary memory for consumer electronics?

HBM is meant to be integrated into the same package as the CPU, so no more DIMM sockets. It also has higher latency apparently.

MarleTangible10 days ago
One of the comments mentioned that they have a bus size of 2048, which may be why the latency is higher.

> These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5.

HarHarVeryFunny10 days ago
Why would you want/need to?

The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU.

So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second!

These are not normal requirements, other than for a GPU.

Dylan1680710 days ago
> Compare that to the 128 bit wide bus of a modern CPU.

Or 256-512 bits on medium to high end consumer CPUs if you're apple.

At least DDR6 is probably widening things 50%.

XorNot10 days ago
Right but if HBM memory is all that people want to produce, then building a CPU which can use it use it would be useful on it's own merits.

But in reality we also already have unified memory architecture systems, integrated graphics etc.

bobmcnamara10 days ago
Intel and IBM have already done this almost this with their wide eDRAM caches.

The advantage in going wide is transferring cache lines rapidly, not the CPU bus interface.

fooker10 days ago
SIMD (and especially the modern matrix extensions) can use as much bandwidth you can throw at it.
chessgecko10 days ago
I think people might prefer the lower idle power consumption from lpddr over the better bandwidth in hbm in battery powered stuff. That said right now the price is definitely preventing us from finding out.
fooker10 days ago
Idle yes, but HBM energy consumption / memory operations seems to be a bit better than DRAM.
torginus10 days ago
Mainly bus width. Afaik HBM is like 1024 bits vs DDRs 64 so you need lots of transfers in parallel to saturate the bus, and CPUs kinda want 64 bytes of data as that's the size of a cache line ASAP. So you need a ton of in flight transfers which isn't a thing CPUs provide, maybe multicore workloads.

Buy the way you win with CPUs is with latency, and not bandwidth, which is why Apple M series actually uses DDR with lower latency because of the stacking.

gs1710 days ago
A shame that this should if anything, lead to consumer DRAM prices getting even worse.
nicoburns10 days ago
Why would more RAM supply lead to higher consumer prices?
devy10 days ago
HBM4 and HBM4E DRAM are NOT the DDR4/5 that consumer markets need. Capacity allocation is leaning more to data center grade HBMs so less to produce dedicated DDR4/5 DRAMs. Supply demand will further drive up the consumer DRAM price! Note, the article mentions NO of new fabs is being constructed (all semiconductor manufacturers know that constructing more fabs means the boom/burst cycle will eventually kill them, so no one create more fabs) Perhaps the federal government need to step in here - the market doesn't fit the issue.
matja10 days ago
Depends on what proportion of Samsung's output is current HBM4 and HBM4E DRAM.

Everything in their statement can be true and it be a bad thing for consumers of non-HBM RAM.

"HBM capacity to expand to 250,000 wafers a month"

So let's say their current HBM capacity is 100k wafers/month (pure speculation/random number for illustration), and their total RAM capacity (including HBM( is 300k wafers/month, then non-HBM capacity reduces from 200k to 50k.

mattstir10 days ago
Part of the implication is that factories that could be producing consumer-facing DRAM like DDR5 would be retooled to produce HBM instead, leading to even less total consumer RAM production.
zargon10 days ago
Memory allocated for HBM is memory taken away from DDR production.
kenny1110 days ago
If Samsung is limited in the number of wafers they can process per month and they use more of those to produce HBM they necessarily have less of them left to make other products, like consumer DRAM.
robotnikman10 days ago
Maybe this leads to device manufacturers using HBM memory instead.

Read the full thread on Hacker News →

Related stories