Cloudflare's global network is immense but not limitless. As we look for small ways to trim our resource usage, we sometimes get lucky and we can cut significantly more. Here’s how we reduced one of our Pingora-based…

488 points•f311a•12 days ago•123 comments•

123 comments

zer0x4d12 days ago
Incredibly happy to see this series of CF articles. I was always so proud of devs back in the days where RAM and processing were scarce and who had to get creative to fit even the most basic stuff in the budget. It seemed to me that after RAM and processing became abundant, most gave up on optimization and focused on shipping instead which meant now that even with several cores, a basic notepad or music player failed to work. In a way, RAM becoming more expensive has ushered in a new era of forced optimizations, which I'm really happy for
jfengel12 days ago
I don't remember those days with a ton of fondness. Yes, the challenge was fun, but I really wanted to ship it and get my product in the hands of customers. Now I can spend more time thinking about what they want and less time about what the computer wants.
rstat112 days ago
And its this obsession with shipping things as fast as possible quality be dammed that got us basic weather apps that eat a gigabyte+ of RAM.
switchbak12 days ago
Yes, I remember those too. The costs of manual memory management were real and were not low.

But costs on the cloud are real too, especially now. I’ve been living in JVM land for a very long time, but now it’s especially clear how important lean services are. Especially now that the bar for writing lean code is so much lower: let the borrow checker figure it out, etc.

I just spent a couple days wringing out more performance/memory efficiency for our services. Nice gains to be sure, but it’s still so immensely wasteful compared to something well written running native. If it was my money, I’d be going native for sure.

1vuio0pswjnm711 days ago
"Now I can spend more time thinking about what they want and less time about what the computer wants."

What if they want memory efficiency

robotresearcher11 days ago
Customers often want to pay less, and shareholders often want lower capital costs. A tasteful optimization is a win-win. The opportunity cost should be traded off against new features of course, but tasteful optimization is a good thing for customers.
appreciatorBus12 days ago
Different people are different.

Some of us find production and optimization more interesting than marketing and distribution.

huijzer12 days ago
> In a way, RAM becoming more expensive has ushered in a new era of forced optimizations, which I'm really happy for

Isn't this mainly Cloudflare's scale though? That's literally also what's in the introduction written as the reason why they are doing the optimization

necovek12 days ago
While it's nice to see this focus, while they highlight absolute figures, we are still talking about 1% improvement. For most other software systems, nothing worth putting effort in for.
colechristensen11 days ago
During the Moore's Law years you didn't have to optimize much, if you did a significant release yearly computer hardware grew faster than your optimization problems.
dr_dshiv12 days ago
Cloudflare is truly amazing, they have made so much possible for my main side-project at a price and performance that I can’t really take credit for (http://sourcelibrary.org), I don’t care if their text was written with AI, I just wish I could get my own AI to sing so well about hashing… but wait.. today I noticed Claude trying to use hashing when a timestamp would honestly do, and now I’m really doubting myself, hmm…
ChoosesBarbecue12 days ago
> I don’t care if their text was written with AI, I just wish I could get my own AI to sing so well about hashing… but wait.. today I noticed Claude trying to use hashing when a timestamp would honestly do, and now I’m really doubting myself, hmm…

Tried out the first 1000 words in Pangram, and it seemed happy it was human written. Not surprised either, it has been some of the better writing I've seen out of Cloudflare recently.

terabyteoff12 days ago
An AI would have known that saying, “Hi, mom” in a professional post was a bad idea.
hiddencost11 days ago
Please stop using this stuff. It's snake oil.
davidbarker12 days ago
This is pleasant coincidence. Really like your site and it's queued to send in my newsletter in the morning! Just happened to see your comment here while I was reading. Great work.
dr_dshiv12 days ago
Oh super — I appreciate that!
mitxela12 days ago
What does Cloudflare make possible for your project?
dr_dshiv11 days ago
Super cheap storage and CDN for rapid image access — plus bot control.
antics911 days ago
That’s some great work! It’s interesting to read the LLMs translation of the yoga sutras
vlovich12312 days ago
I would get rid of consistent hashing and ketama for a better system which works save an additional 600TiB.

You use the first N bits of your key hash to pick the server partition so it’s a reasonable number (eg 128 servers per partition). Then use high quality precomputed hashes (first 64 bits of sha256) for the server name as N in H(K + N). Use wymum from wyhash as the H so that you do o(n) integer multiplications while retaining a result that’s still a good hash statistically.

Now you’re using a tournament hash, the small N means O(N) vs O(N log N) doesn’t matter, and also this O(N) is also going to be much less CPU than computing 160 hashes per key as they do now, so much less latency added per request.

MakersF12 days ago
I think they do only a hash per request. The 160*weight hashes were done per server (per feature set), to partition the hash space. Per request you do a single hash and then a lower_bound on a sorted map to find the serving server (again, on the ring appropriate for the features required by the request, so likely a hash map lookup first)
nopurpose11 days ago
Do I understand correctly, that they spent memory storing largish N hash values per server, so that request hash determines which server to send request to using closest higher value of all server hashes?

That in effect boils down to consistently selecting server S with probability P, where P is function of weight and total number of servers?

Surely there must be better way to select server with a given probability without storing a massive lookup table of hashes? Randevouz hashing of some sorts

QuaternionsBhop11 days ago
Plus it's limited to 65k entries. Perhaps a btree where parent nodes sum the weights of child nodes would work well. Using the input hash scaled by total weight, a binary search lookup would compute the partial sums for comparison on the fly. Adding/removing a node would only update the ~8 parents when the btree order is 4. Eytzinger layout and struct-of-arrays could be used to improve cache locality during lookup. This does mean an add/remove could drastically change the overall mapping, perhaps that's why consistent hashing is used instead.
varispeed11 days ago
Who cares if you can buy all the RAM available. To hell with small business and working class who now cannot afford it.
Fordec12 days ago
This sort of thing makes me thing that we're about to enter an era where software development is going to be where most of the jobs fallout will be. You can't one-shot vibe code your way to this. But for proper Software Engineering, those jobs are safe where more and more problems are going to actually need solving by creatively using math because all the problems individuals deliver are just going to be larger. People are just mourning the loss of the low hanging fruit.
killingtime7412 days ago
I think you're speaking like a software engineer, which is understandable, and not like a historian or economist. There's no reason to believe math based jobs would survive. The models regularly do well on math problems. You can auto-research loop ways to optimize memory usage for any particular program.
Fordec12 days ago
The point isn't that "doing math" is safe. Auto-research solves one target variable in one system, doing it at scale where say one developer is SME for the agentically manged 200 microservices down the line, heh I mean you certainly can, but good luck with that token cost of auto-research when that problem space is O(microservice^2). I point at that example yesterday of that optimized database memory with the comments pointing out that the specific problem fit in memory, over optimized and didn't generalize. The problem isn't the work, but the rework. A historian should know that new solutions to problems doesn't lead to "no problems ever again" but only problems with barriers that the new solution doesn't solve.

Read the full thread on Hacker News →

Related stories