190 comments
Having that said I'm desperately trying to get REAL hardware to run that benchmark. With some successes ;)
Few months ago I got Hetzner machine from Kent Overstreet and I was able to finish 3 runs before machine died... Results: https://bartosz.fenski.pl/modern-fs-benchmark/real-hw/
Currently I've got even more interesting machine with tons of disks and I'm running new set of benchmarks but it's really in its initial stage.
https://bartosz.fenski.pl/modern-fs-benchmark/sas-hdd/ 2nd run in progress... one run on REAL hardware takes much more time than on GH runner so it's slow.
But this new hardware has also so many disks that the plan is to try also more complex, tiered cache topologies. I'm working on it.
I'm happy to answer any other questions, sources of every piece of this benchmark are freely available and I'm not saying they are 100% correct. I'm open to improvements.
I've been saying it for months, but eventually I'm going to move the automated builds off the 48 core monster and we'll be able to use that for automated perf testing too. The machine we just got has spindles for EC perf testing, but the Hetzner monster has very high end enterprise ssdd.
Also, just got done with the Rust for Linux conference, still not home but here's slides that still need reformatting: https://evilpiepirate.org/~kent/Kangrejos-2026-bcachefs.pdf
1. Dual Ext4 + external 32GB journal X4 pcie SSD (the prior winner of benchmark surveys)
2. Bare F2FS after a trim and SSD vendor software cache flush operation (it should be slower, but knowing how much slower on identical hardware could be interesting.)
3. DRBD across a 48U 100Gbps host rack (single X4 pcie data drive per host, OS on primary)
4. CephFS across a 48U 100Gbps host rack (single X4 pcie data drive per host, OS on primary)
Best regards =3
5. a ZFS dRaid configuration. There could be very different characteristics there with it using slabs.
Speaking of slabs, MS ReFS of you feel adventurous!
Some remarks:
1. Why does the CoW button remove XFS from the list? Even https://github.com/fenio/modern-fs-benchmark/blob/main/scrip... mentions it has reflink enabled
2. If you have the time, adding XFS + mdraid + dm-integrity [1] (in bitmap mode) as a comparison point against ZFS RAID-Zx might be an interesting data point. That's what I run, personally.
3. Did you give some thoughts to the I/O scheduler choice? Might matter a lot in some cases.
[1] https://www.kernel.org/doc/html/latest/admin-guide/device-ma...
1. XFS reflink is enabled and its reflink/CoW-break measurements do run. The dashboard button currently means “native/full-CoW filesystem family”, not “supports reflink”, but that distinction is not clear from the label. I’ll rename it to “Native CoW” and add a separate reflink-capable filter that includes XFS.
2. The current integrity comparison is XFS on LVM/dm-raid10 with dm-integrity in its default journal mode. It is not mdraid and not bitmap mode, so your suggested stack would be a genuinely different and useful data point. An md RAID5/6 over per-member bitmap-mode dm-integrity comparison against RAID-Z1/Z2 makes sense, with the weaker post-crash bitmap semantics documented.
3. I did not pin or record the scheduler, which is a reproducibility gap. The dedicated SAS machine currently has mq-deadline active on all HDDs and SSDs. I’ll add queue/scheduler metadata to results before considering separate scheduler variants, since it can strongly affect the mixed and latency-sensitive phases.
* 4 HDDs (for example dm-raid has read balancing optimized specifically for HDDs)
* 5 HDDs (classical raid should see no improvement but btrfs and bcachefs should balance the load)
* 4 SSDs
* 3 HDDs + 1 SSD no tiering
* 2 HDDs + 2 SSD no tiering
* 1 drive 10x larger than others (since how bcachefs and btrfs allocators work)
* nocow
Thanks for awesome workDisregard, I just saw you have some RAID10 tests in there so three SSDs won't be enough.
But the reasons I choose filesystems are more about reliability, failure modes, surrounding tooling, and so on.
Btrfs fails in several critical areas:
1. No way to accurately find free space
2. catastrophic failure on write if a volume fills up, the probability of which is greater because of #1
3. repair tools usually do not recover a corrupted volume and in my testing are most likely to render as damaged volume completely unreadable, which makes #2 worse
Put these things together and I can never trust Btrfs again. In the 9 years since I encountered these, I see no effort to fix them, just fooling around witg unimportant side details like performance tweaks.
Fix the critical issues first then make it faster.
Here is my journey: https://forum.cgsecurity.org/phpBB3/viewtopic.php?p=39143
I switched to ZFS and never Bad a Problem again.
For that you have to dig into the methodology, look at the code, look at user reports, etc.
But you can get a pretty good approximation just from the philosophies and attitudes of the engineers and what they're talking about.
The talk I just gave at the Rust for Linux conference was all about that - how do we make the system debugable, the community aspect of how we respond to bug reports and talk to users, the prep work for the Rust conversion and formal verification and how we're approaching all that.
Reliability doesn't come out of nowhere, "all bugs are shallow with enough eyeballs" really doesn't apply to filesystems. You just have to plan for it, come up with a methodology, and do the work.
https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...
<- 5Y ago.
It's not materially better now. The devs are in denial about the problems because lots of big users are saying "works fine on my machine."
Sure, if you have lots of backups, if you have huge volumes on huge disks and they never fill up...
But it's the default in Fedora, Spiral Linux, Garuda Linux, siduction and others. Personal distros for people's own PCs and those are not well-supported enterprise kit.
I think you are sugarcoating the shit show that was bcachefs's history of involvement in the linux kernel. I mean, do I need to mention that the person was subjected to a code of conduct enforcement action due to his long history of abuse and unprofessional behavior?
Most normal users and especially servers have no reason to run latest upstream kernels
I think if you're not using baremetal for such tests, it's likely that the results are simply not comparable at all? What if another tenant is also using the disk?
If it's something as simple as a KVM hypervisor that only runs 1 test VM at a time (with no other load from anything else other than the basic systemd daemons, ssh daemon etc running on the hypervisor), the results could be very close to bare metal.
I can see it being very time consuming and annoying to do repeated manual bare metal OS installs and new partitioning/filesystem creation for such a large variety of tests.
The author does also say that performance isn't really the main thing but rather, data integrity:
Well don't do that then. There's lots of other options. Probably the simplest is a single bare metal install on a simple filesystem on one device. run the filesystems under test on other storage dedicated to testing.
You could also boot into a network install and use local storage exclusively for testing.
In this case, given that the author's own disclaimer (above) already disclaims the numeric readings, I'm not sure how it's possible to make any inference on "shapes and ratios" derived from the numeric readings.
Real hardware is used in: https://bartosz.fenski.pl/modern-fs-benchmark/real-hw/ https://bartosz.fenski.pl/modern-fs-benchmark/sas-hdd/
But unfortunatelly it's much more limited number of actual runs. sas-hdd is still in progress so numbers for it should increase over time.
I'm saying ZFS on another OS.
zfs is shunned. can't really get around that copyright issue. bcachefs just hit a setback, which i am hopeful will eventually be resolved.
Well except for Ubuntu, one of the most popular Linux distros supporting it.
Eh. Just wait until Oracle goes bankrupt due to their AI misinvestments & see where ZFS rights end up.
But there's a LOT of FUD about it.
I could drop bricks on and cord pull those all day and they would not lose data. Which is a small ask for a filesystem IMO.
Read the full thread on Hacker News →
Related stories
- Hacker News · 2 points · 4 days ago
- Hacker News · 2 points · 3 days ago
- Hacker News · 1 points · 4 days ago
- Lobsters · 33 points · about 1 year ago
- Hacker News · 1 points · 9 days ago
- Leaked Gemini 4 Pro Benchmarks Signal Google's Return to the Top of AInokiapoweruser.comHacker News · 2 points · 1 day ago