179 points•farlight•11 days ago•190 comments•

190 comments

fenio11 days ago
The author of the benchmark here. I went over some comments and I'll try to tackle them here. I'm pretty clear that GH runner based benchmark is far from perfect due to noisy neighbours etc. Thus every test first is running so called calibration... to reject completely unreliable VMs. I'm fully aware that this can't completely fix the issue. Can limit it but not fix. But as of now there are 593 runs recorded so average should still be quite meaningful.

Having that said I'm desperately trying to get REAL hardware to run that benchmark. With some successes ;)

Few months ago I got Hetzner machine from Kent Overstreet and I was able to finish 3 runs before machine died... Results: https://bartosz.fenski.pl/modern-fs-benchmark/real-hw/

Currently I've got even more interesting machine with tons of disks and I'm running new set of benchmarks but it's really in its initial stage.

https://bartosz.fenski.pl/modern-fs-benchmark/sas-hdd/ 2nd run in progress... one run on REAL hardware takes much more time than on GH runner so it's slow.

But this new hardware has also so many disks that the plan is to try also more complex, tiered cache topologies. I'm working on it.

I'm happy to answer any other questions, sources of every piece of this benchmark are freely available and I'm not saying they are 100% correct. I'm open to improvements.

koverstreet11 days ago
I went back and forth with Hetzner a couple times, I think we just got a bad machine :)

I've been saying it for months, but eventually I'm going to move the automated builds off the 48 core monster and we'll be able to use that for automated perf testing too. The machine we just got has spindles for EC perf testing, but the Hetzner monster has very high end enterprise ssdd.

Also, just got done with the Rust for Linux conference, still not home but here's slides that still need reformatting: https://evilpiepirate.org/~kent/Kangrejos-2026-bcachefs.pdf

Joel_Mckay11 days ago
Should include:

1. Dual Ext4 + external 32GB journal X4 pcie SSD (the prior winner of benchmark surveys)

2. Bare F2FS after a trim and SSD vendor software cache flush operation (it should be slower, but knowing how much slower on identical hardware could be interesting.)

3. DRBD across a 48U 100Gbps host rack (single X4 pcie data drive per host, OS on primary)

4. CephFS across a 48U 100Gbps host rack (single X4 pcie data drive per host, OS on primary)

Best regards =3

throwaway27092511 days ago
Also

5. a ZFS dRaid configuration. There could be very different characteristics there with it using slabs.

Speaking of slabs, MS ReFS of you feel adventurous!

delamon11 days ago
6. modern nvme drive, preferrably pcie gen5 that can push >14GiB/sec
BoingBoomTschak10 days ago
Good job and clean website interface! Something as exhaustive as this is clearly needed.

Some remarks:

1. Why does the CoW button remove XFS from the list? Even https://github.com/fenio/modern-fs-benchmark/blob/main/scrip... mentions it has reflink enabled

2. If you have the time, adding XFS + mdraid + dm-integrity [1] (in bitmap mode) as a comparison point against ZFS RAID-Zx might be an interesting data point. That's what I run, personally.

3. Did you give some thoughts to the I/O scheduler choice? Might matter a lot in some cases.

[1] https://www.kernel.org/doc/html/latest/admin-guide/device-ma...

fenio10 days ago
Thanks, all three are fair points.

1. XFS reflink is enabled and its reflink/CoW-break measurements do run. The dashboard button currently means “native/full-CoW filesystem family”, not “supports reflink”, but that distinction is not clear from the label. I’ll rename it to “Native CoW” and add a separate reflink-capable filter that includes XFS.

2. The current integrity comparison is XFS on LVM/dm-raid10 with dm-integrity in its default journal mode. It is not mdraid and not bitmap mode, so your suggested stack would be a genuinely different and useful data point. An md RAID5/6 over per-member bitmap-mode dm-integrity comparison against RAID-Z1/Z2 makes sense, with the weaker post-crash bitmap semantics documented.

3. I did not pin or record the scheduler, which is a reproducibility gap. The dedicated SAS machine currently has mq-deadline active on all HDDs and SSDs. I’ll add queue/scheduler metadata to results before considering separate scheduler variants, since it can strongly affect the mixed and latency-sensitive phases.

Grayskull11 days ago
Love to hear that there will be more real hw tests. At the moment I am building NAS and used your benchmark for evaluating the filesystems. I am glad to see that your data roughly matches mine (apart from scrub which on 4x 6tb HDDs took 15 hours for md-raid10 while CoW systems took seconds). Personally I found that array of HDDs behaves very differently than GH runner (my feeling is that since it runs on same disk you are testing theoretical throughput rather than ability to utilize disks). My tests gave an idea for following topologies:

  * 4 HDDs (for example dm-raid has read balancing optimized specifically for HDDs)
  * 5 HDDs (classical raid should see no improvement but btrfs and bcachefs should balance the load)
  * 4 SSDs
  * 3 HDDs + 1 SSD no tiering
  * 2 HDDs + 2 SSD no tiering
  * 1 drive 10x larger than others (since how bcachefs and btrfs allocators work)
  * nocow
Thanks for awesome work
ciupicri11 days ago
How can a scrub take only seconds when it has to read all the data?
zenoprax11 days ago
I have three identical Lenovo SFF PCs with a U.2 SSD in each. I'm currently running them in a Ceph cluster but I'll be tearing that down soon. I could run some benchmarks with three in one box and report back? Would be a one-time thing rather than an on-going commitment though.

Disregard, I just saw you have some RAID10 tests in there so three SSDs won't be enough.

lproven11 days ago
Interesting although I'd have liked more summaries: there's an awful lot there.

But the reasons I choose filesystems are more about reliability, failure modes, surrounding tooling, and so on.

Btrfs fails in several critical areas:

1. No way to accurately find free space

2. catastrophic failure on write if a volume fills up, the probability of which is greater because of #1

3. repair tools usually do not recover a corrupted volume and in my testing are most likely to render as damaged volume completely unreadable, which makes #2 worse

Put these things together and I can never trust Btrfs again. In the 9 years since I encountered these, I see no effort to fix them, just fooling around witg unimportant side details like performance tweaks.

Fix the critical issues first then make it faster.

sandreas10 days ago
This. Btrfs blew up without ANY reason at all in my case. Rebooted and the system won't even recognize any filesystem. All btrfs tools fails to recover a single file.

Here is my journey: https://forum.cgsecurity.org/phpBB3/viewtopic.php?p=39143

I switched to ZFS and never Bad a Problem again.

FrinkleFrankle9 days ago
I've had a ZFS array running since ~2009. It's gone through upgrades and disk replacements, etc. Have not suffered a data loss in that time. ZFS I'd the goat.
burnt-resistor10 days ago
I used ZFS until it self-corrupted and refused to mount rw ever again. Community support was totally unhelpful and there was no resolution except buy another array and use something else. There's way too much ZFS cult fanboy glazing out there it doesn't deserve.
raegis10 days ago
Does the report say any of this? I only see "FAIL" on a few tests with ext4 and one with xfs.
koverstreet10 days ago
It's hard to show with any accuracy how likely a filesystem is to not break when the SHTF or something weird happens, or if they've handled all the weird corner cases, with any kind of automated test.

For that you have to dig into the methodology, look at the code, look at user reports, etc.

But you can get a pretty good approximation just from the philosophies and attitudes of the engineers and what they're talking about.

The talk I just gave at the Rust for Linux conference was all about that - how do we make the system debugable, the community aspect of how we respond to bug reports and talk to users, the prep work for the Rust conversion and formal verification and how we're approaching all that.

Reliability doesn't come out of nowhere, "all bugs are shallow with enough eyeballs" really doesn't apply to filesystems. You just have to plan for it, come up with a methodology, and do the work.

lproven10 days ago
It's not just me and it's had serious problems for years:

https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...

<- 5Y ago.

It's not materially better now. The devs are in denial about the problems because lots of big users are saying "works fine on my machine."

Sure, if you have lots of backups, if you have huge volumes on huge disks and they never fill up...

But it's the default in Fedora, Spiral Linux, Garuda Linux, siduction and others. Personal distros for people's own PCs and those are not well-supported enterprise kit.

KrOctave10 days ago
This is why I have fully switched to Bcachefs, which is just way more stable and performant than btrfs
bhaney11 days ago
Seeing great results from bcachefs just makes me more sad that Kent and the other kernel devs couldn't come to an understanding to keep bcachefs in-tree. I want to use it for my storage arrays so badly, but I'm still stuck with btrfs as the only available in-tree filesystem with modern features.
locknitpicker11 days ago
> Seeing great results from bcachefs just makes me more sad that Kent and the other kernel devs couldn't come to an understanding to keep bcachefs in-tree.

I think you are sugarcoating the shit show that was bcachefs's history of involvement in the linux kernel. I mean, do I need to mention that the person was subjected to a code of conduct enforcement action due to his long history of abuse and unprofessional behavior?

https://lwn.net/Articles/999197/

bhaney11 days ago
No, you probably didn't need to mention that.
RX1411 days ago
Personally I've had 0 issues with the bcachefs dkms packages from distro repos. Unlike zfs it keeps up with upstream kernel releases so it's far less hassle than running zfs dkms.
hnarn11 days ago
> Unlike zfs it keeps up with upstream kernel releases

Most normal users and especially servers have no reason to run latest upstream kernels

khajdamowicz11 days ago
You can use NASty as NAS appliance. It's based on NixOS, offers flexibility, atomic upgrades and all bells and whistles of bcachefs.
koverstreet11 days ago
NixOS. You can't go wrong.
Sha1rholder11 days ago
What do u mean NixOS? It's an OS not a file system
Farmadupe11 days ago
> CI runs use loop devices on shared ephemeral VMs (one VM per filesystem): compare shapes and ratios, not absolute MB/s. Each job records a host-calibration anchor — see the table.

I think if you're not using baremetal for such tests, it's likely that the results are simply not comparable at all? What if another tenant is also using the disk?

walrus0111 days ago
It's a fair point but it's also possible the person running the tests has a dedicated test hypervisor for this , so that different configurations of filesystems and VMs can be created and destroyed quickly in an automated manner.

If it's something as simple as a KVM hypervisor that only runs 1 test VM at a time (with no other load from anything else other than the basic systemd daemons, ssh daemon etc running on the hypervisor), the results could be very close to bare metal.

I can see it being very time consuming and annoying to do repeated manual bare metal OS installs and new partitioning/filesystem creation for such a large variety of tests.

The author does also say that performance isn't really the main thing but rather, data integrity:

https://github.com/fenio/modern-fs-benchmark

toast011 days ago
> I can see it being very time consuming and annoying to do repeated manual bare metal OS installs.

Well don't do that then. There's lots of other options. Probably the simplest is a single bare metal install on a simple filesystem on one device. run the filesystems under test on other storage dedicated to testing.

You could also boot into a network install and use local storage exclusively for testing.

Farmadupe11 days ago
> compare shapes and ratios, not absolute MB/s

In this case, given that the author's own disclaimer (above) already disclaims the numeric readings, I'm not sure how it's possible to make any inference on "shapes and ratios" derived from the numeric readings.

fenio11 days ago
For the main linked benchmark there are no dedicated test hypervisors. It's all based on GH runners with all the limitations and quirks that came with it.

Real hardware is used in: https://bartosz.fenski.pl/modern-fs-benchmark/real-hw/ https://bartosz.fenski.pl/modern-fs-benchmark/sas-hdd/

But unfortunatelly it's much more limited number of actual runs. sas-hdd is still in progress so numbers for it should increase over time.

sippingabonedry11 days ago
So two filesystems that are essentially shunned from the Linux kernel and permanent second-class citizens, and one that was removed from Red Hat and has a questionable history of reliability. Oh boy which do I choose?

I'm saying ZFS on another OS.

em-bee11 days ago
no thanks to red hat. i switched to debian for servers. fedora fortunately still supports btrfs.

zfs is shunned. can't really get around that copyright issue. bcachefs just hit a setback, which i am hopeful will eventually be resolved.

rdtsc11 days ago
> zfs is shunned. can't really get around that copyright issue

Well except for Ubuntu, one of the most popular Linux distros supporting it.

eqvinox10 days ago
> can't really get around that copyright issue.

Eh. Just wait until Oracle goes bankrupt due to their AI misinvestments & see where ZFS rights end up.

p_l11 days ago
The copyright issue is not really that big, except from some devs who decided to sure Ubuntu :V

But there's a LOT of FUD about it.

petre11 days ago
xfs/lvm-raid10 or ext4/lvm-raid10, obviously.
sippingabonedry11 days ago
If I'm forced to use Linux, sure.

I could drop bricks on and cord pull those all day and they would not lose data. Which is a small ask for a filesystem IMO.

andriy_koval11 days ago
would want to have compression..

Read the full thread on Hacker News →

Related stories