Git 3.0 will make SHA-256 the new default content hashing algorithm and it will be an incomprehensibly expensive and ultimately valueless and avoidable global nightmare.

123 points•chmaynard•about 3 hours ago•137 comments•

137 comments

kpcyrdabout 2 hours ago
This article is full of mistakes and misleading claims:

1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.

2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.

3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.

schaconabout 2 hours ago
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.

2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.

3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.

https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...

bawolffabout 1 hour ago
> 1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.

It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.

In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?

kazinatorabout 2 hours ago
The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.

Nothing else matters.

Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.

zygentoma13 minutes ago
Sorry, no.

When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.

Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.

shakowabout 1 hour ago
> Git hashes are not supposed to be a security mechanism

Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?

meinersburabout 2 hours ago
Linus Torvalds in 2007:

> but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, Okay, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get. ... [1]

[1] https://www.youtube.com/watch?v=4XpnKHJAok8&t=56m20s

So Torvalds used SHA-1 purely because he needed a hash function with no other property than identifying content.

zamalekabout 2 hours ago
Exactly. But that's why I think SHA was a mistake. He should have gone with something like murmur to avoid all this frothing at the mouth.
layer811 minutes ago
It was a mistake to assume a fixed algorithm in the repository format and client-server protocol. I remember being surprised when I learned about that choice, having myself been familiar with cryptographic protocols and formats where the hash algorithm is usually a parameter that can vary for each concrete hash.
Someone29 minutes ago
Linus, in 2005, couldn’t have gone for murmur, from 2008.

Was there “something like murmur” in 2005 that’s cryptographically better than SHA1?

bawolffabout 1 hour ago
If that was true, i doubt git would have switched to using the slower version of sha-1 that detects attacks.
kazinatorabout 2 hours ago
Frothers gonna froth, though.
gandreaniabout 2 hours ago
One of my favorite fun facts about Fossil SCM (another source control by the devs of sqlite) is that they patched their use of SHA1 6 days after the shattered attack was published:

"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."

https://fossil-scm.org/home/doc/trunk/www/hundredandone.md

To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.

That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!

6thbitabout 2 hours ago
That's impressive. I suppose they had a more flexible architecture to make that change so fast.

Is there any writeup on why it was easy for them and not for git?

toyminabout 2 hours ago
My guess is that it's less about the architecture and more about the blast radius and the number of users
gandreaniabout 2 hours ago
Hmm probably nothing architecture wise. It's probably just the fact that fossil is developed by way fewer devs.

From the skim I read of this article it seems both projects arrived at the same solution: support both but make SHA-256 the default.

schaconabout 2 hours ago
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.

Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.

schaconabout 2 hours ago
Also, interestingly, Git today does _not_ use a straight SHA1 because of these attacks. It uses `sha1dc`, a slower collision detecting variant that specifically checks for this vector of attacks. So currently, Git's SHA-1 variant is not susceptible to the SHAttered/Shambles attacks.
gandreaniabout 2 hours ago
Agreed! This isn't a tech dig at all.

To me it's more of a reality of creating a tool with a huge active community and a community of contributors and creating a tool with a small team and small community.

fragmedeabout 2 hours ago
It's easier to make world breaking changes when the world is really small. If git could magically just get everything and everyone to cut over and use git 3.0 in a magic instant, it wouldn't be having this problem.
nofunsirabout 1 hour ago
Here's a stoichiometric bird for you:

:%s/git 3\.0/python 3.0/g

amlutoabout 3 hours ago
I don't understand why Git is not making the SHA-1 and SHA-256 modes far more compatible with each other.

SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.

But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.

And now it would be impossible to get a new SHA1 collision in to the repo.

The only new git features needed would be:

a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object

b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit

c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit

schaconabout 3 hours ago
Emily's talk does a pretty good job of summarizing the issues with intermixing the hashes: https://youtu.be/eJJp0RE7cd4
RJIb8RBYxzAMX9uabout 2 hours ago
I skimmed the video, and I didn't quite catch that. Near the end of the video, however, she did mention that interop is in the works[0].

In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.

Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.

[0] https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s

mort96about 2 hours ago
Hm but the date is stored inside of the commit. The only way we can know that a commit's date is authentic... is through its hash. If I can forge commits with any SHA1 hash at will, I can make a repository whose head commit has the same SHA1 as the one in torvalds: /linux but where any commit was replaced by a malicious commit with the same SHA1 and a fake date. You have no way to detect that my repo is inauthentic other than through a deep history comparison. The whole idea behind a merkle tree is that just checking the hash of the top is sufficient to know the identity of the whole tree.

I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.

amluto41 minutes ago
The date that a repo receives a commit is known to that repo. And a repo can stop accepting new SHA1 objects. And a SHA256 object could have a flag that says that no SHA1 objects may ever reference it.
MBCookabout 3 hours ago
So they’ve been talking about this for many years, planning, and finally announce when they’re going to switch the default.

So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?

I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.

This seems like a bunch of Monday morning quarterbacking.

schaconabout 2 hours ago
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
throwworhtthrowabout 2 hours ago
Your argument is persuasive and well illustrated. I think the problem is the intro paragraphs come off as too certain of catastrophe which, when juxtaposed with your claim that "smarter people than me have been working on this", makes it sound like you don't actually believe they're smarter than you. The rest of your essay feels fair and not judgmental.
nofunsirabout 2 hours ago
The plans have been on display in a cellar. Beware of the leopard.
bawolffabout 1 hour ago
This isn't really applicable. The plans have been talked about for a while very publicly.
MBCookabout 2 hours ago
I don’t understand what this is supposed to mean.
kazinatorabout 1 hour ago
A number of years ago when I heard about this, I was pretty angry and made a private fork of git immediately in which I tried to scrub away the SHA-256 bullshit. But that's basically just paddling upstream with a spoon for a oar.

The stewards of Git are going to do whatever they want, and there is nothing you can do about it if you don't have the clout to create a fork that takes the lead.

No amount of discussion will do anything because they've already decided that their view of the situation is correct. Git hashes are not just content identification but a digital certificate mechanism, and their collision resistance is a grave issue that must be fixed, the end.

You will be browbeaten in any discussion; it's not worth the energy in a world replete with issues.

Read the full thread on Hacker News →

Related stories