Git 3.0's SHA-256 default is a costly mistake, blog argues

A post on the GitButler blog, written by an author the text addresses as "Scott", takes aim at a planned change in the upcoming Git 3.0 release: the default hashing algorithm moves from SHA-1 to the stronger SHA-256. The author says he has sat on his objections for a few years because smarter people have been working on the problem, but now believes the change will cost everyone a lot of time and angst for little benefit, and that hardly anyone knows what is coming. He calls it a huge, costly, global train wreck of a change for almost no practical value.
He starts with a primer. Git is a content addressable database: it hashes content and uses the hash as the key, so identical content always gets the same hash and is never stored twice. Each commit encodes the hash of the commit before it, so changing anything changes everything after it. That gives Git what he calls cryptographic integrity. Git has always used SHA-1, which Linus picked in 2005. It has worked well for 20 years. By the author's account, an accidental collision has never happened in the history of every file, tree and commit in every Git repository. For SHA-1's 160-bit output, the birthday bound means you would need about 1.4 septillion random files in a single project before hashes collide by accident.
The problem is that SHA-1 is now considered semi-broken. Collision attacks have been published (SHAttered in 2017, SHA-1 is a Shambles in 2020). The author says they are not practical to exploit in any demonstrated way, but they are theoretically possible. "Broken" here only means that finding collisions is not impossible: with enough money and GPUs, you can on purpose produce two different pieces of content with the same hash. The papers suggest purposeful collisions cost on the order of a few tens of thousands of dollars today, which SHA-256 does not allow because of a property of SHA-1's message schedule.
He then separates two attacks. In a collision attack, the attacker generates a benign and a malicious file with the same hash, hands out the benign one until trusted, then swaps in the malicious one; signed tags or commits could even appear to cover the bad file. In a second-preimage attack, the attacker takes an existing file and crafts a malicious one with the same hash, and the original author need not be involved. The author stresses that second-preimage attacks are the more worrying kind, yet nearly no widely used hash function is susceptible to them. Git could use MD5, which is considered completely broken (SHA-1 is not, he says, and is much stronger), and still be effectively immune. His illustration: if all roughly 3 billion GPUs on Earth were replaced with an RTX 5090 and spent all their time on MD5, brute-forcing a particular preimage would still take about 16 billion years expected (about 11 billion median), roughly the age of the universe. So any realistic attack relies on a collision attack, where whoever introduces the original file has already prepared the malicious twin.
Even if a second preimage could be made in an hour for any known file, he asks two questions that the debate rarely addresses: how do you get the file fetched by people who do not know you and have never pulled an earlier version, and how do you get them to run it usefully? His answer is that hashing is not the mechanism of trust in source control. It provides integrity, but trust rests on "where do you pull from?", an argument he says Linus made at the birth of Git. He pulls rust-lang/rust from GitHub because he trusts GitHub's authentication, not because he checks a GPG signature on a commit hash, and he would not pull from a random .onion address because an email said so.
Real-world attacks, he says, look different. Untrusted code gets in because someone socially engineers write access to an npm package used by millions of projects, which compromises every file in the dependency, not one hard-to-craft binary. He calls this maybe a billion times simpler, cheaper and more likely to succeed than a hash collision (his own rough estimate). Likewise, to get bad code into Android it is easier to bribe or convince the maintainer of a popular downstream project; he asks whether any tired open source maintainer would refuse a $40k lump sum for an unloved, heavily used project. In his words, hash collision attacks are maybe the dumbest possible way to get untrusted code onto a system.
His conclusion is that if SHA-1 is treated as a good-enough, fast-enough key generator in a trusted repository, and the hashes are not used as the basis of trust, there is no real reason to replace it; MD5 would probably be fine. Accidental collisions are nearly impossible, and the rest can be handled by signed content verification, external authentication and social trust. If the community does not accept this, he warns, it will be stuck in a loop: a new theoretical collision paper appears, the whole Git ecosystem migrates again, and after years of moving to SHA-256 quantum computers break it and the cycle restarts. The root error, he says, is conflating cryptographic integrity with trust.
He concedes that "train wreck" is probably hyperbole, but says that however the cost is calculated, the migration will not be cheap or easy. He cites Emily Shaffer's recent talk on how Google is preparing for the problem, and says it is not a pretty picture. The post then begins a section on how the rollout of the new default will go, and the account here ends where that section begins.
Key facts
- Git 3.0 plans to change its default hash from SHA-1 (chosen by Linus in 2005) to SHA-256; the author calls this a costly, global change with almost no practical value.
- SHA-1 is called semi-broken because of published collision attacks (SHAttered in 2017, SHA-1 is a Shambles in 2020); the author says they are not practical to exploit in any demonstrated way, with purposeful collisions costing on the order of a few tens of thousands of dollars today.
- The author argues Git could use MD5 and still be effectively immune from a second-preimage attack: brute-forcing a particular MD5 preimage with all of Earth's roughly 3 billion GPUs replaced by RTX 5090s would take about 16 billion years expected.
- His core claim is that trust in source control rests on where code is pulled from, not on the hash; he estimates that an npm package takeover is about a billion times simpler than a hash collision attack.
- He warns of a repeating loop of migrations whenever a new theoretical collision paper appears, and cites Emily Shaffer's talk on Google's preparations as evidence that the migration will not be cheap or easy.
Why it matters
Git is the version control system underneath most software development, and changing its default hash touches every repository, host and tool that stores or transmits Git objects. The author frames this as a breaking change in Git 3.0 that few people are aware of, and challenges the premise behind it: that a stronger hash meaningfully improves security. His argument is that the cryptographic integrity a hash provides has been mistaken for trust, and that the real-world ways untrusted code reaches a codebase, such as a hijacked npm package or a bought-out maintainer, are far cheaper than any hash attack.
Who it affects
Anyone who maintains, hosts or depends on Git repositories, since the author says the whole industry is about to be forced into the SHA-256 migration. He points to Emily Shaffer's talk about how Google is preparing for it as a sign of the scale of the work. Maintainers of popular packages also come into the argument as the weak point he considers more realistic than hash collisions.
How to use it
The post is an argument, not a how-to guide, and it offers no migration steps in the portion covered here. What it offers is a way of thinking: treat SHA-1 as a good-enough key generator for content in a trusted repository, and rely on signed content verification, external authentication and social trust mechanisms for security. For practical preparation details, the author points to Emily Shaffer's talk on how Google is preparing for the SHA-256 migration.
How solid is it
This is a single author's opinion piece, and the key claims are his own judgements: that collision attacks are impractical, that hash attacks are the dumbest route for an attacker, and that an npm takeover is about a billion times easier (explicitly a rough estimate, not a measured figure). The cryptographic background he gives, such as the 160-bit output, the birthday bound of about 1.4 septillion files and the published collision attacks, is stated plainly and consistently. The scenarios with a $40k payment, a one-hour second preimage and 3 billion GPUs are hypotheticals he constructs to make a point. The claim that no accidental SHA-1 collision has happened in Git is limited to what he is aware of.
Risks and caveats
The author himself calls the train-wreck framing probably hyperbole, and gives no figure for the migration cost; "costly" is his characterisation. His argument downplays the risk that published collision attacks may improve, and it runs against the judgement of the "very smart people" who did a huge amount of work on the SHA-256 change. The post also says its reasoning depends on the repository being trusted in the first place; where a repository is not trusted, his reassurance would not apply. The rollout analysis the post goes on to promise is not part of this account.
“Trust is based on "where do you pull from?" and it always has been.”
— From the GitButler blog post on Git 3.0's SHA-256 default