Why bug blindness makes people defend flawed software
A personal essay argues that most people are effectively blind to software bugs and quality problems, even severe ones, while a smaller group of people notice constantly. The author says they personally observe hundreds to thousands of bugs a week and that almost nothing seems to work properly, yet most people they talk to see nothing like this. The explanation, in the author's view, is not that they use computers differently: it is that most people hit the exact same bugs and simply do not register them. For a non-programmer, not noticing is probably the more comfortable way to live, but the author thinks curing "quality/bug blindness" specifically helps programmers, and says that after pointing out bugs to friends and acquaintances for a few weeks, people who are willing to try tend to start noticing bugs on their own.
Because of this trait, the author says they have had multiple jobs in which directors, VPs and other executives asked them to give an honest opinion on whether something actually works. Sometimes nothing is wrong. More often the problems are mild to moderate. Occasionally they are severe enough that the product effectively does not work at all. What the author finds strange is that severe cases are usually accompanied by a trail of internal comments insisting the product is great and works well, even though using it firsthand means routing around it with a string of non-intuitive workarounds that an ordinary user would never find, and would have such a bad experience that they would tell their friends about it. Having watched products launch and then fail in exactly the way they expected, the author no longer believes they are simply triggering unusual edge cases; they add that LLMs now let them simulate a range of ordinary users and confirm that a given issue reproduces across many different scenarios.
To illustrate the pattern without using job-related examples, the author offers several public ones instead, starting with web search. A prior post of theirs documented poor results from Google, Bing and Kagi: pages heavy with low-quality SEO spam and, in some cases, outright scam sites. The author rates this as "moderate" rather than "severe", reserving "severe" for something like a search engine that errors out half the time or returns mostly scam results, on the reasoning that severe means an ordinary user cannot use the product at all, not merely that they have a bad experience. Almost nobody objected to the Google and Bing characterization, but several people insisted the author was wrong about Kagi and sent their own search results as proof. In every case the author reviewed, those results still lacked a genuinely good link and were full of spam, including a seasonal-forecast query that failed to surface a current forecast. One person shared their Kagi filters and results without arguing either way. The one partial exception was a user who had pinned GitHub to the top of their results, which worked for queries aimed at downloading GitHub-hosted software but, by the author's account, failed completely for the other queries from the original post.
A second example concerns cars. The author bought a Volvo after seeing it perform well in out-of-sample crash tests, then began reading Volvo owner forums. Reliability data going back well over a decade, which the author says is backed by the anecdotal experience of mechanics who service Volvos, points to mediocre-to-poor reliability. Volvo forums, by contrast, are full of owners insisting Volvos are among the most reliable cars on the road and that the data is simply wrong.
A third example is Blackboard, once the dominant course-management platform at universities. The author describes it as, informally, the most widely disliked software in their social circle at the time, trailing only Visual Source Safe in unpopularity, since Source Safe was disliked even more strongly but was not used widely enough to take the top spot. Blackboard's own Wikipedia page states that the company had become one of the most disliked, even detested, companies in education, and cites a Fast Company report from December 2011 that 93% of respondents to an Amplicate customer-opinion survey said they hate the company. Despite that, the author recalls once asking a Blackboard employee, more or less, what it was like to work on software so many people disliked; the employee was confused rather than offended, believing the software was genuinely well liked, and did not accept that the author's description could be true.
A fourth example is Discourse, the forum software, whose employees told the author they believed its web performance was great. Investigating, the author says they found code inside Discourse that deliberately slowed down real page loads specifically to improve reported performance metrics such as LCP. In the author's framing, that goes beyond ordinary benchmark optimization into outright cheating, since it gave users no benefit and actively made their experience worse.
For a non-programming example, the author points to an unnamed basketball player widely seen by fans as the dirtiest of his era. The NBA does not officially track player dirtiness, or genital strikes specifically, but the author argues he almost certainly holds the unofficial record this century for punching, kicking, kneeing or otherwise striking opponents in the genitals, and should lead on an era-adjusted basis too, though he may fall short of the all-time record because play generally was dirtier in the 1980s and 1990s. Fans of his own team, the author says, wave this away with running jokes like "natural rebounding motion" and "natural shooting motion" whenever he contorts himself into an opponent.
The author generalizes from these cases: people have a strong tendency to ignore negatives in things they are fans of, including, especially, their own work or their employer's work. The author says they are unusually free of this particular blind spot, actively seeking out people who can poke holes in their reasoning and generally rating their own published work as full of major flaws. When readers have asked how the author would feel if someone criticized their work, or said outright that it isn't good, the author's answer is that reasonable criticism is welcome and that, overall, calling the work not good seems like a fair assessment, even though certain parts of it strike the author as interesting or worthwhile.
The essay then turns to what it calls habitual mitigations, opening with a childhood story from the era of mechanical computer mice. A friend found the author's mouse almost unusable because the pointer moved nearly randomly, the result of detritus built up on the mouse ball degrading its tracking. The author, using the same mouse right after, had no trouble at all, but noticed they were throwing their hand around in wildly erratic movements to compensate, an adaptation built up gradually as the grime accumulated and never consciously registered. That leads the author to wonder whether they are doing some equivalent, unnoticed compensation today, and into a discussion of specific personal habits developed to work around known bugs, starting with an example involving opening a new Google Doc. The excerpt available here ends mid-sentence at that point, before the example is finished or the essay reaches its conclusion.
Key facts
- The author says they personally observe "hundreds to thousands" of bugs a week, while most people they talk to do not notice the same problems at all.
- Testing Google, Bing and Kagi search results, the author calls the pattern "moderate" rather than "severe", and says every reader-submitted Kagi result they later checked was still spam-filled and lacked a good link, except when a user had pinned GitHub to the top purely for GitHub downloads.
- Volvo forums insist the cars are among the most reliable on the road, even though reliability data covering the last decade or more, backed by mechanics who service Volvos, shows mediocre-to-poor reliability.
- Blackboard's own Wikipedia page cites a Fast Company report from December 2011 that 93% of respondents to an Amplicate customer-opinion survey said they hate the company, yet a Blackboard employee the author once spoke to believed the software was well liked.
- The author says Discourse contained code that deliberately slowed real page loads to game performance metrics such as LCP, which they call outright cheating rather than benchmark optimization, since it only hurt users.
Why it matters
The essay's core claim is that most people, including experienced professionals, are broadly blind to the bugs and quality problems in the software and products around them, and that the blindness gets stronger the more someone likes or is invested in the thing being judged. That matters for anyone who ships software: internal praise is not reliable evidence on its own, and a product can be severely broken in ways only a genuinely fresh, skeptical look will catch. The author frames the blindness as trainable: pointing out bugs repeatedly to another person tends to make that person start noticing bugs unprompted within a few weeks.
Who it affects
Programmers and product teams who rely on internal sign-off before shipping are the direct audience, since the essay argues that a chorus of insiders calling something great is compatible with it being unusable. It also implicates fan communities and company staff who defend a product against documented problems, illustrated with Volvo owners, Kagi users, and people who worked on Blackboard and Discourse. More broadly it applies to anyone judging their own work or their employer's work, where the author says the same blindness is strongest of all.
How to use it
The essay's practical suggestion is to deliberately coach bug blindness out of people: pointing out bugs to friends and colleagues, repeated over a few weeks, tends to make willing people start noticing bugs on their own. The author also describes using LLMs to stand in for a range of ordinary users and check whether a suspected bug reproduces across many different scenarios, as a way to rule out that they are personally hitting an unusual corner case. A third habit is treating a product as genuinely broken when it seems severely flawed under firsthand testing, rather than assuming average users would not run into the same wall.
How solid is it
This is a first-person opinion essay, not a study: the author says the idea sat unpublished for roughly a decade before this write-up, and the evidence is a set of individually recounted cases rather than a controlled comparison. The search-engine and Volvo claims rest on the author's own testing plus unsolicited examples that readers sent in; the Blackboard claim rests on the company's own Wikipedia page, which in turn cites a Fast Company report on a December 2011 Amplicate survey; the Discourse claim rests on the author's own description of code they say they found. The author explicitly withheld what might be the strongest evidence, job-related examples, calling the public ones "less interesting and less well supported" by comparison. The excerpt available for this summary also ends before the essay reaches its own conclusion.
Risks and caveats
Several details that would let a reader verify the specific cases are missing from the text: the author, the basketball player, and the Blackboard employee are all unnamed, and there is no count of how many people sent Kagi results or what share of tested queries failed. Whether the Discourse code the author describes is still in place, or was ever changed, is not stated. The whole piece is built on the author's personal judgment of what counts as a bug and how severe it is, so its force depends on trusting that judgment case by case rather than on independently checked data.
“I easily observe hundreds to thousands of bugs per week and nothing seems to work, but most people I talk to don't see anything like this.”
— the essay's author