On February 1, 2019, domains behind badly behaved DNS servers started failing on resolvers that had been updated. Not because of an outage. Not because of an attack. Because the people who run the world’s DNS resolvers had agreed, months in advance, to stop being nice.
For twenty years, resolvers had quietly been cleaning up after servers that couldn’t follow the rules. On DNS Flag Day, they stopped. And a pile of misconfigured domains found out — all at once — that the safety net was gone.
The workaround that ate two decades
Here’s the thing about DNS: the wire format was frozen in 1987, and RFC 1035 gave you 512 bytes per UDP response. That was fine when a response was an IP address and a TTL. It stopped being fine the moment anyone wanted DNSSEC signatures, or a fistful of IPv6 records, or anything bigger than a postcard.
So in 1999, EDNS(0) arrived — Extension Mechanisms for DNS, originally RFC 2671, now RFC 6891. EDNS is a pseudo-record the resolver bolts onto its query that says, in effect: “I can take bigger responses, and here’s the newer stuff I understand.” It’s the foundation everything modern sits on. DNSSEC doesn’t function without it.
The trouble was that a lot of DNS servers — and, worse, a lot of firewalls sitting in front of them — handled EDNS badly. Some dropped queries carrying an option they didn’t recognize. Some replied with garbage. Some just went silent. A firewall configured in 2004 to block “weird-looking” DNS packets would happily swallow every EDNS query and never say a word.
Resolver authors did the reasonable thing. When an EDNS query got no answer, they’d assume the server was ancient, strip the EDNS bits off, and try again as plain 1987-style DNS. It worked. The domain resolved. Everybody went home happy.
Except “retry without EDNS” is expensive. Each fallback meant waiting out a timeout — often a full second or more — before the second attempt. And it quietly rewarded broken setups: if your firewall was silently eating EDNS, you never found out, because the resolvers of the world papered over it for you, one slow retry at a time.
Twenty years of that. The fallback code got more baroque every release, full of special cases for particular broken behaviors, and it was the kind of thing nobody wanted to touch because nobody fully understood what would break if they did.
Cutting the net
So the major implementers — ISC, who make BIND; NLnet Labs, who make Unbound; PowerDNS; CZ.NIC, who make Knot — plus the big public resolvers at Google and Cloudflare, did something coordinated and slightly ruthless. They picked a date and announced it: as of February 1, 2019, the workaround is gone. If your authoritative server does not answer an EDNS query at all, we treat it as dead. Not “old.” Dead.
BIND dropped the fallback in 9.14. Unbound did it in 1.9.0. Knot Resolver had already tightened up. The public resolvers shipped it on their own infrastructure. The blast radius was narrower than it sounds: a server that correctly answered EDNS — even to say “I don’t support that option” — was completely fine. The ones that got hurt were the servers that met an EDNS query with pure silence. Those domains stopped resolving on any updated resolver — and as the big public resolvers switched over and operators upgraded BIND, Unbound and PowerDNS Recursor (4.2.0), that covered a growing share of users.
It sounds reckless. It wasn’t. The entire point of picking a public date, standing up a website, and building a test page anyone could paste a domain into was to make the failure loud and legible instead of silent and slow. An operator who ran the test and saw red got weeks of warning and one clear thing to fix — usually a firewall rule, sometimes a decade-old DNS appliance. The alternative, carrying the workaround forever, meant those broken configs would never get fixed and everything downstream would keep paying the tax.
This is the part I like. It’s a rare case of internet infrastructure choosing short-term pain to kill long-term rot, on purpose, with a countdown clock. Nobody made them. There is no DNS police. It was a handful of maintainers deciding that enough was enough.
The sequel nobody expected: 1232 bytes
Flag Day 2019 dealt with the servers that couldn’t speak EDNS. Flag Day 2020 went after the ones that spoke it too enthusiastically.
The 2020 problem was fragmentation. EDNS lets a resolver advertise how large a UDP response it’s willing to accept, and for years the common default was 4096 bytes. But a 4096-byte UDP datagram is far larger than the roughly 1500 bytes a typical Ethernet frame carries, so the response gets fragmented at the IP layer into pieces that have to be reassembled on the far end.
IP fragmentation is bad news for DNS specifically. Fragments are easy to lose, easy for a middlebox to drop, and — the alarming part — easier to spoof. An attacker who can force a fragmented response has a much better shot at slipping a forged second fragment into the reassembly, which is a cache-poisoning path. Fragmentation quietly turned a performance annoyance into a security question.
The 2020 fix was almost aggressively plain: advertise a smaller buffer. The recommended number became 1232 bytes, and it comes from arithmetic, not a vibe. Take the minimum MTU that IPv6 guarantees will never need fragmentation: 1280 bytes. Subtract the 40-byte IPv6 header. Subtract the 8-byte UDP header. What’s left, 1232 bytes, is DNS payload that should cross any path in one piece.
The trade-off is honest. If a response is bigger than that, the server sets the truncation bit and the resolver retries the whole query over TCP. Which is exactly why Flag Day 2020’s other demand was that every DNS server — authoritative and recursive — must actually answer over TCP on port 53. “DNS is UDP” was always folk wisdom; TCP has been in the spec since the start. Plenty of operators had firewalled it off anyway, because they’d never needed it. After 2020, they needed it.
What it left behind
The Flag Days didn’t make the news. No CVE, no vendor war room, no breathless postmortem. Just a coordinated decision to stop tolerating broken behavior, executed with a public deadline and a test page.
The lesson isn’t really about EDNS. It’s that compatibility workarounds are debt, and debt compounds. Every “we’ll just handle the broken case” is a small mercy that, multiplied across two decades and a trillion queries, becomes the thing holding the whole system hostage. Once in a while somebody has to declare a flag day, publish the date, and let the broken things break — loudly, on schedule, in front of the people who can fix them.
Most of the internet’s problems never get that treatment. They get a workaround, and then a workaround for the workaround, forever. DNS got two flag days, and it’s better for them.