239 points•allanbreyes•9 days ago•150 comments•

150 comments

chasd009 days ago
It happens. There use to be a joke during the first big DC build out phase that went like if you're ever going into the wilderness take a 1ft length of fiber optic cable with you. If you get lost bury it and a back hoe operator will appear and sever it within an hour. You can get a ride back with them.
john_strinlai9 days ago
reminds me of https://www.cbc.ca/news/canada/saskatchewan/stranded-man-cut...

>A man stranded in the bush in northern Saskatchewan was rescued last week after chopping down four power poles — knocking out electricity to surrounding communities. [...]

>But he had an axe and he knew SaskPower would have to check the downed line, so he went to work.

foxyv8 days ago
What gets me is the Counselor was like "Why didn't he just build a big bonfire" as if a big fire would be a better idea.
castillar769 days ago
Nothing like summoning the North American Fiber-Seeking Backhoe! https://imgur.com/a/DZNhdIG
baud1472589 days ago
I remember a story from an acquaintance working in construction. In one case they had to dug up right where a fiber optic ran, so they had cleared it with its operator and waited for their go-ahead. And waited. And waited, until the boss said something to the effect of 'start digging, we have a schedule to keep'...
Onavo9 days ago
I think it happened many times to electrical lines due to lost campers.
cube009 days ago
> When it went to flip into the backup, we discovered that the backup fiber had a break

Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.

I wonder how long it was down? Days, weeks, months?

newhotelowner9 days ago
All these cable/internet service companies are incompetent as it gets.

I am paying, $1000, $1800 & $1900 for the same service at 3 different location (20 mins from each other).

Two locations, I also have old coax lines that are still active, but not paying for it.

When I bought two businesses, I learned that they were paying for a dedicated fiber but using coax service.

At one of the location, we had fiber, and paying for backup coax and wireless. But if you turn off fiber box, it wont fail over to either one.

I wouldn't surprise it was down for weeks and no one bothered about it.

Neywiny9 days ago
Any possibility of line of site wireless?
toast09 days ago
Sometimes your multiple fiber paths end up in the same bundle, severed by the same backhoe. It's always a fun day when that becomes apparent.
to11mtm9 days ago
I actually got to see one of those once.

For a major trans-oceanic backbone provider, at least 15ish years ago they had a mile or two between Detroit and Chicago where both ends were on the same side of the interstate highway.

But it's more frequent on DAS (Distributed Antenna Systems, AKA small-cell or micro-cell) networks.

Also the challenge of when fibers are leased (if that's still a thing, based on the networks I helped design I'd say 'probably').

They really are analogous to Lamport's "Distributed System" quip; A damaged fiber owned by a company you have never heard of can wreck your day.

unethical_ban9 days ago
That is negligence, practical if not contractual.

Over ten years ago, my employer was spinning up a new DC across the state and had three links between it and the primary DC. Two were pretty direct, but we needed a third because at one point in the 300 mile path, the two main links went within 400 meters of each other.

So how national-security-adjacent critical systems like an airport system doesn't have a larger set of backup links, and validates that they are geographically separate up until the connections, and have realtime alerting on the connection status, is surprising to me.

Havoc9 days ago
>Sometimes your multiple fiber paths end up in the same bundle

For mission critical stuff like airports I would like to think they're go for a more rigorous methodology than hope for the best on paths

koolba9 days ago
Put all your backups in one basket, and then pray that nobody crushes the basket.
konfusinomicon9 days ago
last time this happened in my area a local farmer was burying a cow and took out the whole towns connection
krashidov9 days ago
Could this be sabotage?
kqgnkqgn9 days ago
That seems more than merely plausible given current tensions and the location/timing connection with NYC and UN.
wiml9 days ago
Sure. It could also be the first sign of the invasion of the Mole-Men.
imglorp9 days ago
There are 400,000 to 800,000 utility strikes a year in the US. Sure you could bury (!) an instance or two of sabotage in there with an "oops".

Recently, I tried calling 811 before digging in my yard. The webpage was broken and the hotline kept me on hold forever. I gave up. Small wonder.

https://blackhydrovac.com/underground-utility-strikes-learn-...

pvaldes9 days ago
I would put my money on it.

Similar events happened in Europe disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero so... either is an honest accident and is cleared in a few days, or is sabotage

jasonjayr9 days ago
I drive by a VZ pedestal that has had it's service doors wide open for about 3 weeks.

The punchline is that it's about 2000 ft from the local CO, where (I believe) half the town's lines terminate.

readthenotes19 days ago
A lot of people don't check their backups until they need to restore.

A lot of people are incompetent.

andrewjf9 days ago
Nobody wants backups as a feature. The feature is restore.
jimt12349 days ago
I've worked on backup/failure systems since the mid-90s, and I've found there's one universal truth: If you don't fully test your backup/failure system, you don't have a backup/failure system.

There's generally two wrong responses: (1) We spent a lot of money on 'blah blah blah', a lot of other companies use it, so yeah, we've got a backup/failure system. And, (2) inadequate testing - either, we tested 1 of 50 services, and it worked, so the whole system can be restored; or, we gracefully tested, and it worked, so it will obviously work during not-graceful incidents.

And the root cause of this is generally that no one gets promoted for implementing an adequate backup/failure system, or it's extremely rare.

thewebguyd9 days ago
A backup without a restore test isn't a backup at all
ranger_danger9 days ago
Besides what the others have said, speaking from a neteng perspective, sometimes backup lines (and the core infrastructure in general) are engineered in such a way that it can't easily be tested properly without taking other things down, or manually rolling a truck specifically to test it in isolation with extra equipment.

Not saying that's what is going on here, just that it's possible.

kqgnkqgn9 days ago
Building multiple diverse paths and monitoring for fiber cuts is not hard. Having 2 fiber paths is insufficient for even only moderately important workloads at a tech company. Overlapping fiber cuts happen. For something with significant economic and safety impact this is just crazy.

Sometimes the level of incompetence / lack of care in organizations like this astounds me. I understand issues like this can be complicated and systemic but it honestly makes me think very poorly of the technologists building these systems in government.

woodruffw9 days ago
They had two diverse paths, but apparently did not have a (reliable?) process for ensuring the backup path was functional during fail-over. I imagine there will be some soul-searching over that.

Nothing I've ever seen or experienced with FAA indicates a lack of care; the parsimonious explanation is almost always that people who care a great deal are working within complex systems that don't always have a consistent or externally legible set of priorities. Or more intuitively, the failures we see are the "acceptable" ones versus the unacceptable ones (like planes falling out of the sky).

markdown9 days ago
> Nothing I've ever seen or experienced with FAA indicates a lack of care

Hard to fathom someone still holding that view after the 737MAX disasters.

The NTSB has repeatedly and publicly criticized the FAA for failing to implement their recommendations, including as recently as last week (Amazon Prime crash resulting in 5 deaths). They've explicitly called out the FAA's failure to act as a contributing factor in crashes, multiple times.

Regulatory capture ruined the FAA.

OutOfHere9 days ago
The problem is with considering the second path a backup path. It is a common problem with things that are considered to be a backup. In contrast, both paths should ideally have been in regular use. As a general policy, for resiliency, one should always exercise all of one's routes/suppliers/vendors, even if some are suboptimal, not just the primary one.
jcrawfordor9 days ago
The FAA contracts virtually all of its communication infrastructure now, so they probably don't even know much about these details. L3Harris has the contract for microwave radio but the fiber stuff seems scattered depending on when it was built.
rasz9 days ago
Funny its FAA. In planes you normally encounter dual-channel FADECs (dual ECUs) that alternate authority automagically during pre flight check.
RyJones9 days ago
I've told this story so many times I've worn the corners off it. BT was giving a presentation to MSN WAN OPS about the new datacenter buildout in London for us. They kept going on and on about physical security, man traps, and that our fiber left the building on each side and didn't get close to each other for so many km away.

BT guy ends that part of the presentation with "the IRA will really have to get shit together to take you off the net"

This was during the troubles. Same planet, different world.

dgellow9 days ago
I’m pretty sure that type of project is contracted to the private sector and isn’t caused by government technologists in any way…
kqgnkqgn9 days ago
That sounds pretty plausible. But you still have to have enough internal competency to oversee a contractor, and ask the right questions/build proper requirements and verify they are meeting them.
sc68cal9 days ago
There is also a new ATC system being rolled out, as early as today

https://www.airwaysmag.com/new-post/faa-smart-first-deployme...

cmiles89 days ago
There’s been “a new ATC system” rolling out for the last 25 years.
jcranmer9 days ago
Over the weekend, I got linked to https://admiralcloudberg.medium.com/reaping-the-whirlwind-in..., which is an exhaustive look at the causes of the first airspace collision in the US in several decades. After reading that, my conclusion is that, if the software being rolled out is what I think it is, it would be a borderline impeachable offense to not first trying it out in the DC airspace.

The fundamental problem with DCA, and one of the two main causes of the crash [1], is that they are required by political pressure to operate at a higher operational tempo than they can safely operate at. DCA has essentially 1½ usable runaways--large planes can only use the larger runway, and given the capacity restrictions, airlines have been pushing to use fewer small planes at the airport. This shift in plane size means the effective safe slot capacity has gone down, but people still keep citing the same number as justification for safe numbers, and the politicization of the issue has shut down everyone who complained that the actual traffic just couldn't be safely handled.

If DCA's theoretical slot capacity (36 landings and takeoffs each per hour) is to be reached one just one runway, you have about 100 seconds to go from plane 1 touchdown; exit the runway, letting plane 2 on to take off; plane 2's wheel leaving the runway, clearing plane 3 touchdown. That's doable, but with essentially 0 margin for error. Alternating the runways used for each landing would give you closer to 30s of margin, but if only 10-20% of the planes can use the alternate runway, you can't divert enough planes to use the alternate runway.

Now DCA and the FAA aren't stupid enough to actually schedule 36 landing slots every hour, but the problem is that because of a few various factors, what was scheduled as, say, 30 slots for an hour (giving 120 seconds between landings, probably sufficient margin) ends up being 12 slots used in the first half-hour and 18 slots used in the second half-hour, which means a lot of the actual operation ends up having no safety margin even though on paper you have sufficient margin.

One of the ways you can rectify that is to assign slots further in advance so that you don't get the bunching. Effectively saying "oh, if you leave right now, you'll arrive at 3:30 with three other planes, but if I hold you for 10 minutes, I can push you into a less busy arrival time." Doing this requires good, accurate prediction of the actual flight travel times, and my understanding is that this is what the new software is meant to provide.

[1] The other main cause is essentially that the DoD's aviation practices in the area is a giant clusterfuck that endangers lives, and unfortunately that sentence is not relegated to the past tense.

kj4211cash9 days ago
That system wouldn't have replaced the one that went down. It's good to see investment in this area but that system is not going to fix much TBH.
whatever19 days ago
I suspect it is vibe coded. Contract won in June. Deploying in September.

Good luck to all of us.

gkedzierski9 days ago
I can guarantee you that it was not. Check out DO-278A, and its standards and guidelines that must be followed.
macintux9 days ago
Deploying first to DC? Yeah, that’ll end well.
atonse9 days ago
I don't quite understand this sort of thing happening. Wasn't the whole point of the internet to be a self-healing network where we route around severed cables, etc?

Or is it that these ATC networks are their own air-gapped network with less redundancy? That just doesn't add up. Or maybe there was only one line going to the ATC, with no multiple "ISPs" like a datacenter would have?

bri3d9 days ago
* Yes, TRACON have their own dedicated links. You can look into ASTERIX and STARS to learn some of the cursed ways the data processing and dataflow work.

* There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.

The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.

cmiles89 days ago
Reporting is saying that the backup was cut a while ago and they only discovered it when the primary failed. Apparently nobody was ping testing the backup link.
wrs9 days ago
I’ve had enough double-redundant links fail - and I mean pretty thoroughly double-redundant, different provider, different media, different physical path - that I’m quite surprised something like the air traffic control system only has two links.
readthenotes19 days ago
Do you want air traffic control to be on the general internet?

That seems like an extremely foolhardy thing to do.

ssl-39 days ago
As a tertiary backup: Yeah, maybe I do want that.

It seems like it would present a less-chaotic solution than that provided by having no data communications at all.

kqgnkqgn9 days ago
You certainly don’t need to run this over public Internet to get resiliency over multiple paths.
bell-cot9 days ago
Why not? Assuming it's competently set up - serious encryption, details not blabbered about (to attracted DDoS or whatever), etc.
Hikikomori9 days ago
It is, but you need to set it up correctly.

Read the full thread on Hacker News →

Related stories