150 comments
>A man stranded in the bush in northern Saskatchewan was rescued last week after chopping down four power poles — knocking out electricity to surrounding communities. [...]
>But he had an axe and he knew SaskPower would have to check the downed line, so he went to work.
Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.
I wonder how long it was down? Days, weeks, months?
I am paying, $1000, $1800 & $1900 for the same service at 3 different location (20 mins from each other).
Two locations, I also have old coax lines that are still active, but not paying for it.
When I bought two businesses, I learned that they were paying for a dedicated fiber but using coax service.
At one of the location, we had fiber, and paying for backup coax and wireless. But if you turn off fiber box, it wont fail over to either one.
I wouldn't surprise it was down for weeks and no one bothered about it.
For a major trans-oceanic backbone provider, at least 15ish years ago they had a mile or two between Detroit and Chicago where both ends were on the same side of the interstate highway.
But it's more frequent on DAS (Distributed Antenna Systems, AKA small-cell or micro-cell) networks.
Also the challenge of when fibers are leased (if that's still a thing, based on the networks I helped design I'd say 'probably').
They really are analogous to Lamport's "Distributed System" quip; A damaged fiber owned by a company you have never heard of can wreck your day.
Over ten years ago, my employer was spinning up a new DC across the state and had three links between it and the primary DC. Two were pretty direct, but we needed a third because at one point in the 300 mile path, the two main links went within 400 meters of each other.
So how national-security-adjacent critical systems like an airport system doesn't have a larger set of backup links, and validates that they are geographically separate up until the connections, and have realtime alerting on the connection status, is surprising to me.
For mission critical stuff like airports I would like to think they're go for a more rigorous methodology than hope for the best on paths
Recently, I tried calling 811 before digging in my yard. The webpage was broken and the hotline kept me on hold forever. I gave up. Small wonder.
https://blackhydrovac.com/underground-utility-strikes-learn-...
Similar events happened in Europe disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero so... either is an honest accident and is cleared in a few days, or is sabotage
The punchline is that it's about 2000 ft from the local CO, where (I believe) half the town's lines terminate.
A lot of people are incompetent.
There's generally two wrong responses: (1) We spent a lot of money on 'blah blah blah', a lot of other companies use it, so yeah, we've got a backup/failure system. And, (2) inadequate testing - either, we tested 1 of 50 services, and it worked, so the whole system can be restored; or, we gracefully tested, and it worked, so it will obviously work during not-graceful incidents.
And the root cause of this is generally that no one gets promoted for implementing an adequate backup/failure system, or it's extremely rare.
Not saying that's what is going on here, just that it's possible.
Sometimes the level of incompetence / lack of care in organizations like this astounds me. I understand issues like this can be complicated and systemic but it honestly makes me think very poorly of the technologists building these systems in government.
Nothing I've ever seen or experienced with FAA indicates a lack of care; the parsimonious explanation is almost always that people who care a great deal are working within complex systems that don't always have a consistent or externally legible set of priorities. Or more intuitively, the failures we see are the "acceptable" ones versus the unacceptable ones (like planes falling out of the sky).
Hard to fathom someone still holding that view after the 737MAX disasters.
The NTSB has repeatedly and publicly criticized the FAA for failing to implement their recommendations, including as recently as last week (Amazon Prime crash resulting in 5 deaths). They've explicitly called out the FAA's failure to act as a contributing factor in crashes, multiple times.
Regulatory capture ruined the FAA.
BT guy ends that part of the presentation with "the IRA will really have to get shit together to take you off the net"
This was during the troubles. Same planet, different world.
https://www.airwaysmag.com/new-post/faa-smart-first-deployme...
The fundamental problem with DCA, and one of the two main causes of the crash [1], is that they are required by political pressure to operate at a higher operational tempo than they can safely operate at. DCA has essentially 1½ usable runaways--large planes can only use the larger runway, and given the capacity restrictions, airlines have been pushing to use fewer small planes at the airport. This shift in plane size means the effective safe slot capacity has gone down, but people still keep citing the same number as justification for safe numbers, and the politicization of the issue has shut down everyone who complained that the actual traffic just couldn't be safely handled.
If DCA's theoretical slot capacity (36 landings and takeoffs each per hour) is to be reached one just one runway, you have about 100 seconds to go from plane 1 touchdown; exit the runway, letting plane 2 on to take off; plane 2's wheel leaving the runway, clearing plane 3 touchdown. That's doable, but with essentially 0 margin for error. Alternating the runways used for each landing would give you closer to 30s of margin, but if only 10-20% of the planes can use the alternate runway, you can't divert enough planes to use the alternate runway.
Now DCA and the FAA aren't stupid enough to actually schedule 36 landing slots every hour, but the problem is that because of a few various factors, what was scheduled as, say, 30 slots for an hour (giving 120 seconds between landings, probably sufficient margin) ends up being 12 slots used in the first half-hour and 18 slots used in the second half-hour, which means a lot of the actual operation ends up having no safety margin even though on paper you have sufficient margin.
One of the ways you can rectify that is to assign slots further in advance so that you don't get the bunching. Effectively saying "oh, if you leave right now, you'll arrive at 3:30 with three other planes, but if I hold you for 10 minutes, I can push you into a less busy arrival time." Doing this requires good, accurate prediction of the actual flight travel times, and my understanding is that this is what the new software is meant to provide.
[1] The other main cause is essentially that the DoD's aviation practices in the area is a giant clusterfuck that endangers lives, and unfortunately that sentence is not relegated to the past tense.
Good luck to all of us.
Or is it that these ATC networks are their own air-gapped network with less redundancy? That just doesn't add up. Or maybe there was only one line going to the ATC, with no multiple "ISPs" like a datacenter would have?
* There is supposed to be a primary and a secondary link, in this case the primary failed and the fail-over also failed. It's unclear from the reporting if they were damaged in the same incident or if the failover was not tested or monitored adequately.
The Philadelphia TRACON site has been notoriously unreliable and was supposedly improved in 2025, it's also unclear if these issue actually could stem from that implementation.
That seems like an extremely foolhardy thing to do.
It seems like it would present a less-chaotic solution than that provided by having no data communications at all.
Read the full thread on Hacker News →
Related stories
- Hacker News · 3 points · 10 days ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 1 points · 4 days ago
- Hacker News · 1 points · 6 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 1 points · 11 days ago