I believe "our site works when the engineers go on holiday" thing is not fair. Of course that application is less stable when it's actively modified. There's no way to make it more reliable on weekdays than on weekends except for stopping the development altogether or maybe deploying on weekends.
No way at all? FB has a lot of servers and a lot of users. They have opportunities for quality practices that few other organizations get.
Here's a simple example:
Split the user population into 365* groups. Assign them to days of the year. Test new code only the group whose day has come up. Follow them until they stop having problems. Now you can deploy that to a month's worth of groups. All good? Deploy to everyone.
Yes, that means that you can't have more than 365 changes in simultaneous development. Tough.
*Yes, yes, leap years. Take a day off from deploying.
Ack. Also, the system introduces a new problem: If you are deploying on weekend, most devs are not around to help solving the problem. Thus, the outages would be longer.
They might already be doing A/B testing (or in your example A1/A2/.../A365). It's not clear what's the definition of an "incident". At my workplace, even if a bad code push affects say 0.1% of users, it would be classified as an incident.
What other engineering discipline would say "there's no way to improve our reliability except stopping work altogether"? There's always a way to improve reliability. Arguably, with formal verification, you could ensure large parts of your system are perfectly reliable given simple assumptions.
The problem isn't that it's impossible -- it's that it's more expensive than just hiring one more engineer to keep papering over the problems.
What other engineering discipline would say "there's no way to improve our reliability except stopping work altogether"?
All of them. Are your roads more reliable when they're constantly being changed or when they are just being maintained? Is NASA achieving its reliability by constantly changing the designs of their ships, or by reusing the same design over and over?
Arguably, with formal verification, you could ensure large parts of your system are perfectly reliable given simple assumptions.
Yes. But a fixed formally verified system will still be more reliable than a formally verified system being constantly changed.
What was said wasn't that FB couldn't be more reliable. It's that they are already so reliable that only new changes introduce problems. Sure you can still work on minimizing those problems, but that's a different point.