This is likely the reason the large cloud companies all do testing that actively causes outages. I'm over-simplifying on purpose here: this is something that requires a lot of thought.
At first glance that seems foolish, but to quote you, "complex network topologies" are very prone to falling over badly. Since they all seem to be a one-off custom setup these days, how can you be sure it won't fall over?
2. Google does it. They have a team that goes around unplugging network cables and monitoring how fast the engineers can find and fix the problem. I can't dig it up but it was only a few months ago - hey, Google, your search engine can't find an article about you. :)
At first glance that seems foolish, but to quote you, "complex network topologies" are very prone to falling over badly. Since they all seem to be a one-off custom setup these days, how can you be sure it won't fall over?
Here are the testing approaches I know about:
1. Netflix Chaos Monkey: http://techblog.netflix.com/2012/07/chaos-monkey-released-in... - but that doesn't mean Netflix has it all together. They still have outages.
2. Google does it. They have a team that goes around unplugging network cables and monitoring how fast the engineers can find and fix the problem. I can't dig it up but it was only a few months ago - hey, Google, your search engine can't find an article about you. :)