Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yes, but next day support makes all the difference. When I worked as an HPC sysadmin, we saw about 2-3 DIMM failures a week in our 10,000 core cluster. I'm sure we could have gotten the hardware cheaper, even with replacement DIMMs for the 3 years period of service, but Dell made it quite frictionless. A simple web login and form and the new DIMM would show up the next day. We ship the old one back (using their shipping label) and move on. Our time was worth more than we would have saved building it ourselves.

500 TB in 1 or 2 TB drives is a lot of hardware. <complete back of the hand calculation> If we have 250 drives with a 1 000 000 hour mtbf, then (according to an exponential distribution), you've got around ~5% chance of a drive failure each week. Not huge, but worth noting. </ complete back of the hand calculation>



Well, a 10k core cluster is a far cry from a 500TB array. The latter fits comfortably into 2-3 racks (if your DC allows the power density). The former sounds like 15+ racks - entirely different ballgame.

I'm absolutely not denying that you should have a dedicated admin to look after such a deployment. But by today's standards it's just not a lot of hardware anymore. And definitely not the size where you must spend a fortune on "platinum" support contracts that can easily cost as much (per year) as a fully fleshed out spare shelf...

FWIW, I'm personally running a few moderately sized storage clusters (the largest is ~200T, 3x3 MD1000) and we're seeing disk failures at a rate of 2-3 per year.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: