
A drive failed, at 4:00 on a Friday (2 weeks ago)… Then at about 11 that night, the rebuild failed. I thought it was the end of the world; especially since we’re severely overextended with our storage.
I’ve never experienced anything like that on a SAN, so I was very relieved to see it really didn’t phase it much. I was afraid the whole thing would collapse and I would have to rebuild everything from scratch.
But we got lucky and the bad blocks belonged to only one important volume, the rest were snapshots and backups I hadn’t bothered to delete. The only one in use was left as online-bad-blocks while the other backup volume was left offline-bad-blocks (but it was only a temporary replica from last year).
I just deleted the damaged snapshots and the old backup volume. But the one in use (one of the first volumes I created) seemed to be working fine… the EQL tech said it will only throw an error to the OS if it tries to read one of the bad blocks… and we cant really tell if they were in use or just free space in the VM… but if they are just overwritten by the OS, then the array will mark them as good and move on. I migrated the 2 production VMs on that volume to another and it didn’t complain.
I’m about to pull the “bad” drive… though the array still thinks its good, the tech and I both agreed that it had too many errors in its stats… so they sent me another. Decided to wait until the weekend before we chanced anything.
Now I’m trying to decide if I want to convert the RAID50 to a RAID6 (once the rebuild is complete of course) EQL seemed to think there wouldn’t be any discernible difference in terms of performance. And I’d rather have protection against any 2 drive failures than hope the two that fail aren’t in the same set. With drives as large as they are today, its just an accident waiting to happen.
All things considered, we’re very happy with our purchase… now we just need another.