Drive Crash 2: The fscking: Difference between revisions

no edit summary
m (+link)
No edit summary
Line 21: Line 21:
== The Fixing of Fsrv ==
== The Fixing of Fsrv ==


At this point, everything was considered good to go - we had a full, working backup on a redundant array, and would only need a couple of days to rebuild the array on Fsrv and move everything back. So the array was rebuilt as a 3TB software RAID.
At this point, everything was considered good to go - we had a full, working backup on a redundant array, and would only need a couple of days to rebuild the array on Fsrv and move everything back. So, we shut down fsrv with the intention of rebuilding the array in the RAID BIOS.


== The Balls-up ==
== The Balls-up ==


During the rebuild, a [[docs:High Impedance Air Gap | High Impedance Air Gap]] developed between the power cable providing power to backup and the power supply of backup, causing an expected power-down. Upon restart, it was discovered that the EXT4 partition holding all the data had become corrupt. It is suspected that the drive being over 95% full didn't help with this.
As fsrv posted, there was one immediate concern with the POST data - the RAID controller insisted that the existing array had failed, as all four of the drives connected to it were no longer there. However, they were replaced with four completely different drives, despite the fact they were identical down to the serial numbers. It was at the first stage of investigating this that a [[docs:High Impedance Air Gap | High Impedance Air Gap]] developed between the power cable providing power to backup and the power supply of backup, causing an expected power-down. At this time, it was not noticed that the noise heard was that of backup restarting, and was disregarded whilst fsrv was told to start building a new array on its "new" drives.
 
Of course, once sufficient array building was done to fsrv to make it unlikely anything was ever coming back out of it, the source of the earlier electrical crackling noise was investigated. Once ystvbackup was brought back up, it was only natural for our worst fears to be correct - the backup of fsrv's recently wiped data was corrupt.
 
== The "Recovery" ==
Over the following weeks, multiple attempts were made to recover the data from what was left of this filesystem using dd images, but no luck was to be had. Alex Williams is currently known to be working on writing Python script to manually repair the inodes - we know for a fact the data is still on the disk, it is just the drive metadata that is lost.
 
Since this was just after the annual NaSTA deadline, this wasn't as catastrophic as Drive Crash Classic, but meant that the content that won us Best Broadcaster 2014 made its way happily to the judges. For much of the content that was on the drives, producers expressed relief at no longer needing to find time to edit the content. For the content that was important, most of it was on local tempvideo drives on the edit machines or still had the original recordings from Kenobi or Vidsrv, so the only real loss was show resources and finished shows, the latter of which can be reconstructed using data from web and playout.
 
== The Conclusion ==
Never, ever have Just One copy of you data. Especially when you ask yourself "Should I make an extra copy on this spare disk I have too?" - the answer should not have been No to that question.
44

edits