ImageSilo Service Bulletin Update

 This is an update pertaining to the ImageSilo Service Bulletin issued the evening of Feb 5, 2010 and the subsequent Emergency Maintenance opened the following morning. Full service and full redundancy has been restored and operational since yesterday. Two separate issues came to light: (1) a fairly seldom seen Microsoft SQL Server bug and (2) a hardware failure with fibre equipment linking a backup SQL server to its database storage array. Below is a more detailed outline of the events of Feb 5 and 6:

On Feb 5, 2010, at 7:38 PM MST ImageSilo’s internal monitoring systems alerted engineers of an issue with a SQL database cluster. Initially (as communicated in the Feb 5 Service Bulletin) , it was thought that this issue had affected all ImageSilo customers. Instead, it turned out that the issue had affected only a single customer’s database. However, by 7:44 PM MST, the issue began affecting other databases. ImageSilo engineers immediately made the decision to manually fail the database cluster over to its redundant partner node. As a result of this issue and the resultant failover (which took considerably longer than expected), access interuption for all customers began at 7:45 PM MST and service was restored by 7:51 PM MST (except for the single customer database that was still being affected).

Engineers worked through the night, enlisting the help of escalated Microsoft support, to determine the cause of the problem. By early Saturday morning (Feb 6), it was determined that the cause of the initial problem was a well known (within Microsoft) but seldom seen bug in SQL Server. An Emergency Maintenance Notification was issued to provide the oppportunity to apply a Microsoft-supplied patch to address the bug, as well as to identify and address a new issue that appeared on the database node that was carrying the problematic node’s load (which was also the cause of the extended period of time taken for the cluster failover the previous evening).

By 11:57 AM MST (Feb 6), engineers had identified that the current database node was experiencing an issue with one of the fibre links connecting it to its database storage array. The hardware vendor immediatley dispatched parts to replace the failed/failing equipment and the bad link was fully restored and tested by 2:02 PM MST. Engineers then went back to work to apply the Microsoft patch to address the SQL Server bug. By 8:59 PM MST the patch had been installed on all cluster nodes and multiple tests had been performed to ensure that the software and hardware was operating as expected.

As of the time of this writing (1:34 PM MST, Feb 7), the issues appear to have been completely addressed by the patch applied and the replacement of the failed hardware. If the status should change, another Service Bulletin Update will be posted. Otherwise this issue is considered closed.

As always, we appreciate your continued support and look forward to continuing to provide you with the high level of service you’ve come to enjoy as an ImageSilo customer. If you have any questions or comments, feel free to contact ImageSilo Administration.

ImageSilo Administration
866.374.3569 (303.493.6900)
siloadmin@imagesilo.com

Leave a Reply

Your email address will not be published.