This is just a reminder, we will be conducting a major system upgrade in 5 hours. All systems will go offline.
The maintenance has been scheduled for 0000 CST on Friday February 22nd 2013.
This is just a reminder, we will be conducting a major system upgrade in 5 hours. All systems will go offline.
The maintenance has been scheduled for 0000 CST on Friday February 22nd 2013.
We’re experiencing a major system outage due to human error at Rackspace. We’re working to fix this as soon as possible.
Hi Everyone,
Just wanted to let you know that we need to take our systems off line for one hour on Friday, February 22nd at midnight CST. The nature of this maintenance is complicated, as it puts the finishing touches on the upgrades that have been under way for the past 8 months. Unfortunately, this means connectivity to your website will be affected.
All websites will be unavailable during this time period. We suggest you suppress your monitors during this time, as there will be nothing we will be able to do about it.
Connectivity will be restored by 1:30AM CST or 0130 CST hours at the latest.
We’ve found some issues with our memcached cluster, which we’re working to resolve as soon as possible. Symptoms of this are slower sites, and sometimes pages that will return a “504 Timeout” page.
Sorry for the issues, we’re working to resolve these issues ASAP.
Update February 12, 2013 8:46 PM: The problem was fixed at 7:30PM CST, we’ve been monitoring the situation for the past hour, and things have been stable. We consider the issue resolved.
We’re currently aware of an issue inside our database cluster that’s causing some slowness/unavailable sites. We’re currently working on the issue and will update you with more details as available.
UPDATE 10:45AM CST: The database connection issues are still ongoing and we’re continuing to investigate the cause.
We’re starting to get things under control. We’ve blocked 2832 unique ip addressess so far. We’re continuing to monitor the situation and isolate the customers who were affected by this, from the customer who was being attacked.
What we know so far
What are we doing to bring customers back?
One of the websites hosted with us is under a major denial of service attack. We’re working with our network security team to isolate the traffic to this site, so we can restore service to normal.
Currently we’re seeing 10,000 requests/second just to this domain on our load balancers.
We experienced intermittent issues with the content delivery network we have. The symptoms of which are that your site doesn’t look correct, all the text loads, but you won’t see images or your stylesheets.
The outage was for about 20 minutes. All systems are functioning now.
As of 2:53 PM on January 27, 2013 All systems are functioning normally again. We had intermittent issues across our network.
Here’s what happened.
One of our 4 memcached servers had run out of memory, and in the process locked up. This made it so that our database servers were seeing 8x the average calls. Since our monitors started telling us about higher than normal database activity, we started investigating the issue there.
What did we learn?
It turns out, our monitoring on the memcached systems isn’t as good as we thought it was. Had we known that the one of the memcached server was out of commission, we would’ve been able to identify the problem, and fix it. Rather than investigating what was causing the spike in the database usage.
We are currently experience an issue that causing sites to display database connection errors.
We are working on having this resolved as soon as possible and will update with more information as it is available. Update: