We’re removing a plugin that is known to put malware onto our customers sites off our systems. This will be in effect until the plugin has been re-instated at WordPress.org. Please see this blog post by Sucuri for more details on what this plugin does.
Caching Server Issues
There’s an issue with some of our caching services causing intermittent site issues and downtime. We are working with our providers and operations team to get everything online. This blog will be updated with more information when it becomes available.
UPDATE 7:55PM CST: At this time we’ve reconfigured our caching setup to remove the impacted hardware from production. We’re still evaluating the root cause, but expect sites to be stable now.
Scheduled Maintenance Sunday April 7 0030 – 0230 (12:30AM – 2:30AM) CST
In our ongoing efforts to give you the best WordPress Hosting on the planet, we need to make some system wide upgrades to our infrastructure. These upgrades will be happening on Sunday morning at 12:30AM CST, and will last 2 hours, till 2:30AM.
During this maintenance window, all websites will show a generic message letting you and your visitors know that the website is down for maintenance. We’re sorry for the downtime, but this is absolutely necessary, to keep up with our growth, and to continue to serve you in 2013.
For the geeks: We’re upgrading our NAS again, adding drives that are much faster, which should result in “Zippier” websites.
All systems back up and functional
Thank you for your patience. The upgrade has gone smoothly, and all systems are up and running. We’ve gone through an internal check list.
There is a very small chance that we missed something, if something is not working for you, please open a ticket, and we’ll look into it immediately.
Reminder about scheduled maintenance in 5 hours
This is just a reminder, we will be conducting a major system upgrade in 5 hours. All systems will go offline.
The maintenance has been scheduled for 0000 CST on Friday February 22nd 2013.
All systems down, due to an error at Rackspace we’re fixing the issue now
We’re experiencing a major system outage due to human error at Rackspace. We’re working to fix this as soon as possible.
Scheduled Maintenance – February 22nd 2013 – 0000 hours
Hi Everyone,
Just wanted to let you know that we need to take our systems off line for one hour on Friday, February 22nd at midnight CST. The nature of this maintenance is complicated, as it puts the finishing touches on the upgrades that have been under way for the past 8 months. Unfortunately, this means connectivity to your website will be affected.
All websites will be unavailable during this time period. We suggest you suppress your monitors during this time, as there will be nothing we will be able to do about it.
Connectivity will be restored by 1:30AM CST or 0130 CST hours at the latest.
Caching layer degradation – Fixed and Stable
We’ve found some issues with our memcached cluster, which we’re working to resolve as soon as possible. Symptoms of this are slower sites, and sometimes pages that will return a “504 Timeout” page.
Sorry for the issues, we’re working to resolve these issues ASAP.
Update February 12, 2013 8:46 PM: The problem was fixed at 7:30PM CST, we’ve been monitoring the situation for the past hour, and things have been stable. We consider the issue resolved.
Database and Service Interruptions
We’re currently aware of an issue inside our database cluster that’s causing some slowness/unavailable sites. We’re currently working on the issue and will update you with more details as available.
UPDATE 10:45AM CST: The database connection issues are still ongoing and we’re continuing to investigate the cause.
Update on the botnet attack of February 7, 2013
We’re starting to get things under control. We’ve blocked 2832 unique ip addressess so far. We’re continuing to monitor the situation and isolate the customers who were affected by this, from the customer who was being attacked.
What we know so far
- A customer’s website is under a botnet attack, where we are seeing 190,000 requests/second made to one ip address.
- These requests seem to be coming from about 3000 unique ip addresses.
- Our firewall was reaching a CPU max of about 90% while this was happening, our alarms go off when it hits 51%.
- Blocking all 3000 ips on the firewall is not a good idea, so we’ve “null routed” the destination ip address.
What are we doing to bring customers back?
- We are assigning new ips to the affected customers (several hundred) who shared the same ip address with this customer.
- If we control your dns, this change will happen within the next 30 minutes. If we don’t, we’ll be contacting you to let you know what the ip address should be.