Chicago Datacenter Issue – RESOLVED

Howdy,

We just finished dealing with an issue in our Chicago datacenter that was causing several other clusters to experience instability. Our “Ursa” cluster was taking on an extreme amount of traffic that looks to be, in large part, a bot attack.

This happened as a result of the “Ursa” cluster having a set of tools not running appropriately that detects and mitigates issues like this.

We’ve cleared this up and all sites are now back up and running appropriately.

If you have any questions or concerns, please contact our help desk via your https://my.pressable.com control panel.

Thank you!

RESOLVED: Chicago and Virginia Datacenter Outages.

Due to our upstream provider having connectivity issues, we are currently experiencing downtimes at our Chicago and Virginia Datacenters. We are currently working with our provider to correct this issue. We will continue to update this post as we get more information.

UPDATE 4:41 PM CST: After further investigation, it appears that this outage is only affecting customers in our Chicago Datacenter.  We are still trying to gather more information from our provider so we can provide a possible ETA. 

UPDATE 5:15 PM CST: We are still working with our provider in order to diagnose this issue. We should have a more detailed update very soon. 

UPDATE 6:30 PM CST: We are seeing some clusters begin to online. We are working on rolling out the rest of our clusters to full functionality now. We will update once that is done. 

UPDATE 7:47 PM CST: We have restored functionality across our systems. If you are continuing to see issues with your site/s, please submit a ticket via your my.pressable.com control panel. 

 

RESOLVED: WooCommerce SQL Injection Vulnerability

Earlier today the Wordfence Security team released the details of a WooCommerce SQL Injection Vulnerability. Our systems are already at work patching this popular plugin across sites on our systems. We’ll provide an update when the process has been completed.

UPDATE March 14th, 7:55AM CST: At this time all sites on our systems have been updated to the latest (patched) version of WooCommerce. If you have any questions, please don’t hesitate to reach out.

Rackspace Scheduled Critical Maintenance

This is a notice that Rackspace will be performing critical security related updates to many cloud server host machines in order to patch vulnerabilities in Xen Hypervisor.

You can read more about this maintenance here:

https://community.rackspace.com/general/f/53/t/4978

These patches/updates will require host machines to be rebooted, subsequently causing cloud servers hosted on them to require a reboot as well.

As it relates to our customers, here are the maintenance windows that we have been provided with and can expect we will begin seeing server reboots occur based on cluster:

  • Hyperion, Pegasus, Cartwheel Clusters
    • Tuesday, March 3rd 01:00 – Tuesday, March 3rd 05:00 EST COMPLETE
  • Galaxy01, Thor, Bode, Ursa, Hydra Clusters
    • Wednesday, March 4th 22:00 – Thursday, March 5th 06:00 CST
    • Thursday, March 5th 22:00 – Friday, March 6th 02:00 CST

To find out which cluster your sites are on, please reference our knowledge base article on identifying which cluster your site is on.

We definitely understand these kinds of outages are not ideal but we are hoping this early notice is helpful in the way of being able to notify your users, visitors, and customers.

If you have any questions, please feel free to contact the help desk via your my.pressable.com control panel.

Thank you!

IAD Datacenter Issues

We are experiencing an issue with sites in our IAD Datacenter that is causing them to not load appropriately.

We believe this is related to issues that Rackspace is currently having with their cloud block storage services. We are working with them directly and awaiting further details regarding this issue and will update the status blog with more information as it becomes available.

If you have questions or concerns, please create a help desk ticket via your https://my.pressable.com panel or join us in our community lounge at http://chat.pressable.com for updates while we await further details.

UPDATE Feb 27, 2015 @ 5:00 AM Central: We were able to confirm that this is an issue occurring at Rackspace with their Cloud Block Storage service. You can find more details and information on their status page: https://status.rackspace.com/index/viewincidents?start=1425013200

We will update again as more information is available.

UPDATE Feb 27, 2015 @ 5:55 AM Central: Rackspace has resolved the issue on their end and we are now working on re-establishing stability on our end. We will update again as soon as this is taken care of.

UPDATE Feb 27, 2015 @ 6:10 AM Central: We have now restored functionality across our IAD datacenter and all sites are now functioning normally. If you continue to experience problems, please submit a help desk ticket via your https://my.pressable.com panel.

Rackspace Scheduled Maintenance, February 18th 12:00am-6:00am CST

This is a notice that Rackspace will be performing software upgrades on network switches in our Chicago datacenter. These upgrades will require a reboot of these network switches, which will result in approximately five minutes of downtime for customers in this datacenter.

The maintenance is scheduled for Wednesday, February 18th during the window of 12:00am-6:00am CST.

Please note that this window can not be moved to another time due to the nature of our Provider. We apologize in advance for any inconvenience this may cause.

If you have any questions, please feel free to contact the helpdesk, via your my.pressable.com control panel.

Slow connections causing issues with sites.

We are currently experiencing an attack similar to the attack we had 2 weeks ago, we have isolated the target which is towards the cluster “galaxy01”.

We are working with our provider now to resolve this as quickly as we can.

UPDATE: 4:35 P.M.- We’re working on identifying which specific set of servers are being attacked and will have a fix in place as soon as we can identify this.

UPDATE: 5:07 P.M.- We have identified the IP that was causing this attack. We have banned that IP which has resulted in connections returning to normal, which in turn means that sites should be loading properly now. If you are not seeing your site/sites coming online, please submit a ticket via your my.pressable.com control panel. 

If you have any other questions, please feel free to submit a support ticket via your my.pressable.com control panel.

All Systems Operational

We’re happy to report that all systems are back online and operational. We will be coming forth with a detailed explanation of what happened in the coming days, but this is what we can share so far.

  1. This was a coordinated attack on our systems.
  2. This attack used a modified version of the “Slow-Loris” attack against our platform.
  3. Due to the sophistication of this particular attack, it went undetected by the network security team at our provider Rackspace.  It made it look like our infrastructure was being overloaded, when it was not.
  4. We identified this was an attack at 1:00AM on January 24th 2015, by 5:30AM, we had a solution in place that was blocking the majority of the attacks, this is when some customers on “Bode” started noticing their websites working again.

As of 1:30PM January 24th 2015, we have the majority of the attacks blocked, and have pushed the rules to block these attacks throughout our infrastructure.

We are working as fast we can to answer tickets specific to your site, and will keep you posted.

Currently, our systems are reporting at 100%, any issues you may be experiencing now are not related to this outage, and we encourage you to create a  support ticket so we can help you.

Once again, we’re very sorry for this to have happened, we’re working to find out why we were targeted and by whom, but more importantly, we’re working to ensure we are protected against this in the future.

Credits/Refunds/Recompense?..

We will be reaching out to all of our customers who were affected, sometime next week to make this right.  At this current moment, we have some ideas, but our focus is currently on stability and prevention.

Current Outage Breakdown and Full Information

Howdy,

We have been flooded with help desk requests, tweets, emails, and phone calls requesting more information and have had a bit of a difficult time keeping up and getting responses out to everyone and answer all questions.

We are starting a new status blog post to help answer as many of the most frequently asked questions as we can. You can continue to receive the latest updates at the bottom of this post.

For customers who have questions/concerns regarding the outage, please join us in our chat lounge to discuss: hipchat.com/g5gQ8vl9S

Exactly what happened and what is going on?

Yesterday morning, we encountered issues with caching servers that had previously been built and optimized to handle load from the outage early last week. It was determined that this was a result of a bandwidth limitations of our internal network traffic.

As a result, we built new servers that had double the network throughput and continued to have the same issues as before. From here, we decided to subdivide caching traffic based on cluster and saw significant improvement in the situation.

After seeing improvement, we began to see internal bandwidth limitations on database servers and are currently working on adding additional database servers to help with this.

What is being done to fix it?

The first step in addressing the original issue from yesterday morning was to subdivide internal caching traffic based on cluster. Doing has helped significantly and has helped us see bottle necks in other places, most importantly our database servers.

As a result, we are working to add additional database servers as a means of addressing these internal bandwidth limitations as they pertain to databases.

What is the ETA on completing a fix?

Providing an ETA in this situation is very difficult. It relies on us knowing exactly how quickly we can get new servers up, optimized, and working reliably. These times are unknown because it needs to include some time to monitor the implementation and verify improvement.

We are all hands on deck and all working very hard to have stability restored to the system.

We do not currently have and will not likely provide an ETA in this situation. The best thing to do is to keep checking the current status at the bottom of this post.

We expect things to be working normally within a number of hours.

What is the current status?

As of 9:35AM, Jan. 21, 2015: Sites are currently up and down, intermittently. We are currently working to re-provision portions of our architecture but the rate at which we can add servers is currently limited and we are working with our provider to have our rate limits pushed up. Once new servers are up, we will begin seeing sites stay up consistently and running at normal speed.

UPDATE 10:45AM, Jan. 21, 2015: Our provider has raised our rate limits and we are able to provision servers at a more rapid pace. We are still working on getting new database servers up and will update again as soon as we have more information to share.

UPDATE 1:50PM, Jan. 21, 2015: We are currently finalizing the deployment of several new machines and are monitoring these for improvements. We are expecting to see improvements in several clusters as soon as these deploys are finished. We will update again as soon as we have more information to share.

UPDATE 5:10 PM, Jan. 21, 2015: Our team is still working to fine tune the new hardware deployments brought online. We’re continuing to monitor the situation and make adjustments as needed.

UPDATE 8:00 PM, Jan. 21, 2015: Our team is still working to get the new hardware deployments brought online and in rotation.We are expecting to see improvements in several clusters as soon as these deploys are finished. We’re continuing to monitor the situation and make adjustments as needed.

UPDATE 11:30 PM, Jan. 21, 2015: We’re continuing to monitor the situation and make adjustments as needed. Sites are currently up and down, intermittently until the new hardware deployments are brought online. We will update again as soon as we have more information to share.

UPDATE 9:15 AM, Jan. 22, 2015: After discussions with our provider, we are in the middle of rolling out changes that we are hoping will help resolve this in the near future. Thank you for hanging in there with us as we look for a resolution.

UPDATE 10:30 AM: We have isolated a single cluster that was causing trouble for the others. For the time being, all other clusters except that one are up and running “normally.” As we work on that other cluster and test changes/fixes, though, the others may be affected by it. For customers on our bode cluster, we are working on a fix right now  and looking to possibly move customers off this cluster as soon as possible. We are still working on a finalized course of action for these customers.

UPDATE 1:35 PM: We have created a new cluster and have moved a large chunk of customers off of bode and onto a new cluster named “hydra.” If you previously had a site on bode and that has been moved, you will likely see it begin working within the next 1 – 2 hours as DNS changes over. We are finalizing plans for customers that do not have DNS pointed at us and will communicate this as soon as we know what we will be doing with this set of customers.

For customers who have questions/concerns regarding the outage, please join us in our chat lounge to discuss: hipchat.com/g5gQ8vl9S