Category Archives: Uncategorized

RESOLVED: Email Outage Across All Systems

On December 30th, 2013 there was a system wide email outage that impacted delivery of email from all of our systems. This includes customer WordPress sites, notification emails from My Pressable, support requests submitted through My Pressable, and email inquiries through our own website.

What Happened: On December 27th we were notified through some older (zippykid.com) channels of an issue with our account at our provider. On December 30th there was a miscommunication with a configuration change and our account was temporarily suspended. This suspension resulted in the immediate halting of all email sent from our systems.

We were able to identify and correlate the issue through several customer issues. Once we determined the issue was with the changes at our provider we worked with them to restore services. At approximately 5:00pm CST our account was restored, and messages began to queue for delivery. However, complete services were not restored until 9:00pm CST.

What We’re Doing: We’ve made some adjustments to the account control for this service provider so that more people are aware of changes made. We’ve also evaluated fallback options should we need to implement a new solution in the future.

While this was a unique issue, we do apologize for the scope and time of the outage. We’re continuing to see what else we may be able to do to improve the reliability of this in the future as well as prevent any further issues.

Please also keep in mind, that if a request for support was submitted through the My Pressable control panel during this time, we would NOT have received that request. We apologize for that, but please re-submit your issue if you still require assistance.

If you have any questions, please feel free to contact our support team.

Scheduled Maintenance – February 22nd 2013 – 0000 hours

Hi Everyone, 

 Just wanted to let you know that we need to take our systems off line for one hour on Friday, February 22nd at midnight CST. The nature of this maintenance is complicated, as it puts the finishing touches on the upgrades that have been under way for the past 8 months.  Unfortunately, this means connectivity to your website will be affected. 

All websites will be unavailable during this time period. We suggest you suppress your monitors during this time, as there will be nothing we will be able to do about it. 

Connectivity will be restored by 1:30AM CST or 0130 CST hours at the latest. 

Caching layer degradation – Fixed and Stable

We’ve found some issues with our memcached cluster, which we’re working to resolve as soon as possible. Symptoms of this are slower sites, and sometimes pages that will return a “504 Timeout” page. 

Sorry for the issues, we’re working to resolve these issues ASAP. 

 

Update February 12, 2013 8:46 PM: The problem was fixed at 7:30PM CST, we’ve been monitoring the situation for the past hour, and things have been stable. We consider the issue resolved. 

Database and Service Interruptions

We’re currently aware of an issue inside our database cluster that’s causing some slowness/unavailable sites. We’re currently working on the issue and will update you with more details as available.

 

UPDATE 10:45AM CST: The database connection issues are still ongoing and we’re continuing to investigate the cause.

Update on the botnet attack of February 7, 2013

We’re starting to get things under control. We’ve blocked 2832 unique ip addressess so far. We’re continuing to monitor the situation and isolate the customers who were affected by this, from the customer who was being attacked. 

What we know so far

  1. A customer’s website is under a botnet attack, where we are seeing 190,000 requests/second made to one ip address. 
  2. These requests seem to be coming from about 3000 unique ip addresses.
  3. Our firewall was reaching a CPU max of about 90% while this was happening, our alarms go off when it hits 51%. 
  4. Blocking all 3000 ips on the firewall is not a good idea, so we’ve “null routed” the destination ip address. 

What are we doing to bring customers back?

  1. We are assigning new ips to the affected customers (several hundred) who shared the same ip address with this customer.
  2. If we control your dns, this change will happen within the next 30 minutes. If we don’t, we’ll be contacting you to let you know what the ip address should be. 

 

January 27, 2013 all systems functioning normally again

As of 2:53 PM on January 27, 2013 All systems are functioning normally again. We had intermittent issues across our network.

Here’s what happened. 

One of our 4 memcached servers had run out of memory, and in the process locked up. This made it so that our database servers were seeing 8x the average calls. Since our monitors started telling us about higher than normal database activity, we started investigating the issue there.  

What did we learn?

It turns out, our monitoring on the memcached systems isn’t as good as we thought it was. Had we known that the one of the memcached server was out of commission, we would’ve been able to identify the problem, and fix it. Rather than investigating what was causing the spike in the database usage.