back
156 comments
> Goddamn AWS seems to be down again 4th time this month alone

I feel obligated to point out that this is a very editorialized title, which is against site guidelines. Not... that I don't sympathize, just that it's perhaps a bit out of line even for providing context.

It may be editorialized but it's also exactly the words I would expect to come out of the mouths of some of my most experienced, senior, and clueful colleagues when asked to describe the present reliability of AWS US-EAST-1
True, but it's funny... perhaps next time a "Tell HN" with the link?
I think it would have been fine if it had Tell HN in front of it. Without it the title would be against the rules and guidelines. One could argue the word Goddamn seems to be provocative. Which could be appropriate if it was really down. But it is not. I think @dang need to edit the title.
It certainly could be editorialized. Alternatively, AWS may be in fact damned by god.
I'm honestly sick of seeing these posts. Every single post I've seen on HN, I have had no down time in EC2, S3, Workspaces, FSx for Windows, Directory Service, Console, etc.

Over 200 services, 86 availability zones and 26 regions. We might as well post every 5 minutes if we post every time something on AWS goes down. And yes AWS is more dependent on us-east-1 but one of the outrages posts was for a minor outage in us-west-1.

nice try, Jeff ...
You felt obligated to the benefit of who, exactly?
I got worried when I saw this then checked https://status.aws.amazon.com/ and all seems good, I'm relieved ;)
I hope you are kidding because Amazon has been very late with updating that page in the past despite problems being obvious.
For once the AWS status page is actually telling the truth, there is no outage.
There is no AWS outage, just routing issues with shit asian ISPs affecting connectivity to many providers. Amazon has nothing whatsoever to do with this.

Imagine Comcast or Verizon fucking up their configs, you might not be able to access AWS stuff but that doesn’t mean that AWS is down.

But hey, the real story wouldn’t make the front page, so lets just stick to the fabricated narrative.

> 10:41 AM PST Between 8:59 AM and 9:32 AM PST and between 9:40 AM and 10:16 AM PST we observed Internet connectivity issues with a network provider outside of our network in the AP-SOUTH-1 Region. This impacted Internet connectivity from some customer networks to the AP-SOUTH-1 Region. Connectivity between EC2 instances and other AWS services within the Region was not impacted by this event. The issue has been resolved and we continue to work with the external provider to ensure it does not reoccur.

https://status.aws.amazon.com/#AP_block

AS6453 is having a mega outage according to https://www.thousandeyes.com/outages/. Lot of services for users in India were inaccessible for ~1h or so.
So not AWS, but internet outage in general in India?
Some cloudfare seem to be down too?
These threads always remind me of Umberto Eco's essay Sports Chatter. He describes the degress of participation in sport:

1. Playing the game yourself.

2. Watching other people play.

3. Talking about people playing the game.

4. Listening to someone else talk about people playing the game (i.e. sports commentators).

5. Talking about what sports commentators have to say.

In sports and in threads like this we're mostly at 3, 4, or 5. Even people who work at AWS posting here may only be "watching," i.e. working with second-hand information. While it's entertaining to nerds to speculate about root causes and possible solutions, and the dire predicted fallout from problems (especially when they happen to companies they dislike), it's just chatter and doesn't actually address any real issue, or inform any decisions that mean anything.

It’s Christmas Eve, I aint answering any PagerDuty notification today or tomorrow.
It takes years to build trust with your customers, and very little to disrupt it. I think these recent events will have people rethink their cloud strategy. This could be a good opportunity for Google to take on Amazon.
I would be curious to hear what any current or former AWS employees think the internal consequences become for this.

From the outside, it feels like they are going to have to do something different to get some customer confidence back. Some sort of "mea culpa" with an explanation of what they are going to change.

Complete speculation; but I wonder if how Amazon works their employee's so hard is the root cause of this. Layer on COVID as an additional stressor and give it a few years and this could be the result. Curious to hear from those at Amazon, what's it been like during COVID?
…or there are Internet issues that are causing issues for people. Do we actually know AWS is having issues or people are speculating?

Not seeing any issues here, but am seeing people reporting broader internet issues at the moment. Post title seems a bit quick on the trigger to point blame.

Why isn't there any investigative journalism looking into wtf is happening at AWS? What the point of all these tech journalists writing all those puff pieces for access if they're not going to use it to pierce the NDA shield at times like this?
This is getting to be a massive joke. Any reputation AWS has for being robust has been wiped out over the past month or so. I feel bad for the various engineers at AWS and other companies who are having to work on their day off.
I've always heard about technical debt at AWS, with thousand line functions everyone is terrified to touch, demotivated employees doing whatever it takes to close tickets without dealing with the root causes, and so on.

I presume that this rash of downtime is just the rickety structure inevitably creaking and breaking. It's probably too late now to fix things — dealing with legacy code requires patience, discipline, and understanding which the management of Amazon doesn't have. They'll just yell at people louder, and hold people "accountable" by punishing anyone where anything breaks.

So this is the 4th time an uninteresting update like this has come to the front page... Why? What can be gathered from it? Get the AWS post-mortem on here, for that I'll be piqued. HN is not a status page. I don't instinctively check HN when my service is down because why would I?
looks like it's a connectivity outage across all services - https://www.thousandeyes.com/outages/
Amazon’s SLAs are quickly becoming only useful for kindling.

At least your fireplace will give some warmth for your relatives. Merry Christmas!

I have been wondering if Charlie Bell leaving AWS has anything to do with these recent outages. His ops meetings were fucking brutal. I’m curious how they are since he left.
Anyone deploying to production Christmas eve, or even the week of Christmas, should have to get direct approval from the CTO of the company to do so. Just a terrible, terrible idea.
According to ThousandEyes, it looks like it was AS16509 that was affected. The outage lasted for 28 min; however, it didn't seem to affect many applications for too long.

I think it's because there was a backup server that kicked in for AS16509; however, it also went down but only for 8 min.

https://www.thousandeyes.com/outages/

I don't know about "4th time this month alone" which seems a bit editorialised, but our company is off AWS for now, and most likely for good.

We're not in the cloud or tech business per se, and as such our customers are not really understanding of technical issues which unfortunately means they are blaming us, and our own reputation is on the line because of AWS' shortcomings.

We did consider Google but for now OVH is the only major provider which is both reliable and secure (w.r.t. court-and-gag letters from government and intelligence agencies) as far as we're concerned. We still use Hetzner and Scaleway for some older stuff and also because of Scaleway's low prices, but it's likely we're moving everything to OVH in the future.

That's because they're running on-premise.
Wonder how much of this is caused by engineers rushing to fix log4j issues at AWS so they can enjoy the holidays? Their stack is heavily Java-based isnt it?
Happens when you delegate, don't give them money.
There's a video that's being going around HN lately about civilization collapsing and the loss of the ability to maintain robust systems.

Now that I'm looking for it I see it everywhere.

Pro-Tip: Get the hell out of US-East-1.
212 cases doesn't seem to be a lot. Does anyone know if it's legit? If they are burning left and right the 9s are going to leave soon.
It's almost as though it were a bad idea to put all the eggs in one basket
meanwhile my raspberry pi cluster at home on residential FTTP has 100% uptime over the last 2 years

(not entirely serious)

This looks to be an ISP issue and not AWS.
what the hell this is unacceptable. is this some sort of sabotage? AWS may have to live this down for years in RFPs
I never ever want to hear people disparage DigitalOcean or Linode ever again. The "cloud" has now been objectively proven to be just as flakey.