The cloud is a murky, ambiguity-laden concept though. Both Netflix and my 92 year old grandmother on Facebook 'use the cloud', but the former is much more sophisticated in their network and data management practices. My grandmother just wants to see fun pictures of her family and great grandkids.
No one realizes that the Cloud is running on the same crap we've always had and is vulnerable to the same issues as everything else. MAYBE the company is better at data management, MAYBE the employees take pride in their job and do it properly, but that's all MAYBE MAYBE MAYBE, and could just as well be "no" and you're entrusting your data to people who really have nothing to lose if it goes into the garbage tomorrow.
I think it is sort of like a delivery system. If USPS or FedEx or USP started losing a massive number of packages (I know that they do lose some) then they would get abandoned, just like "Cloud" companies have an incentive to maintain a baseline level of quality. The alternative is that every business would have to create their own shipping services. I think it makes sense to assume that in most cases, unless the business is already massive enough to warrant it, that it is cheaper and more reliable to use the aggregate, dedicated ones for hire. The "Cloud" will be cheaper than individual implementations, and it won't be nearly as suspect to individual implementation errors because the identical system will have been proven by many other customers (otherwise it would be abandoned).
If amazon loses more than 3 datacenters (only total loss of external connectivity for all of your instances in an entire availability zone, or total loss of hard disk access, again only counts if all your instances completely lose hard disk/EBS access) for more than 45 minutes in a month you get 10% of what you pay as a voucher for future ec2 usage. If they lose it for more than 7 hours you get 30%.
So no, Amazon, or at least their legal department, does not trust their own competency. Or at least, they're not willing to risk any revenue on that, but they're willing to give you a small future discount to encourage you to restart using the service. Oh and you only get that if you explicitly ask for it.
If they lose your data on EBS/S3/Dynamo/..., you get nothing. So having any data exclusively on any Amazon service should be cause for getting fired, and this of course also means that using Dynamo for storing anything non-trivial is a big no-no from a disaster recovery standpoint.
So I have to say, I would suggest you do not trust Amazon with either your data, nor with keeping your site online. Yes, historically their performance has been better than this, but ...
This reads worse than the SLAs on internet connectivity from places like level3 and cogent (pay 10% less if they fuck up completely for more than 2 days).
Second, you've completely ignored the actual key points. For one thing, "Cloud" companies make their business by providing a stable service. You have the "guarantee" based on thousands of other business using the exact same infrastructure without serious service failures. That is a huge amount of statistical reliability. Compared to hiring your own IT department and cobbling together your own system, that is actually really good indicator. Second, the cost difference is potentially massive. Again, it is for similar reasons that shipping via UPS is a much better deal than shipping via your own private distribution network. You might have to still pay some people to handle your own inventory from its source (like you'd have to have some people to work on your system in the cloud) but you'd be taking advantage of a much larger, more efficient system instead of having to build and maintain your own.
So if it's all the same, I'd rather have a decent SLA. Furthermore this sounds a lot like Amazon's not in fact giving me anything.
Your point is that they'll do the right thing because otherwise their customers would leave. Customers you said in the previous paragraph they give "the most conservative amount they can get away with, and they are "getting away" with it just fine so why up it?".
Sounds like they really care about customers doesn't it ?
> You have the "guarantee" based on thousands of other business using the exact same infrastructure without serious service failures.
I can get that guarantee at 1000 datacenters and colo providers, at least. Some of which have a decent SLA. But even among cloud IAAS, both Azure and Google provide both Amazon's guarantee, and better SLAs.
I used Amazon as an example. If you actually read what I'm saying, about how Cloud businesses in general depend on meeting their guarantees and not screwing over businesses, how their superior quality is because of scale and specialization, how they are reliable because one failure would doom them and they haven't failed yet, you could see that this has nothing to do with Amazon at all.
You keep arguing that Amazon is a bad provider. So what? I was never interested in that at all. I'm not comparing them to Azure or Google or the supposed "1000 datacenters and colo providers" you seem to know of. I don't care who is better or worse, I was talking about using cloud services in general.
Pay attention to the topic, pay attention to what my arguments were. Amazon's SLA is utterly irrelevant to anything, what are you even trying to convince me of? None of anything you've said is remotely relevant to my point. It's like arguing about whether Ford or Toyota makes better hybrids in a discussion about whether electric cars are a good idea, I just don't care.
Amazon has superior quality ? They have at best average quality as a vps provider, unless you accept their products that cause lock-in. At which point you're at their mercy, and they have even less reason to treat you well. Amazon doesn't match, say, digital ocean (especially not in the transparency in billing department. WTF). There are other reasons to pick amazon of course, but quality, not one of them. Price ... not one of them. Service ? Not one of them. Stability ? Not one of them. Geographical reach ? At the moment Amazon does better (not that it matters unless you're in Asia).
One failure would doom them ? Just from memory I know two big amazon cloud failures that you could not protect from with availability zones, the ones in a single datacenter, they don't even publish.
The fact that they refuse to publish single cluster failures is probably another aspect of that superior quality you mentioned.
Also, you can get fucked on an ongoing basis just by getting scheduled on a machine. I guess that's part of their superior quality (a lot of VPS providers of course have this problem, others are better at it).
The Netflix/Amazon relationship reminds of the "you owe the bank $100, the bank is your problem, you owe the bank $1bil, you are the bank's problem" sentiment. Netflix is probably such a big business that they are dependent on each other.
On the other hand, Amazon seriously screwing a small business would be like a bank failing a normal customer's withdrawal from their deposit. The second that information went public, the bank would essentially be dead.
http://www.theregister.co.uk/2015/09/20/aws_database_outage/
AWS is the new Microsoft / IBM, nobody ever got fired for picking AWS.
There is a difference in a discussion about trusting the Cloud with your data and services between it going down briefly on occasion (somewhat acceptable, within very narrow limits) and actually losing data or longterm traffic because of a service failure. Seriously breaching the SLA causes compensation as well as a big loss of reputation and business, going down for a couple hours once a year is hardly the type of instability that would terrify most online businesses, nor is it something that individual companies are able to avoid themselves.
The key piece that's missing here is the idea that risk is something you have to compare, and can combine in interesting ways, then trade off against costs.
There are a bunch of ways in which you can do compute, storage, and networking. You get to pick zero or more of these ways. One of them is "buy a bunch of iron and make a pile of it in your bedroom". Major risk factors here are your house burning down, you getting evicted, or there being a power cut. Another is "rent those services from an infrastructure provider". Risk factors here are much harder for you to visualise, but include things like "governments ban that company from operating in your country".
You can look at the risk of any of these options, and quantify it with an SLO, like "we intend for this compute resource to be available 99.99% of the time in a given quarter". You can then have an SLA that defines what will happen if that objective is not met, and measure how often this is complied with over time. There are lots of ways to analyse this information, but let's suppose that you can reduce it to a single number measuring how safe the resource is for your use case.
If you only look at a single option, and say "this has a safety of X", then the only thing you can get out of this effort is anxiety. This only becomes interesting when you start looking at differences between alternatives, like "the safety of servers in my bedroom is X, but the safety of buying resources on GCE is Y, so I can get this much of an improvement by spending that amount of money", or "by doing both of these things I improve my safety to Z, and I am willing to pay the additional cost of doing so". Or perhaps your position would be "this option is less safe but much cheaper and I'm willing to accept the extra risk".
The problem I have with the "fuck the cloud" article is that it doesn't do any of this. All it says is "the safety of this option is only X, you should experience anxiety". Is X higher or lower than that pile of iron in your bedroom? You still don't know.
(Realistically, unless you have the ability to build a system in your bedroom that has continental diversity for storage, N+2 of everything for hardware failure, etc, your bedroom is likely to be far less safe than the major cloud services - unless you live in a country which regularly bans American companies from doing business with you, which a sixth of the world's population does.)
There exists just about every price point and combination of services you can imagine. It's awesome, too, because I can grow a business and can enter into the market with a whole "rack" of servers for almost no cost at all.