I’ll try to edit this page tomorrow when I’m at a computer, but there’s much more information in Derek’s talk at NEXT [1]. They (rightfully) didn’t want to get into a detailed “this is what we saw on <X>”, but Derek alludes to their careful benchmarking across providers.
While you should always assume smart people make economically reasonable decisions, Derek’s point about savings is about list price differences that result from total system performance (and not any sort of special discounting). I’m hoping a follow-up talk will let us say more about how the migration is going, while this talk was focused on the decision to move part of (!) their Hadoop environment to GCP.
Due to our new allegiance to google cloud I’ve been given a little more privileged access to engineers (after the fact) and I can tell that I definitely made the right choice, the people I spoke to favoured having a clean/clear backend with actual quality of life features; but since they don’t go nicely into feature comparison charts people often think that GCP is less mature or featureful.
I’m a real convert, and our company uses all three of the big cloud providers in some fashion- but my team only deals with GCP and we’ve had the least headaches.
What I’m trying to say is: you guys are doing great. I’m really happy with the product for my use-cases.
- >500k cores
- >300PB storage
- >12,500 cluster size
- >1T messages per day
And there's also this other talk, "How Twitter Migrated its On-Prem Analytics to Google Cloud" - focused on their migration to BigQuery:
- https://www.youtube.com/watch?v=sitnQxyejUg
- 20 TB/ day of raw log data, >100k events/sec
- Loading ~1TB/hour into BigQuery.
- Serving 5,000+ complex queries / second. p99 ~300ms
Disclosure: I'm Felipe Hoffa and I work for Google Cloud https://twitter.com/felipehoffa.
> Disclosure: I work at Google Cloud (and directly with Derek and the Twitter team).
Thanks for making this community awesome. I want to work at Google or with Google one day. One of the few companies I have always admired and always will.
It is possible that while GCP gives a good value based on the writing on the tin, someone else might beat that with a good margin under special terms.
That being said, I have reasons to think very few could beat Google on compute or storage game and definitely much harder to beat the big G on Networking side of things.
https://getpolarized.io/2019/01/03/building-cloud-sync-on-go...
I like the google storage API the most honestly. Firebase is really nice but the storage API is very polished and basically does exactly what you need it to do without the complexity of S3.
Obviously that's the thing you'd choose to highlight in a public Google-hosted presentation after a strategic partnership that is a bonanza in marketing for Google Cloud. It wouldn't work so well to tell a room full of engineers "We moved to Google cuz they gave us a fat discount, one you can't get because you're not Twitter!"
The fact that you allocated a top level URL for it is embarrassing (https://cloud.google.com/twitter/). Can you imagine if Amazon advertised https://aws.amazon.com/netflix/?
It is true that any cloud provider would get good publicity by being able to say that Twitter runs on their cloud, but that is precisely the same reason why we shouldn't just shrug this off as a publicity/business decision, because just like Google might have tried to persuade Twitter other cloud providers will most certainly have tried their best too. If Twitter still went with GCP then I think it is only fair to assume that there must have been some real advantage in going with GCP. I don't think anyone who says this must have been a pure business decision is entirely honest with themselves, because a cloud which cannot handle Twitter's volume is no good even if it was entirely for free.
I have nothing to do with Twitter (actually don't even like Twitter because it is such a toxic social media platform) and I don't like many things which Google does, but as an engineer myself I can only say that after using Azure, AWS and GCP for many years in a commercial setting that GCP is indeed years ahead. The quality, speed and reliability of GCP is second to none from my experience and I honestly couldn't say the same about AWS and especially not at all about Azure.
Google's management is what scares me. I'd never build a business around a company with so little follow-through, commitment to product longevity, or focus.
Can you be more specific on what reliability GCP is years ahead of the others? I'm guessing numbers will be pretty comparable on all platforms so if we want to declare outright winners it would be useful to cite a number to indicate where you experienced this.
I hope this is a joke. AS someone who started on GCP, and got tired of the TERRIBLE quality and reliability of their docs among other issues saying that they are ahead is an absolute joke.
I've found some corner cases with AWS, but got quick resolution even on things they have in "beta" - and consistency from documentation through to system is high.
Secondarily, AWS seems to support even old tech FOREVER. I have a simpledb based app. That tech is 9 years old now. Still ticking. When you are building stuff up over time and can't afford to rebuild on the new hotness every other year, this is nice.
Anyone know the revenue AWS and GCP generate? Seriously hard to believe GCP is so many years ahead in this space from my own experience.
"You can transfer up to 100PB per Snowmobile, a 45-foot long rugged sized shipping container, pulled by a semi-trailer truck."
For this use case, we setup 800 Gbps of interconnect with Google.
First you can try to replicate some data and move read instances. After that step by step the rest.
Its really not possible to do that overnight or months even.
Fyi: im currently migrating twice the Twitter size one in the world infrastructure from aws to azure now.
The video linked on the page is also from last August: https://www.youtube.com/watch?v=T1zjmNAuMjs
Throw a 2018 tag on this?
This strikes me as a team (or teams) outgrowing the storage options that were offered internally and choosing to outsource to fit their needs. Isolated use for one business case, and not indicative of a broad movement to the cloud on the part of the whole org.
I could be wrong, of course.
I'd like to know more. And I'd certainly like to know what the engineers at Twitter think.
Relevant technical bits:
> Today, we are excited to announce that we are working with Google Cloud to move cold data storage and our flexible compute Hadoop clusters to Google Cloud Platform. This will enable us to enhance the experience and productivity of our engineering teams working with our data platform.
> Twitter runs multiple large Hadoop clusters that are among the biggest in the world. In fact, our Hadoop file systems host more than 300PB of data across tens of thousands of servers.
Their credit card failed at Google Payments, acc gets marked for fraud and no one in support can help them recover. Twitter is gone for good.
Realistically, it can't be considered as the same service, can it?
Both AWS and Google clouds are perfectly capable of running Twitter, and running any software that Twitter uses, even if the number of machines is different. The only "benchmark" applicable is actually the total negotiated cost of the machines required to get the job done.
So it's not about being able to do things on Google cloud with fewer processors or less RAM or faster hard drives - Google was willing to give them a lower total cost of ownership for reasons known only to Google.
Do not assume that Google (or anybody else) will give you a similar preferential deal. Ignore the "benchmarks".
Hosting your own infrastructure makes less and less sense going forward, the cloud is the new king now.
What was this data, and why did they feel they needed to keep it?
I mean - 300PB is nothing to sneeze at. I can understand wanting to keep important data around (IP, for instance, or recent server logs for the last 30 days or even a year) - but this amount seems to go well beyond that.
In fact, it almost seems like they have stored every single tweet ever written. In addition to probably all of their logs, and who knows what else.
Why are they keeping all of that information? If it is all the tweets ever written - they don't have copyright to it, so such information wouldn't belong to them...
I'm not a big twitter user - can I search for a tweet on their system that I (or someone else) made say...5 years ago or longer? If this is a bunch of tweets - is that why they keep them around?
Are people in general aware of this? That is - if they are tweets - that Twitter has stored all of them, forever and ever, and that they are all still accessible (if not by the general public, then by law enforcement at the very least)?
If they are tweets - do users have any recourse at being able to remove those tweets (aside from those that they likely are legally bound to keep - I am thinking of the current POTUS's tweets, for instance)?
What is the ultimate purpose behind keeping all of that data? Even if it wasn't a bunch of tweets, whatever it is can't really be that useful otherwise, except for maybe the most recent stuff created in the past year, at best. Things like that - logs, etc - might be destined for a rotated "cold storage" and would ultimately "age out". Accounting records (financial data, etc) probably isn't that great a percentage of the data (if at all?) - but whatever was there, too, would likely need to be saved for a greater range - but even there maybe only 10-15 years worth (and it certainly wouldn't even comprise a fraction of the bulk).
So what is it? Why is it being kept? To what purpose? While 300PB isn't that great of an amount in physical terms (that is, the storage medium probably doesn't take up a great amount of space in a datacenter), it still isn't anything tiny, either (especially depending on what medium was used - I mean, I doubt they are storing it all on a ton of 256GB microSD cards - though that does bring up an interesting idea/thought to mind - but I digress).
I'm of the thought that companies shouldn't just "throw away" data - but it should be tempered by "importance to humanity over time". Not everything should be kept - but, for instance, it would have been nice if we (humanity) had kept around the original transmissions from the surface of the moon, or the blueprints to the rocket that took humans there. But we probably don't need all the various interoffice memoranda and fluff at the various NASA contractors (we probably don't need the accounting records - but BOMs might be useful).
That's just a couple of examples I can think of off the top of my head, but history is rife with instances of companies imploding or being bought out or restructured in such a manner that data is just "destroyed" instead of kept. Or lost, or otherwise made unretrievable to humanity.
Here with Twitter - an active company to be sure - we have the seemingly exact opposite - almost like r/datahoarders were in charge.
...and that should concern people a bit, regardless of which company it is - but especially a company like Twitter.