back
268 comments
The interesting and yet-to-be-explained part is why google flagged the account?

Put all the timestamps you want in the post mortem about what you observed, but you haven't addressed the root cause.

The "this doesn't make sense" part of the story likely has a real explanation that nobody wants to reveal yet.

This exact thing happened to me when I ran https://www.fatherly.com/ circa 2017. Google just shut down our account without notice. We were spending like $10k/month. It also locked us out of our premium support account, so we couldn't even get anyone there to notice that they'd locked us out.

After about 8 hours, a random Google support tech said it was because we were mining bitcoin, which was laughably untrue. We had CPU usage graphs and logs for the whole time and there was no spike. At around 12 hours, they turned it back on, said it was "misconfiguration of our abuse detection" and gave us like $100 in credit.

Absurd. Say what you will about AWS, they would never do that to a customer without a rep reaching out to you first. I have not trusted GCP since.

Shouldn't Google answer this if they are unhappy with this incident report? Are we even sure that Railway knows?
I don't think you're typically told why for these things, and it's mostly automated from what I can tell. The automated systems make mistakes but more importantly they're completely opaque. Nobody, not even Google, knows how they work exactly.
That‘s the point where Google tells you they won’t tell you the exact reason because of security reasons
Who is the "You" in "you haven't addressed the root cause"? If you are asking Railway to spend effort doing this rather than simply moving away from GCP, I'm not sure why they would unless they want to sue GCP to recover damages to brand and long term customer retention.

The moment GCP shut off without any forewarning, its done deal, no need to ask any further questions.

Top comments as usual are buried in deep hate for Google, I doubt that will pressure anyone at Railway to address this.
Railway has not had the best month in the tech press have they? And in both cases it was an automated process belonging to some other party that put them there, damaging their reputation.

I was going to talk to our google rep about their killing the Gemini cli but this is way more concerning.

In the case of them giving AI admin credentials to delete their production database, and it deleted their production database: that's on them. They were the only ones who put the admin account credentials into their AI.

Then they took no personal responsibility. That definitely damaged their reputation. Here, they are taking at least some responsibility. Props to them on improving.

Also, GCP does indeed have serious reliability issues, and Google does indeed have serious customer support issues.

EDIT: It has been brought to my attention below that the first 2 paragraphs are misattributed, and were not Railway, but rather a customer of theirs. Sorry, Railway!

Building on someone else's platform is always gonna be a risky move, and building a platform on top of someone else's platform is even riskier.

My company used to use a hosting provider that was basically AWS plus some extra guarantees. We just finished migrating onto regular AWS because they now offer what we need directly.

Unfortunately we had to make emergency migration off to Azure yesterday due to this. Thankfully our DB was not hosted on Railway and we were back up in a couple hours.

As much as we loved the simplicity they provided us, there's just been too many mishaps and shortcomings for us to continue running a B2B enterprise app on their infrastructure.

Sad day :(

Question: for a smaller SaaS tool, or even internal product. If a team doesn't want to manage AWS or another IaaS provider, what are the best alternatives for the following

1.) Vercel - having a bad month

2.) Supabase - having a bad month

3.) Railway - now having a bad month

> May 19, 22:10 UTC - Our automated monitoring detected API health check failures and paged our on-calls, who started investigating the issue.

> At 22:20 UTC on May 19, Google Cloud placed Railway’s production account into a suspended status incorrectly, as part of an automated action.

If the timestamps are accurate, what was causing the errors 10 minutes before the account was suspended?

The simplest explanation is just that one or the other of these timestamps is wrong, which wouldn't be a big deal. But if the timestamps aren't known with certainty, it seems very odd to include them in the writeup as though they are certain, even though they are very obviously inconsistent with each other.

This should be a warning to anyone running GCP. They suspend accounts left right and centre without even thinking about what they're doing. It seems like they use Gemini 3.1 Pro to run their production decisions.

TK has a history of absolutely destroying the culture of the place like in OCI and has done something similar in GCP from what I've heard. GCP and Google are completely different entities with how they work. Don't expect Google quality from the name. It's just like those old brands which now have cheap licensed products like Nokia (An exaggeration I know but not far from truth).

Not only that they are known to shut off their services randomly giving you like 6 months to migrate. They have lots of engineers not doing anything, so they put them on migrating internal users off those services, most of their clients don't. There was a brilliant article on this by an ex-GCP employee that I can't find right now.

Avoid GCP like plague if you are serious about your business.

Edit: Gemini (unironically) found the article on this, a very good read: https://steve-yegge.medium.com/dear-google-cloud-your-deprec...

This isn’t the first time Google Cloud has seriously messed with a customer’s account: https://cloud.google.com/blog/products/infrastructure/detail...
"Railway owns our vendor choices, and we ultimately own this one. Your customers don't care whether the failure was Google or Railway; they see your product. Your uptime is our responsibility, and we'll keep delivering on it."

Kudos to them for acknowledging it and not doing PR speak. It shows it was an architectural failure from their part of trusting GCP, and they are working to fix it. Should they have seen it coming? Yes. But better late than never.

"Finally, we are in planning to remove Google Cloud services from our data plane’s hot path, and keeping them only for secondary/failover."

That's pretty clear. Google can no longer be trusted as a B2B service provider.

What drives Google to apply these actions so completely and immediately, versus a more deliberate approach, with notification and delay before action, manual review for paying customers, or a warning to resolve within X hours/days? Once or twice could be errors or bad implementation, but these can't explain away the pattern.

It would seem that Google's counsel has deemed that whenever _____ is detected, the company must immediately and completely sever the business relationship. What is that driving concern? Is it sanctions enforcement? CSAM? Something else?

> As a side effect, Terms-of-service acceptance records were also reset, prompting users to re-accept on their next visit to the dashboard.

Don't get me wrong- the rest of this mess falls pretty clearly on Google Cloud, but this one feels like something Railway did to themselves.

It's reassuring to know they will ban a million dollar enterprise customer just like they will ban your GMail of 20 years.
I've read all the threads and their main page and I still don't really understand what this service is. Is this like a commercial alternative to Gerrit? What do people use this for?

I'm not a developer, just curious what this is.

"Your customers don't care whether the failure was Google or Railway; they see your product. Your uptime is our responsibility, and we'll keep delivering on it." - Thanks Claude!
It's highly unlikely that GCP banning their account without telling them is true, but GCP is probably not going to go public with the real reason.
Sadly, My Railway project is still having issues 24 hours later. Already started emergency migration away from Railway backend :(
The RCA and preventive measures was a pleasant read. I got a lot of respect for companies putting a lot of effort into incident reports like these. Makes them appear very professional rather than just blaming the cloud provider outright.
Had similar experience with GCP. Terminated VMs six times, and responded zero times.
So, what was the reason for the account suspension. Why did it happen? I know Google can be a bit stupid with their automatons but I am bit skeptical here. There are sites more critical than Railway hosted on GCP.
How many trains were delayed or incorrectly routed as a result?
Even if it ultimately turns out to be "Google's fault" (as this report seems to be saying), Railway say they own the incident but make no apology here.
Unfortunately we've also had a litany of problems with our GCP deployment and chose to remove them completely as a service provider.
> Railway’s production account into a suspended status incorrectly, as part of an automated action.

Be it individuals or companies, this time is the best time to ditch all dependence on anything clouds or SaaS since all are using automated AI, more and more of these incidents will occur.

All my services have lost their databases. All their AI agent is offering me is to delete all the data (thousands of data) and create a new database. It insists there's no other way. But I have clients, I have hundreds of courses on the platform!
Google has a culture problem. This is not something that can change easily nor will it change when it’s not recognized as being an issue within their organization.

Between my peer c-suites, the conversation is that GCP cannot even be in the consideration set until such a time as a several-year period has elapsed without this kind of incident.

back to on-prem
>Google Cloud placed Railway’s production account into a suspended status incorrectly, as part of an automated action.

There is no justification given on why this action was incorrect. It's possible they actually did something wrong.

Nov 28, 2012, user fbuilesv posted: Google Compute Engine is on Limited Preview right now. If you're planning to offer a service that you care about you should consider this.

Do we know if GCP has ever left limited preview..??!

Google, the new Microsoft!
I will definitely not be signing up on GCP because of this.
Related discussion during the incident:

https://news.ycombinator.com/item?id=48201484

stories like this are why i self-host most things on proxmox instead of depending on a single cloud provider. Ok I have to do the maintenance but at least this way no one can suspend my entire stack with not even an explanation.
It's not my proudest moment, but at least being banned and suspended all the time brings some wisdom.
19 minutes from detection to getting the google account restored is pretty awesome honestly.
They forgot to get reimbursement for downtime. A free month of GCP is better than nothing.
Flagged by some AI automation.
Now given the logic that you can't be dependent on any one service to run your SaaS, how does Railway convince its customers to run their SaaS on a single service?
Major infra provider -> has no backups/game plan if GCP goes down