> Excluding scheduled maintenance windows, Datadog will use commercially reasonable efforts to maintain 99.8% availability of the hosted portion of the Service for each calendar month during the term of this Agreement. The Service will be deemed “available” so long as Authorized Users are able to login to the Service interface and access monitoring data. Excluding planned maintenance periods, in the event the Service availability drops below 99.8% for two consecutive months, Customer may terminate the Service in the calendar month following such two-month period upon written notice to Datadog. To assess uptime, Customer may, if under a Paying Plan, request the Service availability for a prior month by filing a support ticket through the Site.
Doesn't seem like that SLA could be defined as a "low bar" to me honestly, 99.8% in writing is impressive. It's public as well, meaning if you need a better SLA they aren't the ones for you.
New Relic is 98.5% https://docs.newrelic.com/docs/licenses/license-information/...
The first is that it doesn't cover key platform features. I don't see anything about error rates on metric ingestion or error rates/timing on sending out alerts. Being able to log in and look at metrics is like 4th or 5th on my list of things I care about. It also doesn't preclude a severely degraded service being considered up (e.g. a 25% error rate, but refreshing enough times will get it to load). DynaTrace, by comparison, does count the service as unavailable if it's unable to receive any inbound data.
The second is that their SLA doesn't give out credits, it just allows you to cancel your contract in the calendar month following 2 months of not hitting their SLA. In other words, using their SLA means finding a new provider and migrating within ~30 days. It also means there's no real penalty to them for violating their SLA, since customers upset about the uptime would just not renew their contract. This just lets that happen at an accelerated rate.
99.8% is also not that high of an SLA (especially with what it covers). That's ~1.5 hours of downtime per month, which I would consider pretty average or even mediocre. That's almost a half hour outage per week. To me, 99.9% is good (~45 minutes/month) and 99.99% is impressive (~4 minutes/month).
- It needs to be missed 2 consecutive months before it applies
- You can't see the uptime, have to submit support tickets to get it
- And then you only get to cancel a bit earlier (after 2 months of fuckups), not even a service-credit or refund
It's a completely useless SLA