What is our primary use case?
My use case for PagerDuty Operations Cloud is from the SRE and DevOps team. We use PagerDuty Operations Cloud for specific alerting purposes and for the pipeline process. When we build a pipeline and it suddenly fails due to some job and issues, we receive an error. We set up PagerDuty Operations Cloud with our monitoring services, which we are currently using, Datadog. Datadog is connected with PagerDuty Operations Cloud, and whenever Datadog receives an alert or a spike or anything critical, it will trigger an alert to PagerDuty Operations Cloud, and we quickly get a notification. We are currently using this process, and we are also maintaining our on-shift call rotation. For example, on Monday, Wednesday, and Friday, I am working as a shift lead, and then on Tuesday, Saturday, and Sunday, someone else is the shift lead. Regarding MTTR and all those statistics, we can see how many alerts we received, how many alerts we acknowledged this month, and we have a timeline as well. One of the valuable parts of PagerDuty Operations Cloud is that in our team, we can have a competitive environment. For example, if I resolved the most alerts triggered and resolved this month, then someone else can do it next month, and whoever resolves the most critical alerts on time receives appreciation every month.
What is most valuable?
One feature of PagerDuty Operations Cloud that I find valuable is the on-call schedule. We can manage our on-call scheduling, and we have various alert and notification delivery methods available, including mobile. We can receive phone calls, emails, SMS, and push notifications. For example, if someone missed the notification, they will get a phone call, which is very straightforward. We also have incident automation, making collaboration with any third-party monitoring services we use very straightforward, such as Datadog. We can seamlessly automate things with PagerDuty Operations Cloud. The AI features are also beneficial; for example, noisy alerts that trigger regularly and false positive alerts get suppressed. It checks the past month's alerts, showing us that this alert triggered 60 percent, this alert triggered 20 percent, this alert is rare, and this alert is not rare. The escalation policy is excellent as well, as if I did not pick up the call, my manager will get the call; if my manager did not pick up, then his manager gets the call. These are some of the most valuable parts we use in PagerDuty Operations Cloud.
In Datadog, we have multiple dashboards and monitoring systems where we see our spikes and alerts. When we integrated with PagerDuty Operations Cloud, we got better signal and less noise. When we are seeing a spike that is concurrent, in PagerDuty Operations Cloud, the AI feature already signifies that alert as a noisy alert, and it suppresses that alert. This significantly improves our workflow with both Datadog and PagerDuty Operations Cloud. We have faster response and faster escalation. Previously, in Datadog, we did not get notifications, and people would refresh it and check the spike every hour. Now that we integrated PagerDuty Operations Cloud, any alert triggers, and we quickly get a notification or a phone call. Therefore, we do not sit in front of a computer and refresh repeatedly. Additionally, we have a centralized incident workflow; PagerDuty Operations Cloud and Datadog feed into PagerDuty Operations Cloud incident timeline, so we see everything there. We do not need to open Datadog again and again, and if we need to deep dive into an alert from Datadog, we can click the link inside PagerDuty Operations Cloud, redirecting us to the Datadog dashboard where everything is noted down and visible.
In PagerDuty Operations Cloud, AI suppressing our alerts has helped streamline repetitive tasks. For example, very noisy alerts get suppressed automatically, aiding smarter routing. When we have new joiners in our team, they see alerts already suppressed, allowing them to focus on the critical ones instead of the lower ones. Additionally, alert prioritization is present; we receive critical alerts, high alerts, and then low alerts. The faster prioritization facilitated by AI enhances our alert management processes. Also, the root cause historical pattern assists us; if we get an alert similar to one from last month, it tells us how we resolved that alert previously. Historical patterns using AI greatly aid us in alert management.
What needs improvement?
I have already used PagerDuty Operations Cloud, and my previous monitoring tools were very poor for alerting. I had a good impression of PagerDuty Operations Cloud, but I believe it can improve with deeper root cause insights. I know there is automation to detect recent deployments causing incidents, but a deeper root cause analysis could provide more details. If PagerDuty Operations Cloud offers more information, we will not need to jump into the main dashboards where the alert triggered. For instance, if we get more insights directly in PagerDuty Operations Cloud, we would not need to check the Datadog dashboard. Additionally, I think a sandbox mode would be helpful for new team members, allowing us to guide them in simulating alerts, performing escalation policies, and creating PagerDuty Operations Cloud channels.
For how long have I used the solution?
I have been working with PagerDuty Operations Cloud for five years. I worked on two different projects, and in both projects, we use PagerDuty Operations Cloud.
What do I think about the stability of the solution?
In my previous project, we utilized the flexible incident command system to coordinate large-scale incidents, but in my current project with only Datadog, we have not received many alerts or incidents in the last couple of days.
How are customer service and support?
I do not have direct contact with PagerDuty Operations Cloud tech support or customer service teams, but my senior team members have connected with them when we received an alert related to our team failing to set it up properly. The customer support team promptly gave us insight and helped us within 24 hours.
How would you rate customer service and support?
Which solution did I use previously and why did I switch?
I am currently working with PagerDuty Operations Cloud. Previously, on my previous project, we were on BigPanda, but we faced multiple issues during BigPanda. At that time, there was no call schedule feature, and there was no alert triggered feature for BigPanda. We then moved it to PagerDuty Operations Cloud, and suddenly everything was smooth. We got a phone app as well; we set up PagerDuty Operations Cloud on the phone as well. Whenever any alert triggered for us, we used to quickly check from our phone to see if it was a false positive, a true P1, P2 alert, a major alert, or a critical alert. We then quickly jump into the alert and work on it. PagerDuty Operations Cloud changed the process and the flow in our team very smoothly.
How was the initial setup?
I found the initial setup of PagerDuty Operations Cloud straightforward; I did not face any complexities during the setup for alerts or during the initial configuration.
What's my experience with pricing, setup cost, and licensing?
Regarding pricing for PagerDuty Operations Cloud, I am currently a software engineer and a senior software engineer, so I do not handle the pricing aspect. However, I hear from my manager that the pricing is very high for PagerDuty Operations Cloud, and only a few of us have the main business tier accounts. Many of us have low tier accounts that restrict us to acknowledging and viewing alerts, while a few have the ability to create and trigger alerts. Therefore, I do not think much about pricing, but I do believe it is somewhat high. However, I think this is valid because PagerDuty Operations Cloud provides a vast amount of benefits compared to other alerting systems.
Which other solutions did I evaluate?
Regarding the key differences, pros and cons of PagerDuty Operations Cloud compared to competitors, some pros include alert grouping, AI functionality, and the ability to easily integrate with Slack for quicker resolution. Additionally, we receive phone notifications and push notifications, which many of the other competitors do not provide. The pricing of PagerDuty Operations Cloud is also reasonable for the functionalities it offers compared to its competitors. These are some benefits I see in PagerDuty Operations Cloud, including helpful alert insights and direct links to dashboards we have integrated, such as Datadog and Grafana, which allow us to resolve issues quickly.
What other advice do I have?
The recommendation I share, based on my experience with PagerDuty Operations Cloud, is that it is one of the best platforms for synchronizing with your monitoring tools. It will improve your flow, and your team will definitely benefit from PagerDuty Operations Cloud compared to other competitors, as it offers numerous advantages. I give this review a rating of ten out of ten.
Disclosure: PeerSpot contacted the reviewer to collect the review and to validate authenticity. The reviewer was referred by the vendor, but the review is not subject to editing or approval by the vendor.