What is our primary use case?
My main use case for PagerDuty Operations Cloud is integrating with ITOM, incident management, and also using PagerDuty for automation.
A quick specific example of how I use PagerDuty Operations Cloud for incident management is that I commonly use it for critical server down alerts for certain application performance issues, database capacity alerts, network device failures, etc.
In addition to that, for a network device failure when a core switch or a router goes offline, SolarWinds detects the node status is down alert. So whenever an alert is detected, PagerDuty Operations Cloud opens a P1 incident and relevant teams like the NOC team, network team, and incident management team get notified. They automatically get paged through PagerDuty Operations Cloud, and if the incident is not acknowledged, it gets escalated to the managers. This is the automation use case which I created for the network device failure alerts.
The feature I find myself using the most is the best use case from Datadog to PagerDuty Operations Cloud to the Linux team, which I have created. Datadog detects CPU memory or disk or service failure alerts for the Linux server, and PagerDuty Operations Cloud automatically creates the incident and notifies the Linux on-call engineer.
What is most valuable?
According to PagerDuty Operations Cloud, the best features offered are intelligent incident management, AI Ops and alert noise reduction, automation and runbook execution, cloud and hybrid infrastructure monitoring, ChatOps integration, AI-powered operations, service ownerships, and business visibility.
PagerDuty Operations Cloud has positively impacted my organization by reducing mean time to resolution, automatically routing incidents to the correct on-call engineers, reducing alert fatigue, whereas PagerDuty Operations Cloud AI Ops correlates duplicate and related alerts. It also provides 24/7 operational coverage, faster incident management, increased automation, better visibility and accountability, and improved knowledge management.
What needs improvement?
PagerDuty Operations Cloud can be improved by eliminating alert storms, automating common incident resolutions, improving major incident management, using AI for faster troubleshooting, and improving operational metrics.
I could add a feature to automate common incident resolution where engineers perform repetitive actions like restarting servers, clearing disk space, collecting logs, or running diagnostics.
For how long have I used the solution?
I have been using PagerDuty Operations Cloud for almost six plus years.
What do I think about the stability of the solution?
PagerDuty Operations Cloud is stable.
What do I think about the scalability of the solution?
PagerDuty Operations Cloud is highly scalable and is desired for organizations ranging from small operation teams to larger enterprises managing thousands of services, responders, and alerts.
How are customer service and support?
PagerDuty's customer support is generally considered strong from enterprise customers, particularly those running mission-critical operations and requiring 24/7 incident management.
Which solution did I use previously and why did I switch?
Before adopting PagerDuty Operations Cloud, I primarily relied on monitoring platforms such as Datadog and SolarWinds and native cloud monitoring solutions for alert generation.
What was our ROI?
Organizations commonly see ROI from PagerDuty Operations Cloud in both time-saving and downtime reduction, especially when integrated with monitoring platforms and automation workflows.
What's my experience with pricing, setup cost, and licensing?
From my experience evaluating and implementing PagerDuty Operations Cloud for incident management and on-call operation, PagerDuty Operations Cloud offers tiered plans that cost varies based on features such as incident management, AI Ops, event orchestration, automation, and AI capabilities.
Which other solutions did I evaluate?
During the evaluation phase, I considered other incident management platforms such as ServiceNow, Opsgenie, etc., before selecting PagerDuty Operations Cloud.
What other advice do I have?
PagerDuty Operations Cloud helps my team focus on core tasks, which means it helps engineers spend less time on operations overhead and more time on engineering work.
My team has leveraged automation within PagerDuty Operations Cloud by integrating it with monitoring platforms such as Datadog, SolarWinds, and cloud monitoring tools to automate the end-to-end incident management process. Alerts are automatically correlated, routed through the appropriate on-call engineers, escalated when required, and tracked through resolution. The team also leverages automated runbooks and AI-assisted triage to reduce manual intervention during incidents.
PagerDuty Operations Cloud's alert reduction feature has significantly reduced the risk of costly outages in my organization by ensuring critical issues are identified, routed, and escalated to the appropriate responders in real-time. Automation, notifications, intelligent alert grouping, and escalation policies help prevent incidents from being overlooked, reducing downtime and business impact.
I use workflows to standardize the incident response process by automatically engaging the correct support teams, creating collaboration channels, notifying stakeholders, and tracking incidents through resolution. I have also utilized event orchestration to filter, enrich, suppress, and route alerts from monitoring tools such as Datadog and SolarWinds.
Organizations get the highest value from PagerDuty Operations Cloud when they use it for full operation platforms combining incident management, event orchestration, AI Ops, automation, governance, and service ownership.
Regarding governance and security, PagerDuty Operations Cloud provides strong controls that help the cloud operations team maintain compliance, accountability, and operational resilience.
For cloud operations, AI is valuable only if it is accurate, reliable, and actionable. PagerDuty Operations Cloud improves this through an AI-first operation platform, which is built on large-scale operational data, event correlation, and incident response workflows.
I rate this product a 10 out of 10.
Which deployment model are you using for this solution?
Public Cloud
If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?
Google
Disclosure: PeerSpot contacted the reviewer to collect the review and to validate authenticity. The reviewer was referred by the vendor, but the review is not subject to editing or approval by the vendor. The reviewer's company has a business relationship with this vendor other than being a customer: Partner
Nice work Jeremy