Production Support Engineer - Database at a wholesaler/distributor with 201-500 employees
Real User
Top 10
Jun 23, 2026
My usual use cases with PagerDuty Operations Cloud involve handling incidents through a full flow. When there is an outage, an incident is created that can be either severity one or severity two. As the person on call that day, I receive a page from PagerDuty on the app and three calls on my cell phone. When I pick up the call, PagerDuty IVR asks me to acknowledge the incident. Once I acknowledge the incident from the call, I go to PagerDuty through the website, which is much easier to navigate than the mobile app. I then page other teams responsible for the incident, as well as the stakeholders and product owners. PagerDuty is integrated with Microsoft Teams, so I open a Teams bridge call to resolve the issue and update all incident details in PagerDuty notes, which automatically integrates with ServiceNow incident and sends messages to stakeholders' phone numbers. One time while hanging out in the mountains, there was no internet signal but there was cell reception. An incident happened while I was on call that day. Normally, without internet, I would not be able to know about it, but because of PagerDuty, I was paged three times on my cell phone as well as through text message. I managed to call another colleague from my work and told him to take care of the incident. This helped me avoid breaching the SLAs on incident acknowledgment and allowed me to access remote incidents without relying solely on the internet.
My main use case for PagerDuty Operations Cloud is to manage critical alerts and incidents across our production systems. It helps our team route alerts to the right people, manage on-call schedules, coordinate responses, and reduce downtimes. We also use its integrations with our monitoring and collaboration tools, so issues are identified and addressed quickly before they impact our customers.
Operations Lead at a tech vendor with 10,001+ employees
Real User
Top 10
Jun 12, 2026
I have been using PagerDuty for the last nine years, but PagerDuty Operations Cloud for over one and a half years. We work directly with merchants and need to trigger immediate alerts whenever there are 5xx errors or business errors like 4xx issues, as well as payment failures. We have configured every alert on a data log in some other monitoring tools that are integrated with PagerDuty. We receive alerts very immediately and trigger calls and Slack notifications. We integrate everything with PagerDuty and get notifications instantly, after which we start our triage process. One use case I can mention is when we have an auth rate dip. Whenever there is an auth rate dip, we run into revenue losses with the merchants or partners that PayPal currently works with. Since everything is integrated, PagerDuty Operations Cloud catches when there is an auth rate dip for particular merchants and immediately triggers a notification for us. We then immediately dive into what the problem is and figure out how to fix the issue with the help of engineering teams.
Site Reliability Engineer at a tech vendor with 10,001+ employees
Real User
Top 20
Jun 12, 2026
I primarily use PagerDuty Operations Cloud for alert management and incident call rotations. In my earlier firm, we managed rotation shifts across three time zones: EMEA, APAC, and New York time. All rotation and shift management was handled through PagerDuty Operations Cloud schedules. Application monitoring was also updated through PagerDuty Operations Cloud. According to the schedule, we updated people's contact information so that in case of any issues, the contact would be transferred to the respective shift member. We also managed escalations with five layers of escalation. If a first team member missed an alert, it would go to the second team member after 10 minutes, then to the next person after five minutes, continuing according to the priority of the service.
I usually use PagerDuty Operations Cloud for the notification of high-priority incidents within the infrastructure. I also use it for escalating to the on-call members, scheduling the priority of incidents or issues within the infrastructure, and creating scheduled rotations for team members.
As a cloud operation team, I was a user who set the alerts, and whatever important incidents or anomalies were detected that needed to be immediately taken care of were bifurcated through our APM tools that we integrated with PagerDuty Operations Cloud. As a cloud operation team, we supported the platform for rotational shifts. My roles involved setting the person in the shift according to the shift roster, so whenever any incidents triggered, they would get the call. The primary use was supporting production operations and cloud activities. Our multi-environment consists of AWS infrastructure, Linux servers, Kubernetes clusters, and customer-facing applications. PagerDuty Operations Cloud was mainly used for incident management and alerting. We integrated it with AppDynamics, Instana, and CloudWatch, where it would monitor the patterns and platform, and then PagerDuty Operations Cloud would generate the critical alerts that the appropriate support team who was working in that present shift would get notified of immediately. This platform really helped us manage production incidents beyond service outages, mostly high CPU utilization where we set alerts, application failures, pod issues in Kubernetes, and infrastructure-related alerts. We configured all kinds of alerts, which ensured that alerts were routed to the correct on-call person, helping us reduce response time in critical situations.
My main use case for PagerDuty Operations Cloud involves working in the AIOps team, which is an operations team. We have a monitoring tool called Checkmk, and we have integrated it with PagerDuty for incident management. We monitor many servers across different teams, including the Linux team, network team, Windows teams, and database team. All of these servers are monitored in Checkmk, tracking live CPU, memory, and file systems. Upon reaching certain thresholds, Checkmk generates events, which we integrate into PagerDuty Operations Cloud console. There, we set conditions so that if an event is critical or a warning, it converts into an incident. We then route the incidents to respective teams, who handle the details. A specific example of an incident where PagerDuty Operations Cloud played a key role involves automations we created within PagerDuty. Rundeck is a job workflow tool where we can implement scripts or schedule jobs. If a server meets its threshold, it triggers PagerDuty Operations Cloud. We create scripts in Rundeck to handle issues, such as clearing a full file system. We utilize a feature called Automation Actions in PagerDuty Operations Cloud, and whenever an incident comes that matches specific conditions, that job will automatically run in Rundeck. This incident management cycle is effectively managed in PagerDuty Operations Cloud, allowing jobs to run and resolve incidents automatically, ensuring the server is healthy again.
My main use case for PagerDuty Operations Cloud is monitoring multiple platforms. In cloud operations, whenever any application or device has an issue, it triggers PagerDuty, and the on-call shift engineer can immediately check on it. For a specific example of how I have used PagerDuty Operations Cloud to handle an incident, if any media connect flow breaks or applications deployed on cloud instances have issues, we start receiving alerts. This can be configured using PagerDuty schedules to inform the on-call engineer to take immediate action. Regarding my main use case, I would add that we do not need to be on the dashboards or monitor manually. Instead, our APIs work in the back-end. They check for things working on the platform's cloud, and if any action is required, the appropriate team can be aligned using PagerDuty.
The main purpose of PagerDuty Operations Cloud is to receive alarms for real incident production critical issues. Whenever an incident happens, we get an alarm call or a phone call on our phone or through an application call, so that we are aware of the situation and know that we have to be vigilant about it and resolve that particular issue. This is the main use case of PagerDuty Operations Cloud, and it is the use case where most of the company is using it.
Senior Network Operations Center Engineer at Accolite
Real User
Top 20
May 19, 2026
I use PagerDuty Operations Cloud for core tasks only because my client completely belongs to healthcare. Healthcare is a high priority where we need to help people. I primarily use it for core tasks and sometimes for primary tasks as well.
Principal Incident Commander at a tech vendor with 5,001-10,000 employees
Real User
Top 20
May 19, 2026
My main use case for PagerDuty Operations Cloud is targeted alert paging. A specific example of how I use targeted alert paging with PagerDuty Operations Cloud is that we were able to tie our services and use PagerDuty AIOps to understand dependencies so we get paged a minimal number of times as opposed to an alert storm.
Sr. Specialist at a tech vendor with 10,001+ employees
Real User
Top 20
May 3, 2026
PagerDuty Operations Cloud serves as an alerting tool for our organization. Whenever applications or any protection systems are down, we get notified via email and mobile. My current role in PagerDuty Operations Cloud is an admin role. Whenever there is a new user or any higher management requirement to grant email access to a particular server or program, we receive the request via the team, and they mention which team they are from and their role. Based on that information, if I get one email ID with a name, I have to enable the ID, then I have to add the name to PagerDuty Operations Cloud for the particular program. When the application is down or any issue is triggered, the user will be notified via mobile and email.
Senior Executive at a consultancy with 11-50 employees
Real User
Apr 26, 2026
I handle level two operations. Whenever a major incident occurs or there is an outage, I am informed first via Splunk, DataDog, and PagerDuty. PagerDuty Operations Cloud is used for alertness, and we have configured threshold values within it. I have the mobile application installed on my phone, so I receive information about any outage as soon as it occurs. I work at Vodafone Intelligence Services, which is a subsidiary of Vodafone. We are a UK-based company that performs level one, level two, and level three operations for all European countries and some countries in India, including South Africa, Ghana, Spain, Egypt, Hungary, and the UK. These are our major customers. As part of operations, we have a team of about 15,000 people who manage the different markets and customers. PagerDuty Operations Cloud is used everywhere across our organization, along with Splunk and AppDynamics.
My main use case for PagerDuty Operations Cloud is for cloud-based operations, including incident management and resolving incident responses to reduce downtime and improve reliability. PagerDuty Operations Cloud provides a central command center that collects data signals from various IT systems, which helps us detect high-priority incidents and reduce the noise. In addition to my main use case, we are able to perform on-call scheduling and routing with the help of PagerDuty Operations Cloud very easily.
My main use case of PagerDuty Operations Cloud is in incident management . If there was any issue that occurred, I used to be the point of contact to resolve that particular issue, and in some cases I used to integrate that as well. My usual use cases for PagerDuty Operations Cloud were alert notification, on-call scheduling, and escalations. If an incident arose and we wanted to escalate that issue to a senior person within the service level agreement, we used to schedule escalation policies so that we could ensure coverage at all times. Some of my use cases were setting up an alarm, integrating and monitoring with Slack or other tools so that we could monitor and resolve issues better because of using the application.
We took a subscription and started integrating many applications with it. We have integrated it with ServiceNow and also use our wiki repository by Spotify called Beacon. We onboard multiple services and integrate them with PagerDuty so that different engineering teams can update their line of escalation policy and primary, secondary, and tertiary users who are on call 24/7 or following a Follow the Sun model. Additionally, we have integrated PagerDuty for incident commanders who can acknowledge any incident page they receive and update their responses. If they need to hand it over to someone else, they can do that as well. This is how we are actually using PagerDuty. We primarily leverage it for service onboarding. For instance, if you have created a product and need to support it, which might run on any cloud such as EC2, AWS, or Azure, the service teams or triage engineering teams have their members added into PagerDuty along with different playbooks, runbooks, and SOPs integrated with any ticketing tool such as ServiceNow or Jira, whichever you are using. On these fronts, we effectively use PagerDuty. PagerDuty Operations Cloud has been reliable as it has never gone down in my experience; I have never seen it fail. I rely on PagerDuty Operations Cloud for on-call support for any high severity incidents or sev zero scenarios. This is a great feature, and I can update it using the mobile app because we also use PagerDuty Operations Cloud mobile app. Occasionally, I may be on call during weekends, and something might come up. For example, on October 20th, I was on leave but used PagerDuty Operations Cloud to stay in sync, even without my laptop. I could join Teams and Zoom calls and simultaneously update required documentation via Copilot and ChatGPT on PagerDuty Operations Cloud regarding different incidents. This capability was extremely helpful.
L1 SecOps Analyst at a tech vendor with 10,001+ employees
Real User
Top 20
Feb 23, 2026
PagerDuty Operations Cloud is used to understand whether there are false positives or false alerts because it is integrated into a system where there would be a phone number, such as anyone from the team's phone number or perhaps the shift supervisor's phone number. Once those triggers are there, it would hit that number until somebody picks it up and acknowledges or escalates the alert or threat. In a recent incident through PagerDuty Operations Cloud, there was one issue with one appliance that was continuously going wrong. Even if we were acknowledging it and closing it and had denoted it as a false positive and a false trigger or false alarm, it was still continuously hitting us. There was some issue with the appliance or some issue with the server. Once we understood that, we escalated this to the company's SecOps team, and they had to go inside to find out more details. PagerDuty Operations Cloud team was coordinated with them. Once they were coordinated, they could dig in deep and find out what the issue was. A high-priority P1 ticket was raised for that as per ITIL principles. PagerDuty Operations Cloud was being used inside the office only, and if we enter anybody's number, for example, it would continuously be hitting at any time whenever the alerts are there. Whoever's numbers are added, such as a shift supervisor or shift people, those would keep on hitting back.
I have been working in my current field for over seven years as a DevOps and site reliability engineer, and my primary experience involves managing the reliability of infrastructure platforms hosted in multi-cloud and on-premises environments. I have predominantly worked with systems hosted in AWS services, setting up infrastructure, CI/CD, observability, and completely establishing the release process where I utilize PagerDuty Operations Cloud for triage and other SRE operations. I have been using PagerDuty Operations Cloud for over four years, and I have utilized it in multiple ways. One involves using PagerDuty Operations Cloud through enterprise services via a subscription model, and I have also used it in a project at Intel where I utilized PagerDuty Operations Cloud from AWS for approximately one to one and a half years. After that period, I have been using it as a subscription currently at IBM. One of the main use cases for PagerDuty Operations Cloud involves handling the operation center, particularly concerning incident resolutions and triaging different incidents as part of the score platform engineering team within a central IBM cloud where various IBM cloud services are hosted. To ensure continuous reliability, automated incidents are created in PagerDuty Operations Cloud and incident management automation is heavily utilized as part of the project. Previously, I worked on integrating PagerDuty Operations Cloud with default AWS services to create incidents for different AWS services as part of the host infrastructure at Intel. Currently, I am creating different incident workflows within IBM internal cloud operations to ensure an effective incident management process, utilizing integrations with different LLMs as part of incident management, along with agentic SRE tasks that have arisen in the project. Since I am part of a larger platform engineering team and SRE operations team, there are many incidents and services that my team handles. I handle over 26 IBM cloud score services hosted in our internal platform, where there have been many incidents related to service downtime, reliability issues, and update issues. A dedicated SRE team handles end-to-end incident management, and we wanted to automate the incident management process, especially since we receive hundreds of incidents per day, up to thousands of incidents during critical release times of different services. Thus, the manual on-call process has been automated through utilizing PagerDuty Operations Cloud.
We receive a notification if there are any failed jobs or operations. We have some Bamboo agents working, so if one of the jobs fails on one of these servers, PagerDuty Operations Cloud creates an incident and notifies us. We use PagerDuty Operations Cloud for monitoring purposes, and it works great for our current needs.
My main use case for PagerDuty Operations Cloud is monitoring and on-call management for downtime. Recently, we had a service go down last week, and we were alerted via PagerDuty Operations Cloud of the issue. One of our on-call engineers responded to the page and quickly resolved the problem through PagerDuty Operations Cloud app.
Senior Business Application Analyst (Product Team) at Kotak Mahindra Bank
Real User
Top 5
Nov 27, 2025
My main use case for PagerDuty Operations Cloud is to set up alerts for any failures, such as when one server is down, a particular service is down, or when APIs are not responding due to technical issues, with PagerDuty triggering an alert and also calling my personal mobile number to notify me about the issue, allowing me to acknowledge that I am looking into it and take necessary actions. I can give an example of a situation where PagerDuty Operations Cloud helped us handle an incident, such as when our payment system was about to go down. During that time, we usually monitor the system manually, but there are incidents where an automated system works more efficiently than a human. PagerDuty Operations Cloud identified the issue first by alerting us that something went wrong with the servers or services, which enabled us to contact the DevOps and Dev team to identify the exact issue in our banking app, highlighting how helpful PagerDuty Operations Cloud has been from the beginning. PagerDuty Operations Cloud is very helpful for monitoring purposes, allowing us to set up multiple alerting methods such as SMS alerting, email alerting, and call alerting, all of which we commonly use, proving its usefulness across various banking services, with teams including Dev, DevOps, and SecOps relying on it heavily.
Initially, I started using it for managing on-call schedules. As the tech stack developed, we began using it for service alerts and event routing, and then transitioned to operational views and dashboarding. It eventually became central to our alerting systems, where all monitoring tools would send information to PagerDuty, enabling event management and routing to whoever was on call.
The primary use case of the solution is to alert the on-call person when there are any critical errors or when the servers are down. It is also used for the on-call scheduling of personnel.
Principal Architect at a energy/utilities company with 10,001+ employees
Real User
Sep 19, 2022
It's mainly for IT call scheduling, emergency contacts, events, and those kinds of things. It's integrated with AWS, MS Teams, Remedy, and other solutions.
Compliance, Security & Testing Manager at a financial services firm with 11-50 employees
Real User
Oct 8, 2020
We are a 24-hour online business. We use it for scheduling our on-call engineers and making sure that there is follow-the-sun or round-the-clock coverage for alerting and network operations. It ingests all our alert paths, i.e., anything that generates an alert of any description, such as, Splunk, AWS, and internal applications. We feed all our events into it, then it generates alerts which need a response from an engineer with a description. Another thing is it is built-in scheduling is pretty much hands-off for our on-call engineers unless somebody goes on holidays. That is the only time that we have to jump in there and make any changes.
VP of Engineering at a comms service provider with 201-500 employees
Real User
Jun 25, 2020
We mostly use it for our on-call engineers, for schedules, alerting, and critical alerts. And, of course, we use it for the management of an issue, so that people acknowledge the alerts, reassign them, etc.
Tier 4 Support Team Leader at a comms service provider with 10,001+ employees
Real User
Mar 1, 2020
The most common use case is the result of alerts coming from a monitoring system, like New Relic or Nagios, alerts that we define as critical. They are alerts where we need someone to get on a bridge or to start working on them during the night. Once such an alert is firing, it fires a PagerDuty alert and it triggers the current on-call who is scheduled in PagerDuty's schedule. The on-call person acknowledges the alert and looks into it to understand what is going on and to update, via PagerDuty, what the status is. The update will be sent to all the groups that are part of the PagerDuty schedule until the issue is resolved. We mostly integrate it with other monitoring tools like New Relic or Nagios, or we are using their email integration for on-call processes to page people in groups. We also use it for Sev 1 issues that are coming from alerts from New Relic or from Nagios or other monitoring systems.
PagerDuty Operations Cloud focuses on efficient incident management, featuring advanced alert and notification systems, mobile alerts, and AI-driven functionalities that facilitate streamlined on-call schedules and integrations with major monitoring tools.PagerDuty Operations Cloud offers comprehensive incident management with real-time alerts and notifications via mobile, SMS, and calls. This empowers teams to respond swiftly and reduce missed incidents. Efficient on-call management through...
My usual use cases with PagerDuty Operations Cloud involve handling incidents through a full flow. When there is an outage, an incident is created that can be either severity one or severity two. As the person on call that day, I receive a page from PagerDuty on the app and three calls on my cell phone. When I pick up the call, PagerDuty IVR asks me to acknowledge the incident. Once I acknowledge the incident from the call, I go to PagerDuty through the website, which is much easier to navigate than the mobile app. I then page other teams responsible for the incident, as well as the stakeholders and product owners. PagerDuty is integrated with Microsoft Teams, so I open a Teams bridge call to resolve the issue and update all incident details in PagerDuty notes, which automatically integrates with ServiceNow incident and sends messages to stakeholders' phone numbers. One time while hanging out in the mountains, there was no internet signal but there was cell reception. An incident happened while I was on call that day. Normally, without internet, I would not be able to know about it, but because of PagerDuty, I was paged three times on my cell phone as well as through text message. I managed to call another colleague from my work and told him to take care of the incident. This helped me avoid breaching the SLAs on incident acknowledgment and allowed me to access remote incidents without relying solely on the internet.
My main use case for PagerDuty Operations Cloud is to manage critical alerts and incidents across our production systems. It helps our team route alerts to the right people, manage on-call schedules, coordinate responses, and reduce downtimes. We also use its integrations with our monitoring and collaboration tools, so issues are identified and addressed quickly before they impact our customers.
I have been using PagerDuty for the last nine years, but PagerDuty Operations Cloud for over one and a half years. We work directly with merchants and need to trigger immediate alerts whenever there are 5xx errors or business errors like 4xx issues, as well as payment failures. We have configured every alert on a data log in some other monitoring tools that are integrated with PagerDuty. We receive alerts very immediately and trigger calls and Slack notifications. We integrate everything with PagerDuty and get notifications instantly, after which we start our triage process. One use case I can mention is when we have an auth rate dip. Whenever there is an auth rate dip, we run into revenue losses with the merchants or partners that PayPal currently works with. Since everything is integrated, PagerDuty Operations Cloud catches when there is an auth rate dip for particular merchants and immediately triggers a notification for us. We then immediately dive into what the problem is and figure out how to fix the issue with the help of engineering teams.
I primarily use PagerDuty Operations Cloud for alert management and incident call rotations. In my earlier firm, we managed rotation shifts across three time zones: EMEA, APAC, and New York time. All rotation and shift management was handled through PagerDuty Operations Cloud schedules. Application monitoring was also updated through PagerDuty Operations Cloud. According to the schedule, we updated people's contact information so that in case of any issues, the contact would be transferred to the respective shift member. We also managed escalations with five layers of escalation. If a first team member missed an alert, it would go to the second team member after 10 minutes, then to the next person after five minutes, continuing according to the priority of the service.
I usually use PagerDuty Operations Cloud for the notification of high-priority incidents within the infrastructure. I also use it for escalating to the on-call members, scheduling the priority of incidents or issues within the infrastructure, and creating scheduled rotations for team members.
As a cloud operation team, I was a user who set the alerts, and whatever important incidents or anomalies were detected that needed to be immediately taken care of were bifurcated through our APM tools that we integrated with PagerDuty Operations Cloud. As a cloud operation team, we supported the platform for rotational shifts. My roles involved setting the person in the shift according to the shift roster, so whenever any incidents triggered, they would get the call. The primary use was supporting production operations and cloud activities. Our multi-environment consists of AWS infrastructure, Linux servers, Kubernetes clusters, and customer-facing applications. PagerDuty Operations Cloud was mainly used for incident management and alerting. We integrated it with AppDynamics, Instana, and CloudWatch, where it would monitor the patterns and platform, and then PagerDuty Operations Cloud would generate the critical alerts that the appropriate support team who was working in that present shift would get notified of immediately. This platform really helped us manage production incidents beyond service outages, mostly high CPU utilization where we set alerts, application failures, pod issues in Kubernetes, and infrastructure-related alerts. We configured all kinds of alerts, which ensured that alerts were routed to the correct on-call person, helping us reduce response time in critical situations.
My main use case for PagerDuty Operations Cloud involves working in the AIOps team, which is an operations team. We have a monitoring tool called Checkmk, and we have integrated it with PagerDuty for incident management. We monitor many servers across different teams, including the Linux team, network team, Windows teams, and database team. All of these servers are monitored in Checkmk, tracking live CPU, memory, and file systems. Upon reaching certain thresholds, Checkmk generates events, which we integrate into PagerDuty Operations Cloud console. There, we set conditions so that if an event is critical or a warning, it converts into an incident. We then route the incidents to respective teams, who handle the details. A specific example of an incident where PagerDuty Operations Cloud played a key role involves automations we created within PagerDuty. Rundeck is a job workflow tool where we can implement scripts or schedule jobs. If a server meets its threshold, it triggers PagerDuty Operations Cloud. We create scripts in Rundeck to handle issues, such as clearing a full file system. We utilize a feature called Automation Actions in PagerDuty Operations Cloud, and whenever an incident comes that matches specific conditions, that job will automatically run in Rundeck. This incident management cycle is effectively managed in PagerDuty Operations Cloud, allowing jobs to run and resolve incidents automatically, ensuring the server is healthy again.
My main use case for PagerDuty Operations Cloud is monitoring multiple platforms. In cloud operations, whenever any application or device has an issue, it triggers PagerDuty, and the on-call shift engineer can immediately check on it. For a specific example of how I have used PagerDuty Operations Cloud to handle an incident, if any media connect flow breaks or applications deployed on cloud instances have issues, we start receiving alerts. This can be configured using PagerDuty schedules to inform the on-call engineer to take immediate action. Regarding my main use case, I would add that we do not need to be on the dashboards or monitor manually. Instead, our APIs work in the back-end. They check for things working on the platform's cloud, and if any action is required, the appropriate team can be aligned using PagerDuty.
The main purpose of PagerDuty Operations Cloud is to receive alarms for real incident production critical issues. Whenever an incident happens, we get an alarm call or a phone call on our phone or through an application call, so that we are aware of the situation and know that we have to be vigilant about it and resolve that particular issue. This is the main use case of PagerDuty Operations Cloud, and it is the use case where most of the company is using it.
I use PagerDuty Operations Cloud for core tasks only because my client completely belongs to healthcare. Healthcare is a high priority where we need to help people. I primarily use it for core tasks and sometimes for primary tasks as well.
My main use case for PagerDuty Operations Cloud is targeted alert paging. A specific example of how I use targeted alert paging with PagerDuty Operations Cloud is that we were able to tie our services and use PagerDuty AIOps to understand dependencies so we get paged a minimal number of times as opposed to an alert storm.
PagerDuty Operations Cloud serves as an alerting tool for our organization. Whenever applications or any protection systems are down, we get notified via email and mobile. My current role in PagerDuty Operations Cloud is an admin role. Whenever there is a new user or any higher management requirement to grant email access to a particular server or program, we receive the request via the team, and they mention which team they are from and their role. Based on that information, if I get one email ID with a name, I have to enable the ID, then I have to add the name to PagerDuty Operations Cloud for the particular program. When the application is down or any issue is triggered, the user will be notified via mobile and email.
I handle level two operations. Whenever a major incident occurs or there is an outage, I am informed first via Splunk, DataDog, and PagerDuty. PagerDuty Operations Cloud is used for alertness, and we have configured threshold values within it. I have the mobile application installed on my phone, so I receive information about any outage as soon as it occurs. I work at Vodafone Intelligence Services, which is a subsidiary of Vodafone. We are a UK-based company that performs level one, level two, and level three operations for all European countries and some countries in India, including South Africa, Ghana, Spain, Egypt, Hungary, and the UK. These are our major customers. As part of operations, we have a team of about 15,000 people who manage the different markets and customers. PagerDuty Operations Cloud is used everywhere across our organization, along with Splunk and AppDynamics.
The major part of PagerDuty Operations Cloud is for communication and pinging people and engaging them and pulling them into the call.
My main use case for PagerDuty Operations Cloud is for cloud-based operations, including incident management and resolving incident responses to reduce downtime and improve reliability. PagerDuty Operations Cloud provides a central command center that collects data signals from various IT systems, which helps us detect high-priority incidents and reduce the noise. In addition to my main use case, we are able to perform on-call scheduling and routing with the help of PagerDuty Operations Cloud very easily.
My main use case of PagerDuty Operations Cloud is in incident management . If there was any issue that occurred, I used to be the point of contact to resolve that particular issue, and in some cases I used to integrate that as well. My usual use cases for PagerDuty Operations Cloud were alert notification, on-call scheduling, and escalations. If an incident arose and we wanted to escalate that issue to a senior person within the service level agreement, we used to schedule escalation policies so that we could ensure coverage at all times. Some of my use cases were setting up an alarm, integrating and monitoring with Slack or other tools so that we could monitor and resolve issues better because of using the application.
We took a subscription and started integrating many applications with it. We have integrated it with ServiceNow and also use our wiki repository by Spotify called Beacon. We onboard multiple services and integrate them with PagerDuty so that different engineering teams can update their line of escalation policy and primary, secondary, and tertiary users who are on call 24/7 or following a Follow the Sun model. Additionally, we have integrated PagerDuty for incident commanders who can acknowledge any incident page they receive and update their responses. If they need to hand it over to someone else, they can do that as well. This is how we are actually using PagerDuty. We primarily leverage it for service onboarding. For instance, if you have created a product and need to support it, which might run on any cloud such as EC2, AWS, or Azure, the service teams or triage engineering teams have their members added into PagerDuty along with different playbooks, runbooks, and SOPs integrated with any ticketing tool such as ServiceNow or Jira, whichever you are using. On these fronts, we effectively use PagerDuty. PagerDuty Operations Cloud has been reliable as it has never gone down in my experience; I have never seen it fail. I rely on PagerDuty Operations Cloud for on-call support for any high severity incidents or sev zero scenarios. This is a great feature, and I can update it using the mobile app because we also use PagerDuty Operations Cloud mobile app. Occasionally, I may be on call during weekends, and something might come up. For example, on October 20th, I was on leave but used PagerDuty Operations Cloud to stay in sync, even without my laptop. I could join Teams and Zoom calls and simultaneously update required documentation via Copilot and ChatGPT on PagerDuty Operations Cloud regarding different incidents. This capability was extremely helpful.
PagerDuty Operations Cloud is used to understand whether there are false positives or false alerts because it is integrated into a system where there would be a phone number, such as anyone from the team's phone number or perhaps the shift supervisor's phone number. Once those triggers are there, it would hit that number until somebody picks it up and acknowledges or escalates the alert or threat. In a recent incident through PagerDuty Operations Cloud, there was one issue with one appliance that was continuously going wrong. Even if we were acknowledging it and closing it and had denoted it as a false positive and a false trigger or false alarm, it was still continuously hitting us. There was some issue with the appliance or some issue with the server. Once we understood that, we escalated this to the company's SecOps team, and they had to go inside to find out more details. PagerDuty Operations Cloud team was coordinated with them. Once they were coordinated, they could dig in deep and find out what the issue was. A high-priority P1 ticket was raised for that as per ITIL principles. PagerDuty Operations Cloud was being used inside the office only, and if we enter anybody's number, for example, it would continuously be hitting at any time whenever the alerts are there. Whoever's numbers are added, such as a shift supervisor or shift people, those would keep on hitting back.
I have been working in my current field for over seven years as a DevOps and site reliability engineer, and my primary experience involves managing the reliability of infrastructure platforms hosted in multi-cloud and on-premises environments. I have predominantly worked with systems hosted in AWS services, setting up infrastructure, CI/CD, observability, and completely establishing the release process where I utilize PagerDuty Operations Cloud for triage and other SRE operations. I have been using PagerDuty Operations Cloud for over four years, and I have utilized it in multiple ways. One involves using PagerDuty Operations Cloud through enterprise services via a subscription model, and I have also used it in a project at Intel where I utilized PagerDuty Operations Cloud from AWS for approximately one to one and a half years. After that period, I have been using it as a subscription currently at IBM. One of the main use cases for PagerDuty Operations Cloud involves handling the operation center, particularly concerning incident resolutions and triaging different incidents as part of the score platform engineering team within a central IBM cloud where various IBM cloud services are hosted. To ensure continuous reliability, automated incidents are created in PagerDuty Operations Cloud and incident management automation is heavily utilized as part of the project. Previously, I worked on integrating PagerDuty Operations Cloud with default AWS services to create incidents for different AWS services as part of the host infrastructure at Intel. Currently, I am creating different incident workflows within IBM internal cloud operations to ensure an effective incident management process, utilizing integrations with different LLMs as part of incident management, along with agentic SRE tasks that have arisen in the project. Since I am part of a larger platform engineering team and SRE operations team, there are many incidents and services that my team handles. I handle over 26 IBM cloud score services hosted in our internal platform, where there have been many incidents related to service downtime, reliability issues, and update issues. A dedicated SRE team handles end-to-end incident management, and we wanted to automate the incident management process, especially since we receive hundreds of incidents per day, up to thousands of incidents during critical release times of different services. Thus, the manual on-call process has been automated through utilizing PagerDuty Operations Cloud.
We receive a notification if there are any failed jobs or operations. We have some Bamboo agents working, so if one of the jobs fails on one of these servers, PagerDuty Operations Cloud creates an incident and notifies us. We use PagerDuty Operations Cloud for monitoring purposes, and it works great for our current needs.
My main use case for PagerDuty Operations Cloud is monitoring and on-call management for downtime. Recently, we had a service go down last week, and we were alerted via PagerDuty Operations Cloud of the issue. One of our on-call engineers responded to the page and quickly resolved the problem through PagerDuty Operations Cloud app.
My main use case for PagerDuty Operations Cloud is to set up alerts for any failures, such as when one server is down, a particular service is down, or when APIs are not responding due to technical issues, with PagerDuty triggering an alert and also calling my personal mobile number to notify me about the issue, allowing me to acknowledge that I am looking into it and take necessary actions. I can give an example of a situation where PagerDuty Operations Cloud helped us handle an incident, such as when our payment system was about to go down. During that time, we usually monitor the system manually, but there are incidents where an automated system works more efficiently than a human. PagerDuty Operations Cloud identified the issue first by alerting us that something went wrong with the servers or services, which enabled us to contact the DevOps and Dev team to identify the exact issue in our banking app, highlighting how helpful PagerDuty Operations Cloud has been from the beginning. PagerDuty Operations Cloud is very helpful for monitoring purposes, allowing us to set up multiple alerting methods such as SMS alerting, email alerting, and call alerting, all of which we commonly use, proving its usefulness across various banking services, with teams including Dev, DevOps, and SecOps relying on it heavily.
Initially, I started using it for managing on-call schedules. As the tech stack developed, we began using it for service alerts and event routing, and then transitioned to operational views and dashboarding. It eventually became central to our alerting systems, where all monitoring tools would send information to PagerDuty, enabling event management and routing to whoever was on call.
We use the solution for incident management.
The solution is used to alert the on-call users if we have priority-one or business-critical issues.
Our use cases include generating alerts from our site 24/7. We are managing the cloud infrastructure there.
The two major use cases were alerts for events and scheduling of engineers to get pages based on incidents.
The primary use case of the solution is to alert the on-call person when there are any critical errors or when the servers are down. It is also used for the on-call scheduling of personnel.
We primarily use this solution to track alerts from our cloud environment and monitor and respond to alerts on our cloud platform.
It's mainly for IT call scheduling, emergency contacts, events, and those kinds of things. It's integrated with AWS, MS Teams, Remedy, and other solutions.
We use PagerDuty for incident managment. We're looking at integrating PagerDuty with Rundeck in the future.
We are a 24-hour online business. We use it for scheduling our on-call engineers and making sure that there is follow-the-sun or round-the-clock coverage for alerting and network operations. It ingests all our alert paths, i.e., anything that generates an alert of any description, such as, Splunk, AWS, and internal applications. We feed all our events into it, then it generates alerts which need a response from an engineer with a description. Another thing is it is built-in scheduling is pretty much hands-off for our on-call engineers unless somebody goes on holidays. That is the only time that we have to jump in there and make any changes.
We mostly use it for our on-call engineers, for schedules, alerting, and critical alerts. And, of course, we use it for the management of an issue, so that people acknowledge the alerts, reassign them, etc.
The most common use case is the result of alerts coming from a monitoring system, like New Relic or Nagios, alerts that we define as critical. They are alerts where we need someone to get on a bridge or to start working on them during the night. Once such an alert is firing, it fires a PagerDuty alert and it triggers the current on-call who is scheduled in PagerDuty's schedule. The on-call person acknowledges the alert and looks into it to understand what is going on and to update, via PagerDuty, what the status is. The update will be sent to all the groups that are part of the PagerDuty schedule until the issue is resolved. We mostly integrate it with other monitoring tools like New Relic or Nagios, or we are using their email integration for on-call processes to page people in groups. We also use it for Sev 1 issues that are coming from alerts from New Relic or from Nagios or other monitoring systems.
Our primary use case of this solution is for alarming and to mitigate threats in our organization.