No more typing reviews! Try our Samantha, our new voice AI agent.
Tony Martinez1 - PeerSpot reviewer
Works at VANTA INC
Real User
Top 20
Oct 1, 2024
Great logging, session replays, and alerting
Pros and Cons
  • "Dashboards are helpful for reviewing occasionally to get a higher-level overview of what's happening."
  • "The UI has a lot going on. It should be simpler and have a better way to onboard someone new to using Datadog."

What is our primary use case?

Our primary use cases include:

  • Alert on errors customers encounter in our product. We've set up logs that go to slack to tell us when a certain error threshold is hit.
  • Investigate slow page load times. We have pages in our app that are loading slowly and the logs help us figure out which queries are taking the longest time.
  • Metrics. We collect metrics on product usage.
  • Session replays. We watch session replays to see what a user was doing when a page took a long time to load or hit an error. This is helpful.

How has it helped my organization?

It's helped us find bugs that customers are experiencing before they're reported to us. Sometimes, customers don't report errors, so being able to catch errors before they're reported helps us investigate before other users find errors

Datadog has helped us investigate slow page loading times and even see the specific queries that are taking a long time to load

Logging lets us see the context around an error. For example, see if a backend service had an error before it surfaced on the frontend.

Dashboards are helpful for reviewing occasionally to get a higher-level overview of what's happening.

What is most valuable?

The most valuable aspects include: 

  • Logging. Being able to view detailed logs helps debug issues.
  • Session replays. They are helpful for seeing what a customer was doing before they saw an error or had a slow page load
  • Alerting. This is an important part of our on-call process to send alerts to slack when an error threshold is crossed. Alerts/monitors are easy to configure to only alert when we want them to alert.
  • Dashboards. It's helpful to pull up dashboards that show our most common errors or page performance. It's a good way to see how the app is performing from a birds-eye-view.

What needs improvement?

The UI has a lot going on. It should be simpler and have a better way to onboard someone new to using Datadog.

The log querying syntax can be confusing. Usually, I filter by finding a facet in a log and selecting to filter by that facet - but I'm not sure how to write the filter myself

The monitor/alert syntax is also somewhat hard to understand.

Overall, it should be easier to learn how to use the product while you're using the product. Perhaps tooltips or a link to learn more about whatever section you're using.

Buyer's Guide
Datadog
September 2026
Learn what your peers think about Datadog. Get advice and tips from experienced pros sharing their opinions. Updated: September 2026.
913,806 professionals have used our research since 2012.

For how long have I used the solution?

I've used the solution for two years.

Which solution did I use previously and why did I switch?

We did not previously use a different solution.

Which other solutions did I evaluate?

We did not evaluate other options. 

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Neil Elver - PeerSpot reviewer
Application Development Team Lead at TCS EDUCATION SYSTEM
User
Sep 20, 2024
Good synthetic testing, centralized pipeline tracking and error logging
Pros and Cons
  • "Synthetic testing has been a game-changer, allowing us to catch potential problems before they impact real users."
  • "I'd like to see an expansion of the Android and IOS apps to have a simplified CI/CD pipeline history view."

What is our primary use case?

Our primary use case is custom and vendor-supplied web application log aggregation, performance tracing and alerting. 

We run a mix of AWS EC2, Azure serverless, and colocated VMWare servers to support higher education web applications. 

Managing a hybrid multi-cloud solution across hundreds of applications is always a challenge. Datadog agents on each web host and native integrations with GitHubAWS, and Azure get all of our instrumentation and error data in one place for easy analysis and monitoring.

How has it helped my organization?

Through the use of Datadog across all of our apps, we were able to consolidate a number of alerting and error-tracking apps, and Datadog ties them all together in cohesive dashboards. Whether the app is vendor-supplied or we built it ourselves, the depth of tracing, profiling, and hooking into logs is all obtainable and tunable. Both legacy .NET Framework and Windows Event Viewer and cutting-edge .NET Core with streaming logs all work. The breadth of coverage for any app type or situation is really incredible. It feels like there's nothing we can't monitor.

What is most valuable?

When it comes to Datadog, several features have proven particularly valuable. 

The centralized pipeline tracking and error logging provide a comprehensive view of our development and deployment processes, making it much easier to identify and resolve issues quickly. 

Synthetic testing has been a game-changer, allowing us to catch potential problems before they impact real users. Real user monitoring gives us invaluable insights into actual user experiences, helping us prioritize improvements where they matter most. And the ability to create custom dashboards has been incredibly useful, allowing us to visualize key metrics and KPIs in a way that makes sense for different teams and stakeholders. 

Together, these features form a powerful toolkit that helps us maintain high performance and reliability across our applications and infrastructure, ultimately leading to better user satisfaction and more efficient operations.

What needs improvement?

I'd like to see an expansion of the Android and IOS apps to have a simplified CI/CD pipeline history view. I like the idea of monitoring on the go, however, it seems the options are still a bit limited out of the box. 

While the documentation is very good considering all the frameworks and technology Datadog covers, there are areas - specifically .NET Profiling and Tracing of IIS-hosted apps - that need a lot of focus to pick up on the key details needed. In some cases the screenshots don't match the text as updates are made. I feel I spent longer than I should figuring out how to correlate logs to traces, mostly related to environmental variables.

For how long have I used the solution?

I've used the solution for about three years.

What do I think about the stability of the solution?

We have been impressed with the uptime and clean and light resource usage of the agents.

What do I think about the scalability of the solution?

The solution was very scalable and very customizable.

How are customer service and support?

Sales service is always helpful in tuning our committed costs and alerting us when we start spending outside the on-demand budget.

Which solution did I use previously and why did I switch?

We used a mix of a custom error email system, SolarWinds, UptimeRobot, and GitHub actions. We switched to find one platform that could give deep app visibility regardless of Linux, Windows, Container, cloud or on-prem hosted.

How was the initial setup?

The setup is generally simple. That said, .NET Profiling of IIS and aligning logs to traces and profiles was a challenge.

What about the implementation team?

The solution was iImplemented in-house. 

What was our ROI?

I'd count our ROI as significant time saved by the development team assessing bugs and performance issues.

What's my experience with pricing, setup cost, and licensing?

It's a good idea to set up live trials to asses cost scaling. Small decisions around how monitors are used can have big impacts on cost scaling. 

Which other solutions did I evaluate?

NewRelic was considered. LogicMonitor was chosen over Datadog for our network and campus server management use cases.

What other advice do I have?

We are excited to dig further into the new offerings around LLM and continue to grow our footprint in Datadog. 

Which deployment model are you using for this solution?

Hybrid Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Microsoft Azure
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Buyer's Guide
Datadog
September 2026
Learn what your peers think about Datadog. Get advice and tips from experienced pros sharing their opinions. Updated: September 2026.
913,806 professionals have used our research since 2012.
Dmitri Panfilov - PeerSpot reviewer
Software Engineer at Redfin Corp
User
Top 20
Sep 20, 2024
Easy dashboard creation and alarm monitoring with a good ROI
Pros and Cons
  • "The ease of dashboard creation and alarm monitoring has helped us not only stay competitive but be industry leaders in performance."
  • "The product can be improved by allowing the grouping of APIs to add variables. That way, any API with a unique ID could be grouped together."

What is our primary use case?

We use the solution to monitor production service uptime/downtime, latency, and log storage. 

Our entire monitoring infrastructure runs off Datadog, so all our alarms are configured with it. We also use it for tracing API performance; what are the biggest regression points. 

Finally we use it to compare performance on SEO metrics vs competitors. This is a primary use case as SEO dictates our position from google traffic which is a large portion of our customer view generation so it is a vital part of the business we rely on datadog for.

How has it helped my organization?

The product improved the organization primarily by providing consistent data with virtually zero downtime. This was a problem we had with an old provider. It also made it easy to transition an otherwise massive migration involving hundreds of alarms. 

The training provided was crucial, along with having a dedicated team that can forward our requests to and from Datadog efficiently. Without that, we may have never transitioned to Datadog in the first place since it is always hard to lead a migration for an entire company.

What is most valuable?

The API tracing has been massive for debugging latency regressions and how to improve the performance of our least performant APIs. Through tracing, we managed to find the slowest step of an API, improve its latency, and iterate on the process until we had our desired timings. This is important for improving our SEO as LCP, INP are directly taking from the numbers we see on Datadog for our API timings. 

The ease of dashboard creation and alarm monitoring has helped us not only stay competitive but be industry leaders in performance.

What needs improvement?

The product can be improved by allowing the grouping of APIs to add variables. That way, any API with a unique ID could be grouped together. 

Furthermore, SEO monitoring has been crucial for us but also a difficult part to set up as comparing alarms between us and competitors is a tough feat. Data is not always consistent so we have been toying and experimenting with removing the noise of datadog but its been taking a while. 

Finally, Datadog should have a feature that reports stale alarms based on activity.

For how long have I used the solution?

I've used the solution for six months.

What do I think about the stability of the solution?

Its very stable and we have not experienced an issue with downtime on Datadog.

What do I think about the scalability of the solution?

Datadog works well for scalability as volume has not seemed to slow.

How are customer service and support?

We haven't talked to the support team. 

How would you rate customer service and support?

Positive

Which solution did I use previously and why did I switch?

We switched to Datadog as we used to have a provider that had very inconsistent logging. Our alarms would often not fire since our services were not working since the provider had a logging problem.

How was the initial setup?

The initial setup was somewhat complex due to the built-in monitoring with services. This is not always super comprehensive and has to be studied as opposed to other metrics platforms that just service all your endpoints, which you can trace them with Grafana.

What about the implementation team?

We implemented the solution through an in-house team.

What was our ROI?

The ROI is good.

What's my experience with pricing, setup cost, and licensing?

Users must try to understand the way Datadog alarms work off the bat so that they can minimize the requirements for expensive features like custom metrics. 

It can sometimes be tempting to use them; however, it is not always necessary as you migrate to Datalog, as they are a provider that treats alarms somewhat differently than you may be used to.

Which other solutions did I evaluate?

We have evaluated New Relic, Grafana, Splunk, and many more in our quest to find the best monitoring provider.

Which deployment model are you using for this solution?

Hybrid Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
reviewer1494894 - PeerSpot reviewer
Senior Manager, Site Reliability Engineering at Extra Space Storage
Real User
Sep 19, 2024
Improved time to discovery and resolution but needs better consumption visibility
Pros and Cons
  • "Several critical dashboards were created years ago and are still in use today."
  • "We would love to see a % consumed and alert us if we are over budget before getting an overage charge 20 days into the month."

What is our primary use case?

The product monitors multiple systems, from customer interactions on our web applications down to the database and all layers in between. RUM, APM, logging, and infrastructure monitoring are all surfaced into single dashboards.

We initially started with application logs and generated long-term business metrics out of critical logs. We have turned those metrics and logs into a collection of alerts integrated into our pager system. As we have evolved, we have also used APM and RUM data to trigger additional alerts

How has it helped my organization?

The solution has surfaced how integrated our applications really are and helps us track calls from the top down, identifying slowness and errors all through the call stack.

The biggest improvement we have seen is our time to discovery and resolution. As Datadog has improved, and we add new features, the depth and clarity we get from top to bottom has been excellent. Our engineering teams have quickly adopted many features within Datadog, and are quick to build out their own dashboards and alerts. This has also led to a rapid sprawl when left unchecked.

What is most valuable?

We started with application logs and have expanded over the years to include infrastructure, APM, and now RUM. All of these tools have been incredibly valuable in their own sphere. The huge value is tying all of the data points together.

Logging was the first tool we started with years ago, replacing our ELK stack. It was the easiest to get in place, and our engineers quickly embraced the tools. Several critical dashboards were created years ago and are still in use today. Over time, we have shifted from verbose logs and matured into APM and RUM. That has helped us focus on fine-tuning the performance of our applications.

What needs improvement?

We need better visibility into our consumption rate, which is tied to our commit levels. We would love to see a % consumed and alert us if we are over budget before getting an overage charge 20 days into the month.

The biggest complaint we hear comes from the cost of the tool. It is pretty easy to accidentally consume a lot of extra data. Unless you watch everything come in almost daily, you could be in for a big surprise. 

We utilize the Datadog estimated usage metrics to build out alerts and dashboards. The usage and cost system page still doesn't tie into our committed spending - it would be wonderful to see the monthly burn rate on any given day.

For how long have I used the solution?

I've used the solution for six years.

What do I think about the stability of the solution?

There have not been as many outages in the past year. We also haven't been jumping into the new features as quickly as they come out. We may be working on more stable products.

What do I think about the scalability of the solution?

It has scaled up to meet our needs pretty well. Over the years, we have only managed to trigger internal DataDog alerts once or twice by misconfiguring a metric and spiralling out of control with costs.

How are customer service and support?

Support has been lacking. Opening a chat with the tech support rep of the day is always a gamble. We are looking into working with third-party support because it has been so rough over the years.

How would you rate customer service and support?

Neutral

Which solution did I use previously and why did I switch?

We used the ELK stack for logging and monitoring and AppDynamics for APM.

How was the initial setup?

The initial setup for new teams has become easier over the years. We are increasing our adoption rate as we shift our technology to more cloud-native tools. Datadog has supported easy implementation by simply adding a package to the app. 

They have really focused on a lot of out-of-the-box functionality, but the real fun happens as you dive deeper into the configuration. We have also begun adapting open telemetry standards. This has kept us from going too deep into vendor-specific implementations.

What about the implementation team?

We did the initial setup via an in-house team.

What was our ROI?

As long as we stay on top of our consumption mid-month, it has been worth it. However, the few engineers we have who are dedicated to playing whack-a-mole with the growing spending could be better utilized in teaching best practices to new users. I suppose our implementation of the rapidly changing tools over the years has led to a fair amount of technical debt.

What's my experience with pricing, setup cost, and licensing?

It is quite easy to set up any specific tool, but to take advantage of the full visibility it offers, you need to instrument across the board—which can be time-consuming. Be careful about how each tool is billed, and watch your consumption like a hawk.

Which other solutions did I evaluate?

We evaluated AppDynamics and Dynatrace.

What other advice do I have?

It's a very powerful tool, with lots of new features coming, but you certainly will pay for what you get.

Which deployment model are you using for this solution?

Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Victor Chen1 - PeerSpot reviewer
Software Engineer at Zip
Real User
Top 20
Sep 30, 2024
Good for log ingestion and analyzing logs with easy searchability of data
Pros and Cons
  • "The feature I've found most valuable is the log search feature."
  • "More helpful log search keywords/tips would be helpful in improving Datadog's log dashboard."

What is our primary use case?

We use Datadog as our main log ingestion source, and Datadog is one of the first places we go to for analyzing logs. 

This is especially true for cases of debugging, monitoring, and alerting on errors and incidents, as we use traffic logs from K8s, Amazon Web Services, and many other services at our company to Datadog. In addition, many products and teams at our company have dashboards for monitoring statistics (sometimes based on these logs directly, other times we set queries for these metrics) to alert us if there are any errors or health issues.

How has it helped my organization?

Overall, at my company, Datadog has made it easy to search for and look up logs at an impressively quick search rate over a large amount of logs. 

It seamlessly allows you to set up monitoring and alerting directly from log queries which is convenient and helps for a good user experience, and while there is a bit of a learning curve, given enough time a majority of my company now uses Datadog as the first place to check when there are errors or bugs. 

However, the cost aspect of Datadog is tricky to gauge because it's related to usage, and thus, it is hard to tell the relative value of Datadog year to year.

What is most valuable?

The feature I've found most valuable is the log search feature. It's set up with our ingestion to be a quick one-stop shop, is reliable and quick, and seamlessly integrates into building custom monitors and alerts based on log volume and timeframes. 

As a result, it's easy to leverage this to triage bugs and errors, since we can pinpoint the logs around the time that they occur and get metadata/context around the issue. This is the main feature that I use the most in my workflow with Datadog to help debug and triage issues.

What needs improvement?

More helpful log search keywords/tips would be helpful in improving Datadog's log dashboard. I recently struggled a lot to parse text from raw line logs that didn't seem to match directly with facets. There should be smart searching capabilities. However, it's not intuitive to learn how to leverage them, and instead had to resort to a Python script to do some simple regex parsing (I was trying to parse "file:folder/*/*" from the logs and yet didn't seem to be able to do this in Datadog, maybe I'm just not familiar enough with the logs but didn't seem to easily find resources on how to do this either). 

For how long have I used the solution?

I've used the solution for 10 months.

What's my experience with pricing, setup cost, and licensing?

Beware that the cost will fluctuate (and it often only gets more expensive very quickly).

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Tejaswini A - PeerSpot reviewer
Application Engineer at Discover Financial Services
Real User
Jul 3, 2024
Consolidates all our logs into a single place, making it easier to find errors
Pros and Cons
  • "The best way it has helped us is by consolidating all our logs into a single place and making it easier to find errors."
  • "Another issue that I have is with the search syntax, it could be simpler and it feels like there are too many ways to do the same things."

What is our primary use case?

We have a tech stack including all backend services written in TS/Node (mostly) and as a full stack engineer, it is crucial to keep track of new and existing errors. Our logs have been consolidated in Datadog and are accessible for search and review, so the service has become a daily tool for my work. 

More recently, session replay has been adopted at my company, but I do not like it so much because the UI elements are not in their place, so it is very hard to see what the users on the web app are actually clicking on.

How has it helped my organization?

The best way it has helped us is by consolidating all our logs into a single place and making it easier to find errors. Previously using AWS Cloudwatch was cumbersome and time-consuming. One issue I do have with logs is the length of time they are on the platform. Some issues happen sporadically, so it would be good to have logs for longer than one month by default or make it a configuration. 

Another issue that I have is with the search syntax, it could be simpler and it feels like there are too many ways to do the same things.

What is most valuable?

Logs search is the most valuable feature because it has consolidated all of our backend services logs into one place. Now we can see the relationship between them as requests are going from one service to other dependencies. 

What needs improvement?

One issue I do have with logs is the length of time they are on the platform. Some issues happen sporadically, so it would be good to have logs for longer than one month by default or make it a configuration. I have yet to try rehydrating logs, so this might be an option I need to try. Another issue I have is with the search syntax, it could be simpler. The syntax is a bit cumbersome and there is not an intuitive to save them to look for similar searches in the future. 

Finally, while my company replaced a different tool for session replay with DataDog's version, I find it clunky and in need of further improvements. For example, when troubleshooting a web portal issue, it is super important to know what the user clicked, but the elements are not where they should be in the replay.

It is also hard to find details about the sessions, and metadata such as user email, account, etc. that exist on other services with replay features.

For how long have I used the solution?

I have been using Datadof for approximately five years.

What do I think about the stability of the solution?

So far we haven't had any issues with uptime and Datadog has been available when needed.

What do I think about the scalability of the solution?

It seems to scale well as we continue to add services that need monitoring.

How are customer service and support?

I haven't had to contact support.

Which solution did I use previously and why did I switch?

Cloudwatch was not a great tool for what we need to do to troubleshoot issues.

What about the implementation team?

We deployed it in-house with intermediate expertise.

What was our ROI?

I am not sure how much we are paying, but I use the app often enough to feel like we are getting a good ROI.

Which other solutions did I evaluate?

I was not involved in the choosing process as a software engineer

Which deployment model are you using for this solution?

Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
reviewer2561892 - PeerSpot reviewer
Principal. Performance Engineering at Invitation Homes
User
Top 20
Oct 2, 2024
A go-to tool for analyzing, understanding, and investigating application performance
Pros and Cons
  • "Log analytics give us a powerful mechanism for error tracking, research, and analysis."
  • "Network device and performance monitoring could be improved, as we've faced some limitations in this area."

What is our primary use case?

The soluton is used for full stack enterprise performance monitoring for our primarily cloud-based stack on AWS. We have implemented monitoring coverage using RUM for critical apps and websites and utilize APM (integrated with RUM) for full stack traceability.  

We use Datadog as our primary log repository for all apps and platforms, and the advanced log analytics enable accurate log-based monitoring/alerting and investigations. 

Additionally, we some advanced RUM capabilities and metrics to track and optimize client-side user experience. We track SLO's for our critical apps and platforms using Datadog.

How has it helped my organization?

We now have full-stack observability, which allows us to better understand application behavior, quickly alert users about issues, and proactively manage application performance.  

We've seen value by implementing observability coordinated across multiple applications, allowing us to track things like customer shopping and orders across multiple applications and services.  

For critical application launches, we've built dashboards that can track user activity and confirm users are able to successfully utilize new features, tracking user activities in real-time in a war-room situation.  

Datadog is our go-to tool for analyzing, understanding, and investigating application performance and behavior.

What is most valuable?

APM accurately tracks our service performance across our ecosystem. RUM gives us client-side performance and user experience visibility, and the rate of new features implemented in the Digital Experience area recently has been high. Log analytics give us a powerful mechanism for error tracking, research, and analysis.  

Custom metrics that we've created allow us to track KPIs in real-time on dashboards. All of these have proven valuable in our organization.  Additionally, Datadog product support teams are responsive and have provided timely support when needed.

What needs improvement?

Agent remote configuration should be provided/improved and streamlined, allowing for config changes/upgrades to be performed via the portal instead of at the host.   

Cost tracking via the admin portal is a bit lacking, even though it has gotten better.  I'm looking for usage trends (that drive cost) across time and better visibility or notifications about on-demand charges.  

Network device and performance monitoring could be improved, as we've faced some limitations in this area.  

The Datadog usage-based cost model, while giving us better transparency, is difficult to follow at times and is constantly evolving.  

For how long have I used the solution?

I've used the solution for three years.

How are customer service and support?

Support has been responsive and helpful.  

How would you rate customer service and support?

Positive

What's my experience with pricing, setup cost, and licensing?

Pricing is straightforward. That said, it's sometimes difficult to estimate usage volumes.

Which other solutions did I evaluate?

We evaluated Datadog and New Relic in detail and chose Datadog due to their straightforward and competitive pricing model, and their full coverage of monitoring features that we desired, and an easy-to-use UI.  

Which deployment model are you using for this solution?

Public Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
SecOps Engineer at Ava Labs
User
Top 20
Sep 30, 2024
Helpful support, with centralized pipeline tracking and error logging
Pros and Cons
  • "Real user monitoring gives us invaluable insights into actual user experiences, helping us prioritize improvements where they matter most."
  • "While the documentation is very good, there are areas that need a lot of focus to pick up on the key details."

What is our primary use case?

Our primary use case is custom and vendor-supplied web application log aggregation, performance tracing and alerting. 

How has it helped my organization?

Through the use of Datadog across all of our apps, we were able to consolidate a number of alerting and error-tracking apps, and Datadog ties them all together in cohesive dashboards. 

What is most valuable?

The centralized pipeline tracking and error logging provide a comprehensive view of our development and deployment processes, making it much easier to identify and resolve issues quickly. 

Synthetic testing is great, allowing us to catch potential problems before they impact real users. Real user monitoring gives us invaluable insights into actual user experiences, helping us prioritize improvements where they matter most. And the ability to create custom dashboards has been incredibly useful, allowing us to visualize key metrics and KPIs in a way that makes sense for different teams and stakeholders. 

What needs improvement?

While the documentation is very good, there are areas that need a lot of focus to pick up on the key details. In some cases the screenshots don't match the text when updates are made. 

I spent longer than I should trying to figure out how to correlate logs to traces, mostly related to environmental variables.

For how long have I used the solution?

I've used the solution for about three years.

What do I think about the stability of the solution?

We have been impressed with the uptime.

What do I think about the scalability of the solution?

It's scalable and customizable. 

How are customer service and support?

Support is helpful. They help us tune our committed costs and alert us when we start spending out of the on-demand budget.

Which solution did I use previously and why did I switch?

We used a mix of SolarWinds, UptimeRobot, and GitHub actions. We switched to find one platform that could give deep app visibility.

How was the initial setup?

Setup is generally simple. .NET Profiling of IIS and aligning logs to traces and profiles was a challenge.

What about the implementation team?

We implemented the solution in-house.

What was our ROI?

There has been significant time saved by the development team in terms of assessing bugs and performance issues.

What's my experience with pricing, setup cost, and licensing?

I'd advise others to set up live trials to asses cost scaling. Small decisions around how monitors are used can have big impacts on cost scaling. 

Which other solutions did I evaluate?

NewRelic was considered. LogicMonitor was chosen over Datadog for our network and campus server management use cases.

What other advice do I have?

We are excited to dig further into the new offerings around LLM and continue to grow our footprint in Datadog. 

Which deployment model are you using for this solution?

Hybrid Cloud

If public cloud, private cloud, or hybrid cloud, which cloud provider do you use?

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Operations Manager at TodayTix
User
Top 20
Sep 30, 2024
Good dashboards, easy troubleshooting, and integrations
Pros and Cons
  • "The dashboards are super convenient to us for a more zoomed out view of what is going on with each integration that we utilize."
  • "There could be more easily identifiable documentation on how to find different things on the platform."

What is our primary use case?

We utilize Datadog mainly to monitor our API integrations and all of the inventory that comes in from our API partners. Each event has its own ID, so we can trace all activity related to each event and troubleshoot where needed.

How has it helped my organization?

Datadog gives non-dev teams insights as to what all is happening with a particular event as well as flags any errors so that we can troubleshoot more efficiently.

What is most valuable?

The dashboards are super convenient to us for a more zoomed out view of what is going on with each integration that we utilize.

What needs improvement?

There could be more easily identifiable documentation on how to find different things on the platform. It can be overwhelming at first glance, and it's hard to find appropriate documentation on the site to lead you to where you need to be. 

For how long have I used the solution?

I've used the solution for about 1.5 years.

Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Felix Flores - PeerSpot reviewer
Staff Engineer at a tech services company with 1,001-5,000 employees
Real User
Oct 30, 2022
Great distributed tracing and flame graphs for debugging with a relatively painless setup
Pros and Cons
  • "We like the distributed tracing and flame graphs for debugging. This has been invaluable for us during periods of high traffic or red alert conditions."
  • "Both the engineering team and the product team are seeing tremendous value from this solution."
  • "Once Datadog has gained wide adoption, it can often be overwhelming to both know and understand where to go to find answers to questions."

What is our primary use case?

We are using a mixture of on-prem and cloud solutions to bridge the gap with healthcare entities in the service of providing patients with the medication they need to live healthy lives.

Since we're a heavily regulated company, a lot of our solutions grew from on-premises monoliths. However, as we scaled out, it became harder and harder to move forward with that architecture. Today, we're investing heavily in transforming our systems from monoliths into distributed systems.

With this change in mind, the ability for us to connect the dots using Datadog has been invaluable.

How has it helped my organization?

We have an API that serves as a critical aspect of our system for generating new requests for us to process in service of a patient. This service has many tentacles, and it was always hard to track down how issues from this API are affecting things downstream. Since we've added more instrumentation in this API, Datadog has changed our status from a reactive posture to a proactive one.

It has also served as a prime example to other applications on what the benefit of a well-instrumented system is for that application and other applications around it. Due to this, more and more people are using Datadog.

What is most valuable?

We like the distributed tracing and flame graphs for debugging. This has been invaluable for us during periods of high traffic or red alert conditions. It has also informed our developers on how our various systems are interconnected and the downstream effects of the problems we might encounter for certain services.

We're still working on getting widespread adoption of these products. Still, we're already seeing a shift in the developer's perspective from application-specific and starting to look at things from a more holistic systems perspective.

While this is not part of the question, this is relevant: Now that I've learned more about RUM, this will be something that we will heavily leverage moving forward to give us a whole complete view of our system from the front and back end perspective.

What needs improvement?

Once Datadog has gained wide adoption, it can often be overwhelming to both know and understand where to go to find answers to questions. Currently, we use a combination of documentation and COPs to ensure that folks know how to leverage what we have in Datadog properly.

While the guides for Datadog go a long way, a way to customize the user experience from "advanced" to "novice" mode would go a long way.

For how long have I used the solution?

I've been using the solution for two years.

What do I think about the stability of the solution?

It has never failed us and therefore I consider it to be very stable.

What do I think about the scalability of the solution?

It's magic. For the most part, we just installed the product and a lot of it just worked out of the box.

How are customer service and support?

Technical support is excellent.

How would you rate customer service and support?

Positive

Which solution did I use previously and why did I switch?

We have used Splunk, Sentry, and a suite of hand-made solutions. We switched since the Datadog solution was both comprehensive and cohesive. It was also easier to onboard people since the solution was well-documented and standardized.

How was the initial setup?

For the most part, it was really painless to set up.

What about the implementation team?

We implemented the solution in-house.

What was our ROI?

We're still early on in our transformation process. That said, we are gaining a lot of steam in terms of adoption. Both the engineering team and the product team are seeing tremendous value from this solution.

What's my experience with pricing, setup cost, and licensing?


Which other solutions did I evaluate?


What other advice do I have?

Adding more tooltips and links to documentation or how-tos within the application would really go a long way for those trying to get their feet wet with Datadog.

Which deployment model are you using for this solution?

Hybrid Cloud
Disclosure: My company does not have a business relationship with this vendor other than being a customer.
PeerSpot user
Buyer's Guide
Download our free Datadog Report and get advice and tips from experienced pros sharing their opinions.
Updated: September 2026
Buyer's Guide
Download our free Datadog Report and get advice and tips from experienced pros sharing their opinions.