Senior Data Engineer at a tech vendor with 10,001+ employees
Real User
Top 10
Jun 22, 2026
I would like to add that for the connectors, there is sometimes limited support for using wildcards to get the items or assets ingested from sources like S3; it does not support very good wildcard filters. Additionally, Data Hub has a problem with column-level lineage support, especially regarding non-pro users or those without any plans. If I talk about the free features of Data Hub open source, those two I found could be improved during my use case. Regarding improvements needed for Data Hub, I have already mentioned the limitations on the usage of wildcards in the ingestion or connectors; that can be worked upon, especially regarding the open-source part of Data Hub. The rest is that I hope the UI is quite good.
We encountered some issues when we wanted to connect our streaming infrastructure to Data Hub, which was somewhat problematic. In our data streaming infrastructure, we had a database CDC'd through Kafka Connect to a Kafka topic, and at the end of the pipeline, it would go to either an OLAP or a data lakehouse. However, the problem with visualizing this data lineage was that while the connection between MySQL and Kafka worked, when we wanted to track data from Kafka to other services, we couldn't track everything back because the IDs were generated randomly and couldn't be connected. We had to fix this manually by stating where the data had gone, which was tedious. Data Hub's GMS service, or General Metadata Service, is a good service that I used regularly, but the CLI version had considerable changes across different versions. When I installed a different version, there wasn't enough consistency to ensure that commands I used would work in future versions of Data Hub's GMS CLI, which was frustrating. I also recall that setting up Kafka without Zookeeper was not possible, which was inconvenient, though I should verify this as I don't remember if they fixed it. At least from my recollection, when I wanted to set it up one and a half years ago, they did not have direct support for KRAFT in their Helm chart.
Lead Business Analyst at a tech vendor with 10,001+ employees
Real User
Top 20
Jun 5, 2026
In my day-to-day work, most consumers raise concerns with respect to the enrichment that we perform on top of the main data set provided by the respective feeding application. This is where most of the work lies for us because consumers come with multiple data-related concerns regarding data quality issues, data mapping issues, or completely missing data during the enrichment process. It is my day-to-day responsibility to ensure that no records are dropped after the main data set is enriched with referentials. On Azure, we use an architecture on Spark that builds this particular enrichment using executors and driver memory. Sometimes when it goes out of memory due to multiple jobs in progress, some records are dropped because the executor is dropped without completing the entire process. A final solution is being built to address this particular use case. On Oracle database, when you needed to create a new attribute or when a specific user needed a particular value available for their reporting use case, whether for regulatory reporting, compliance, audit, or any kind of requirements, we would determine which would be the most appropriate table and try to add this attribute to that particular flow. In this scenario, only the application creating these tables and the user are involved. To ensure that the table does not become bulky and the data inside it is not overloading, such cases are not ensured at all and there is no normalization. When it comes to the Enterprise Data Model, version one was created looking at all the users who currently require this particular data set. Later, when one or multiple users request the addition of a new attribute or addition of a new block of information altogether, a team called the Data Management Office becomes involved. They take care of adding this attribute to the given model, publish a new version, and ensure that the new version created is backward compatible so that downstream impacts are minimized. Data Hub would achieve a perfect rating once all consumers migrate to the Azure database and we address the majority of concerns they have. Since a public data lake is being used on Azure, resources are coming from Azure itself. As a result, we cannot publish any confidential information onto the Azure Data Lake; private spaces have been created for this purpose. Once the public data lake is resilient enough to any online threats or issues, I believe Data Hub using the public data lake will become one of the major use cases where the majority of consumers can adopt it for going forward, instead of using a private lake or the Oracle database, which is entirely private and not hosted on any public platform.
data platform at a tech vendor with 1,001-5,000 employees
Real User
Top 10
May 30, 2026
I think Data Hub can be improved by supporting the open source version better. Many features have moved to the paid version now, making it difficult for small-scale companies to operate on Data Hub because we are required to pay, even though it started as an open source project that is now essentially behind a paywall. One needed improvement for Data Hub would be stronger AI-powered metadata discovery. I understand Data Hub has been investing in AI, but the natural language processing power on Data Hub search is not that good. The search itself is not accurate many times. Another improvement could be enhancing the DBT developer experience, such as surfacing DBT test failures directly in lineage. Additionally, when we change schema, if it could provide a risk scoring of some sort, that would also be beneficial. Lastly, automated cleanup recommendations would help because managing orphan data assets on Data Hub currently takes a lot of manual time.
In terms of improvements for Data Hub, it seems more useful for critical or large data pipelines, as small data architectures can be straightforward to understand without it. Regarding enhancements for complex projects, I have noticed that sometimes Data Hub does not provide a complete picture of the lineage, particularly in complex data pipelines such as when we fetch data from an API to S3 and subsequently to Snowflake. We have to review the metadata in Data Hub closely.
I know that the integrations are not easy to do, and I believe it happens because it's a customized solution. There always needs to be software developers to work on this. It's complicated; every time we want to integrate new things or new sources, we need to generate a ticket or a request to another department. When I had my experience with Atlan, for example, I was able to connect different sources in a very user-friendly way. I just needed to set up some configurations and connect to the source without having to be a software developer or develop any code in the back end. It was just a feature in the data catalog that enabled me to connect with different kinds of sources. That's why I think the disadvantage of having a customized solution. Although I think Data Hub itself is a very good tool, years ago I had the opportunity to work with it, but with a clear interface and the open-source solution, which was very clear and easy to connect. At Uber, we need to have a request when we want to integrate new sources. Regarding Data Hub's intuitiveness, regarding analytics, I would say that some quality dimensions are available for us. For example, for each field name or each column in a table, it's possible to see the frequency, how many values we have for a specific type or category, and we can see if there are new or null values, whether the columns are empty or not, along with some metrics. This is regarding the data quality dimensions, such as nullables and things of that nature. That is all we have for features. I remember when I was working with Atlan, there was a feature I liked very much—the possibility to have a sample. When I clicked on a table, I could see a short sample without needing SQL skills. I just clicked the table and could see some values or what the table represents; the data catalog would show a screen with some rows of the table. This feature was very good, but we don't have it in Data Hub the way it is implemented at Uber. I think it would be a very good feature for analytics, and we don't have it at the moment. The integration part could be better, but again, it's because it's a customized solution. I think if they used the native version of the tool, it would be simpler. The integration part and the process of setting up new data quality rules would be important for data governance players like me.
Data Hub can be improved with more automation; there are some inbuilt automations, such as documenting definitions of data elements using AI, which is useful. I wonder if it can automate the classification exercise, possibly using AI to auto-classify PII direct and indirect items.
The impact is very positive, and there are many benefits for us using Data Hub because it was easier to make data governance, create centralized metadata management, improve data discoverability, and manage data in general. The areas for improvement, in my opinion, are the initial setup and configuration that can be complex without prior experience, especially in large-scale environments. User experience for non-technical users could be further simplified, particularly around advanced metadata concepts. The out-of-the-box governance workflow, for example, approvals and certification, could be more prescriptive for customers at early maturity stages. Data Hub can be improved in the initial setup and configuration that is somewhat complex, and also in operational monitoring that could benefit from more native dashboards and alerts. However, these are not blockers, but areas where additional guidance or product enhancement would further accelerate adoption.
The product cannot be improved in just one area. There are no points in support or documentation that require improvement. There are no improvements needed for Acryl Data that I have not mentioned yet.
Data Hub is an advanced platform designed to streamline data management processes, enhance data accessibility, and provide comprehensive analytics capabilities for informed decision-making. Data Hub offers a unified approach to handling large-scale datasets, empowering organizations to effectively manage, analyze, and extract insights from their data infrastructure. It provides robust features for data integration, storage, and visualization, supporting diverse business needs and driving...
Data Hub can be improved with easy accessibility. I think integration with other environments is needed to enhance accessibility.
I would like to add that for the connectors, there is sometimes limited support for using wildcards to get the items or assets ingested from sources like S3; it does not support very good wildcard filters. Additionally, Data Hub has a problem with column-level lineage support, especially regarding non-pro users or those without any plans. If I talk about the free features of Data Hub open source, those two I found could be improved during my use case. Regarding improvements needed for Data Hub, I have already mentioned the limitations on the usage of wildcards in the ingestion or connectors; that can be worked upon, especially regarding the open-source part of Data Hub. The rest is that I hope the UI is quite good.
We encountered some issues when we wanted to connect our streaming infrastructure to Data Hub, which was somewhat problematic. In our data streaming infrastructure, we had a database CDC'd through Kafka Connect to a Kafka topic, and at the end of the pipeline, it would go to either an OLAP or a data lakehouse. However, the problem with visualizing this data lineage was that while the connection between MySQL and Kafka worked, when we wanted to track data from Kafka to other services, we couldn't track everything back because the IDs were generated randomly and couldn't be connected. We had to fix this manually by stating where the data had gone, which was tedious. Data Hub's GMS service, or General Metadata Service, is a good service that I used regularly, but the CLI version had considerable changes across different versions. When I installed a different version, there wasn't enough consistency to ensure that commands I used would work in future versions of Data Hub's GMS CLI, which was frustrating. I also recall that setting up Kafka without Zookeeper was not possible, which was inconvenient, though I should verify this as I don't remember if they fixed it. At least from my recollection, when I wanted to set it up one and a half years ago, they did not have direct support for KRAFT in their Helm chart.
In my day-to-day work, most consumers raise concerns with respect to the enrichment that we perform on top of the main data set provided by the respective feeding application. This is where most of the work lies for us because consumers come with multiple data-related concerns regarding data quality issues, data mapping issues, or completely missing data during the enrichment process. It is my day-to-day responsibility to ensure that no records are dropped after the main data set is enriched with referentials. On Azure, we use an architecture on Spark that builds this particular enrichment using executors and driver memory. Sometimes when it goes out of memory due to multiple jobs in progress, some records are dropped because the executor is dropped without completing the entire process. A final solution is being built to address this particular use case. On Oracle database, when you needed to create a new attribute or when a specific user needed a particular value available for their reporting use case, whether for regulatory reporting, compliance, audit, or any kind of requirements, we would determine which would be the most appropriate table and try to add this attribute to that particular flow. In this scenario, only the application creating these tables and the user are involved. To ensure that the table does not become bulky and the data inside it is not overloading, such cases are not ensured at all and there is no normalization. When it comes to the Enterprise Data Model, version one was created looking at all the users who currently require this particular data set. Later, when one or multiple users request the addition of a new attribute or addition of a new block of information altogether, a team called the Data Management Office becomes involved. They take care of adding this attribute to the given model, publish a new version, and ensure that the new version created is backward compatible so that downstream impacts are minimized. Data Hub would achieve a perfect rating once all consumers migrate to the Azure database and we address the majority of concerns they have. Since a public data lake is being used on Azure, resources are coming from Azure itself. As a result, we cannot publish any confidential information onto the Azure Data Lake; private spaces have been created for this purpose. Once the public data lake is resilient enough to any online threats or issues, I believe Data Hub using the public data lake will become one of the major use cases where the majority of consumers can adopt it for going forward, instead of using a private lake or the Oracle database, which is entirely private and not hosted on any public platform.
I think Data Hub can be improved by supporting the open source version better. Many features have moved to the paid version now, making it difficult for small-scale companies to operate on Data Hub because we are required to pay, even though it started as an open source project that is now essentially behind a paywall. One needed improvement for Data Hub would be stronger AI-powered metadata discovery. I understand Data Hub has been investing in AI, but the natural language processing power on Data Hub search is not that good. The search itself is not accurate many times. Another improvement could be enhancing the DBT developer experience, such as surfacing DBT test failures directly in lineage. Additionally, when we change schema, if it could provide a risk scoring of some sort, that would also be beneficial. Lastly, automated cleanup recommendations would help because managing orphan data assets on Data Hub currently takes a lot of manual time.
In terms of improvements for Data Hub, it seems more useful for critical or large data pipelines, as small data architectures can be straightforward to understand without it. Regarding enhancements for complex projects, I have noticed that sometimes Data Hub does not provide a complete picture of the lineage, particularly in complex data pipelines such as when we fetch data from an API to S3 and subsequently to Snowflake. We have to review the metadata in Data Hub closely.
I know that the integrations are not easy to do, and I believe it happens because it's a customized solution. There always needs to be software developers to work on this. It's complicated; every time we want to integrate new things or new sources, we need to generate a ticket or a request to another department. When I had my experience with Atlan, for example, I was able to connect different sources in a very user-friendly way. I just needed to set up some configurations and connect to the source without having to be a software developer or develop any code in the back end. It was just a feature in the data catalog that enabled me to connect with different kinds of sources. That's why I think the disadvantage of having a customized solution. Although I think Data Hub itself is a very good tool, years ago I had the opportunity to work with it, but with a clear interface and the open-source solution, which was very clear and easy to connect. At Uber, we need to have a request when we want to integrate new sources. Regarding Data Hub's intuitiveness, regarding analytics, I would say that some quality dimensions are available for us. For example, for each field name or each column in a table, it's possible to see the frequency, how many values we have for a specific type or category, and we can see if there are new or null values, whether the columns are empty or not, along with some metrics. This is regarding the data quality dimensions, such as nullables and things of that nature. That is all we have for features. I remember when I was working with Atlan, there was a feature I liked very much—the possibility to have a sample. When I clicked on a table, I could see a short sample without needing SQL skills. I just clicked the table and could see some values or what the table represents; the data catalog would show a screen with some rows of the table. This feature was very good, but we don't have it in Data Hub the way it is implemented at Uber. I think it would be a very good feature for analytics, and we don't have it at the moment. The integration part could be better, but again, it's because it's a customized solution. I think if they used the native version of the tool, it would be simpler. The integration part and the process of setting up new data quality rules would be important for data governance players like me.
Data Hub can be improved with more automation; there are some inbuilt automations, such as documenting definitions of data elements using AI, which is useful. I wonder if it can automate the classification exercise, possibly using AI to auto-classify PII direct and indirect items.
The impact is very positive, and there are many benefits for us using Data Hub because it was easier to make data governance, create centralized metadata management, improve data discoverability, and manage data in general. The areas for improvement, in my opinion, are the initial setup and configuration that can be complex without prior experience, especially in large-scale environments. User experience for non-technical users could be further simplified, particularly around advanced metadata concepts. The out-of-the-box governance workflow, for example, approvals and certification, could be more prescriptive for customers at early maturity stages. Data Hub can be improved in the initial setup and configuration that is somewhat complex, and also in operational monitoring that could benefit from more native dashboards and alerts. However, these are not blockers, but areas where additional guidance or product enhancement would further accelerate adoption.
I do not have comments on how Acryl Data can be improved.
The product cannot be improved in just one area. There are no points in support or documentation that require improvement. There are no improvements needed for Acryl Data that I have not mentioned yet.