UDAF aggregate method is called for each value
Hi,
I am creating a UDAF that does approximate distinct values counting (hyperloglog).
What this algorithm does is create a 4K memory from a list of values (regardless of the list size). This 4K memory block can be used to estimate the number of distinct values in the original list. In addition, it can be merged with other 4K memory blocks, so that you can estimate the count of distinct elements of several groups.
I've implemented a c++ implementation of this algorithm, and successfully used it both as a mysql plugin and as a Hadoop HIVE and pig UDFs, that worked great performance wise.
I am trying to use the same code as a vertica UDAF. It works very inefficiently.
What I am seeing is that the AggregateFunction.aggregate method is called for each single value. What this does is create a 4K memory block for each value, and later the combine method combine all these 4K blocks into a single block and then estimated.
This can be much more efficient if multiple values would be passed to the aggregate method (which according to the documentation it should). I am not sure why it behaves this way.
I am using community edition on a single node in an ubuntu VM.
Would really appreciate any advice on how to resolve this
Thanks
Amir
I am creating a UDAF that does approximate distinct values counting (hyperloglog).
What this algorithm does is create a 4K memory from a list of values (regardless of the list size). This 4K memory block can be used to estimate the number of distinct values in the original list. In addition, it can be merged with other 4K memory blocks, so that you can estimate the count of distinct elements of several groups.
I've implemented a c++ implementation of this algorithm, and successfully used it both as a mysql plugin and as a Hadoop HIVE and pig UDFs, that worked great performance wise.
I am trying to use the same code as a vertica UDAF. It works very inefficiently.
What I am seeing is that the AggregateFunction.aggregate method is called for each single value. What this does is create a 4K memory block for each value, and later the combine method combine all these 4K blocks into a single block and then estimated.
This can be much more efficient if multiple values would be passed to the aggregate method (which according to the documentation it should). I am not sure why it behaves this way.
I am using community edition on a single node in an ubuntu VM.
Would really appreciate any advice on how to resolve this
Thanks
Amir
Sign up
Already have an account? Login
Welcome to the Rocket Forum!
Please log in or register:
Employee Login | Registration Member Login | RegistrationEnter your E-mail address. We'll send you an e-mail with instructions to reset your password.