Skip to main content

Collation and performances

  • September 14, 2017
  • 3 replies
  • 14 views

Bernard92
Forum|alt.badge.img+2

Hi,
One of our customer is using CHAR and collation=fr_FR@colstrength=3 , and find that performances are not as good as when they use VARCHAR and the default collation.
I can't really see the difference , but is there any reason why there would be one ?
And if there's one , can we measure it ?
Regards
Bernard

3 replies

Moshe
Forum|alt.badge.img
  • Participating Frequently
  • September 14, 2017

Vertica default locale is en_US@collation=binary, which uses binary collation.
Binary collation is faster as V compares binary representations of strings.
When locale collation is non-binary, the GROUP BY on string data will make V to call COLLATION even if the function is not specified in the query.
The reason for performance degradation is the sorting time, because V sorting depends on LOCALE.


Bernard92
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • September 15, 2017

Many thanks Moshe
Have you got any idea of the overhead ? I mean in term of magnitude : is it twice slower or just adds a few microsecond to each query ?


Moshe
Forum|alt.badge.img
  • Participating Frequently
  • September 17, 2017

Depends on the use-case:
The amount of rows, predicate columns number and VARCHAR(s) variable size.