Skip to main content

Does Vertica trust all Parquet file statistics?

  • May 14, 2018
  • 0 replies
  • 3 views

mosheg
Forum|alt.badge.img+2
  • Participating Frequently

Each Parquet file has a footer that stores codecs, encoding information, as well as column-level statistics, e.g., the minimum and maximum number of column values.

To accommodate Uber data’s size (over 5PB), Uber created a new open source Parquet reader for Presto. Uber new reader implements four optimisations which speed up querying from 2-10x faster compared to when they used the original open source reader.
See: “Engineering Data Analytics with Presto and Parquet at Uber”
By Zhenxiao Luo, published on July 11, 2017 at https://eng.uber.com/presto

Does Vertica trust all Parquet file statistics?
Does Vertica employs performance optimisations like Uber open source Parquet reader does for Presto?