Each Parquet file has a footer that stores codecs, encoding information, as well as column-level statistics, e.g., the minimum and maximum number of column values.
To accommodate Uber data’s size (over 5PB), Uber created a new open source Parquet reader for Presto. Uber new reader implements four optimisations which speed up querying from 2-10x faster compared to when they used the original open source reader.
See: “Engineering Data Analytics with Presto and Parquet at Uber”
By Zhenxiao Luo, published on July 11, 2017 at https://eng.uber.com/presto
Does Vertica trust all Parquet file statistics?
Does Vertica employs performance optimisations like Uber open source Parquet reader does for Presto?