We are starting a POC on MapR with Parquet files. We set up a MapR cluster with Vertica and generated Parquet files with TPC-DS data generator writing to MapR-FS and mounted the Parquet files as external tables in Hive/Tez and Vertica. When running side-by-side tests, we found that the performance increase with Vertica was not as expected. I checked the explain plan and found that Vertica is running external table queries on the initiator node only. Now, this is good since Hive/Tez uses the entire cluster to get similar performance. However, we'd like to show that Vertica can be faster to access Parquet files on MapR-FS.
What are current best practices for setting up Parquet external tables on MapR-FS? There is a 3 node cluster in Vertica VPN if you'd like to look at our setup.