Fastest Compression Type and Export Strategy
Moshe Goldberg, OpenText Vertica Principal Solutions Consultant
Abstract
Exporting large tables from Vertica to the Parquet format enables efficient analytics, backups, and cross-environment portability. Parquet’s columnar layout, embedded metadata, and built-in compression make it ideal for such tasks. A major benefit of Parquet exports is that they can be reloaded into any Vertica cluster, regardless of node count, version, deployment model (EON or Enterprise), or environment (on-premises or cloud).
However, export performance can vary widely depending on compression type, parallelism level, and storage target. In this benchmark, we identify the most efficient export strategies using Vertica's EXPORT TO PARQUET command.
Key finding: ZSTD compression consistently delivers top-tier performance - often surpassing even uncompressed exports in speed - while significantly reducing file size.
Why Export to Parquet?
• Efficient Storage: Columnar compression reduces file size significantly.
• Portability: Exported files can be restored on any Vertica setup.
• Open Format: Compatible with many data processing frameworks.
• Future-Proofing: Parquet files retain metadata and can be inspected or reloaded later.
However, exporting large tables to Parquet can be time-consuming, and the process must be manually repeated per table.
Benchmark Objective
This benchmark aims to identify the fastest method to export a large, partitioned Vertica ROS table to Parquet.
The focus is not on querying the resulting Parquet files, but purely on export speed.
A key insight from Vertica’s documentation is that “If you omit the OVER() clause (i.e., no partitioning), Vertica optimizes for maximum parallelism during export.”
We took advantage of this by avoiding partitioning in the export query, maximizing performance.
Benchmark Setup
• Source table: 100M-row table T1, partitioned by a partition_field column (formatted as YYYYMMDD)
• Export targets: Local directories (/tmp/DELME) on each node
• Compression types tested: Snappy, GZIP, Brotli, ZSTD, and Uncompressed
• Parallelism levels tested: 1 and 2
• Metrics captured: Rows exported, disk usage, and duration (request_duration_ms)
Additionally, though not shown in the script output, we observed that exporting to a local directory is nearly 2x faster than exporting to an NFS mount.
The benchmark was executed on Vertica Analytic Database version 24.2.0-1.
Monitoring During the Benchmark
To ensure accurate measurement of export performance, we collected metrics from Vertica system tables while and after each export run. These metrics include:
• Total number of rows exported (per file and overall)
• File size on disk
• Export duration in milliseconds
• Input rows processed (from v_monitor.query_consumption)
• Request duration (from v_monitor.query_requests)
• File-level details (from v_monitor.udx_events)
The following system query was used to gather per-file export metadata, including row counts and file sizes:
SELECT file, SUM(rows) AS exported_rows
FROM v_monitor.udx_events
WHERE udx_name ILIKE 'ParquetExport%'
GROUP BY file;
file | exported_rows
-----------------------------------------------------------------------------------+--------------
/tmp/DELME_PAR1_GRP_1_vHZzJ5Qe/11a316cb-v_eevdb_node0004-139860456752896-0.parquet | 9996781
/tmp/DELME_PAR1_GRP_1_l7zwnxIx/db6eb3fd-v_eevdb_node0003-139859908126464-0.parquet | 9996781
/tmp/DELME_PAR2_GRP_1_91Yf1Pbn/614b5d84-v_eevdb_node0005-140297388521216-0.parquet | 5001375
/tmp/DELME_PAR1_GRP_1_oJ0pOcKL/dbb6aad1-v_eevdb_node0001-139860786149120-0.parquet | 9996781
/tmp/DELME_PAR4_GRP_2_CmuGnZQb/3b66099f-v_eevdb_node0002-140296282986240-0.parquet | 144128
To track export volume over time or analyze performance trends, we also grouped by timestamp:
SELECT created as 'Create Time', SUM(rows) AS exported_rows
FROM v_monitor.udx_events
WHERE udx_name ILIKE 'ParquetExport%'
GROUP BY created
ORDER BY created DESC;
Create Time | exported_rows
------------------------------+--------------
2025-03-31 01:08:34.720853+03 | 4996848
2025-03-31 01:08:34.718646+03 | 4998557
2025-03-31 01:08:34.717313+03 | 5000019
[..]
These monitoring queries provided consistent validation that EXPORT TO PARQUET completed successfully and allowed us to cross-reference the number of exported rows against the input data. This also helped in diagnosing edge cases, such as files with zero rows or unexpected gaps.
The v_monitor.udx_events system table includes not only the number of rows written and file size, but also metadata such as session ID, user, node name, export timestamps (created and closed), and row group counts.
For example:
SELECT report_time, node_name, session_id, user_id, user_name, transaction_id, statement_id, request_id, udx_name,
file, created, closed, rows, row_groups, size_mb
FROM v_monitor.udx_events
WHERE udx_name ilike 'ParquetExport%'
and file::varchar ilike '%333b4292-v_eevdb_node0001-139682225256192.parquet%';
report_time | 2025-03-26 22:43:44.998832+02
node_name | v_eevdb_node0001
session_id | v_eevdb_node0001-2432254:0x2435d3
user_id | 45035996273704962
user_name | dbadmin
transaction_id | 45035996275914370
statement_id | 1
request_id | 0
udx_name | ParquetExport
file | /tmp/DELME_group_1_hkehCTtY/333b4292-v_eevdb_node0001-139682225256192.parquet
created | 2025-03-26 22:43:44.998554+02
closed | 2025-03-26 22:43:44.998816+02
rows | 3
row_groups | 1
size_mb | 0.000998
Benchmark Results

Key Findings
• ZSTD is the best overall performer, balancing speed and compression:
- ~82% faster than GZIP with similar compression ratio.
- Surprisingly, ZSTD is sometimes faster than Uncompressed:
- At parallelism=2, ZSTD took 7.0s vs. Uncompressed 6.5s — only a 7.7% difference,
while ZSTD compressed files were ~5.5× smaller.
• Snappy (the default) offers moderate compression and good speed.
• GZIP and Brotli provide good compression, but are significantly slower (GZIP up to 6× slower than ZSTD).
• Parallelism (x2) improved performance across all compression types, while increasing it further had minimal impact on export duration.
However, in larger clusters with more CPU cores and memory, higher levels of parallelism for EXPORT TO PARQUET may yield additional performance benefits.
• Export to local disk is almost 2x faster than to NFS mounts (observed outside of script output).

Conclusion
If your goal is fast and space-efficient exports to Parquet in Vertica, ZSTD is the clear winner. It combines compression quality close to GZIP with performance exceeding even uncompressed exports. For large-scale exports, avoid the OVER() clause to allow Vertica to leverage maximum parallelism, and prefer local storage over NFS mounts for best performance.
This benchmark provides a repeatable and automated approach to evaluate export performance, making it easy to adapt to your environment or integrate into maintenance workflows.
To highlight the impact of our benchmark findings, let’s look at a real-world example where exporting a large Vertica table to Parquet using GZIP compression, single-threaded parallelism (x1), and writing to an NFS mount takes 10 hours.
By switching to ZSTD compression (which is approximately 82% faster than GZIP),
Increasing to x2 parallelism (providing around 20% speed improvement),
And exporting to a local disk (nearly 2× faster than NFS),
The total export time was reduced to just 43 minutes.

This example clearly shows how we can dramatically accelerate export performance in Vertica. In addition to compression type, parallelism level, and storage target, several other important factors can significantly influence export performance and will be evaluated in our upcoming benchmarks. One such factor is data type optimization - for instance, using INTEGER columns instead of VARCHAR can reduce both serialization overhead and Parquet file size, resulting in faster exports.
Another key factor is cluster-wide parallel execution. Although the EXPORT TO PARQUET command is initiated from a single node, Vertica intelligently distributes the export workload across all nodes in the cluster. This behavior is similar to how Vertica’s COPY statement parallelizes bulk data loading. Increasing the number of nodes enhances both export and load performance by enabling more parallel execution - each node contributes to the work, scaling out compute, memory, and I/O resources for significantly faster processing.
Performance Disclaimer
This benchmark was conducted on a small virtual machine in a non-production environment. The results may not reflect the full performance potential of Vertica in a properly tuned production setup with high I/O bandwidth and larger clusters.
Performance may vary depending on several factors, including the number of columns, data types, field sizes, and your specific data compression ratio.
To determine the optimal configuration for your environment—based on your actual data characteristics and cluster size—it is recommended to run the provided scripts in a QA or staging environment.
