I'm trying to load a CSV table into Vertica from S3 (Vertica 25.3 on RHEL9.6 as a single node instance)
It has just 4 million rows and is 238MB gzipped standard compression, so tiny by Vertica standards.
But the COPY statement below takes hours to complete, far longer to execute than I would expect.
Specifically loading the data is quick, but the commit takes most of the time.
If i monitor progress by `select session_id, table_name, (load_duration_ms/1000)::int as time, accepted_row_count,rejected_row_count, unsorted_row_count, sorted_row_count, parse_complete_percent FROM LOAD_STREAMS;` i can see that it takes just about 10 minutes to read the data and complete parsing, maxing out at `accepted_row_count` is 4,017,771 which is the correct row count and `parse_complete_percent` is 100.
But then it takes over two hours until sorted_row_count starts incrementing, then a while after that to complete. I've tried defining the table with no `ORDER BY` clause and a simple `ORDER BY id` but that makes no difference. Is there a way to speed up the commit phase?
```sql
COPY myschema.mytable
FROM 's3://mybucket.mycompany.com/path/to/project/datafiles/raw/mydata.csv.gz'
ON ANY NODE GZIP
PARSER fcsvparser(type='traditional', header=true)
REJECTED DATA AS TABLE myschema.mytable_rejects
REJECTMAX 10
```