Skip to main content
Question

Parallel loading from S3 bucket

  • January 28, 2020
  • 1 reply
  • 6 views

Bryan_H
Forum|alt.badge.img+2

Does COPY automatically parallelize loads from S3? If so, what are the parameters to set parallelism (# threads, etc.)? Are there limits for file format? I have a customer wanting to land CSV files on S3 and ingest them in parallel with COPY. Will this "just work" or will they need to split the files and COPY each chunk to a specific node to ensure load balancing?

1 reply

asaeidi
Forum|alt.badge.img
  • Participating Frequently
  • January 29, 2020

Yes, we automatically parallelize loads as much as the resource pools and execution parallelism allows, but we do that in differently for various file types. For delimited files, we can split a single file across the cluster using apportioned load, so each node loads a chunk of the file when it is large. For other files such as compressed ones, we cannot do that so if you are loading a single large compressed file only one node can process that.