So, this is less of a question, and more of a log of some research and (hopefully) resolution to an interesting problem we seem to be encountering in a POC at Finicity.
So, Finicity is doing a POC whereby they are extracting data from Redshift, unloading it into s3 (using Redshift's UNLOAD command) and then (attempting) to load the data from s3 into Vertica using COPY.
The initial problem was around Vertica's COPY statement. When running the COPY command in Vertica it generated an error that said "Access Denied". In this particular case, it was referencing a file (0020) out of 80 separate files in the s3 bucket. By default, Redshift extracts files in parallel, and creates many files per table.
This led us to wonder if there was some specific issue with file 0020. To test this, we used the AWS CLI to run a few simple tests.
I could ls the file correctly. But I could not copy it. The syntax is something like:
aws s3 cp s3:\bucket\file .
(Also, for the record, the vertica.log showed that every file in the set were generating the same error. Management console apparently just reports on the last file in the set, so it's a little confusing.)
That produced this error:
fatal error: An error occurred (403) when calling the HeadObject operation: Forbidden
Which led us to this blog:
http://www.stojanveselinovski.com/blog/2016/05/20/aws-s3-object-acl-and-403-error/
That suggested that the issue had to do with which AWS_PROFILE we were using. AWS_PROFILE is a convenient way of defining the AWS ID and KEY.
We then created a file manually (in Linux) and uploaded it to the s3 bucket, using the CLI. Then we copied that file into a small, simple two-column table. That worked perfectly fine.
At first we decided this must be a Redshift issue. The initial thinking was that Redshift was encrypting files during the extract, but that doesn't appear to be the case. While it does appear to be true that Redshift CAN encrypt data during the unload step, it does not appear to be the default behavior.
We also ruled out the bucket itself. The same files placed into different buckets also don't work. So, our current conclusion is that it's the ROLE that Redshift is using to extract the files, that's causing the issue.
We're still investigating. I'll post an update once we figure out a resolution.