Hi all,
Maybe is a dumb question, but I am guessing how our high availability works when we are using HDFS storage in SQL for Hadoop. As I understand, when we store ROS containers in HDFS we are replicating depending on the K-Safety as we would do with Linux FS, right? But because how HDFS works, the containers could be split in different HDFS files and Hadoop nodes, so I am not sure how that affects to our segmentation and replication estrategy in terms of high availability. How can Vertica assure that the data is segmented properly then between nodes?
And Hadoop has its own replication strategy as well, so could we say that Vertica is always being capable of finding projections data because it is going to be available in Hadoop nodes (regarding that the name node is replicated as well, of course)?
Thank you!
David.