Skip to main content

HDFS storage locations and High Availability

  • July 11, 2017
  • 2 replies
  • 17 views

dcanadillas
Forum|alt.badge.img+1

Hi all,

Maybe is a dumb question, but I am guessing how our high availability works when we are using HDFS storage in SQL for Hadoop. As I understand, when we store ROS containers in HDFS we are replicating depending on the K-Safety as we would do with Linux FS, right? But because how HDFS works, the containers could be split in different HDFS files and Hadoop nodes, so I am not sure how that affects to our segmentation and replication estrategy in terms of high availability. How can Vertica assure that the data is segmented properly then between nodes?

And Hadoop has its own replication strategy as well, so could we say that Vertica is always being capable of finding projections data because it is going to be available in Hadoop nodes (regarding that the name node is replicated as well, of course)?

Thank you!
David.

2 replies

Car1os
Forum|alt.badge.img
  • Participating Frequently
  • July 11, 2017

With KSafety=1 you will end up with 6 copies of the data. By default HDFS will store 3 copies.


dcanadillas
Forum|alt.badge.img+1
  • Author
  • Participating Frequently
  • July 13, 2017
Thanks Car1os!!