Just sharing a way to walk around issue "Server not found in Kerberos database (7) - LOOKING_UP_SERVER" when saving Spark DataFrame to kerberized HDFS.
To save Spark DataFrame to kerberized HDFS through our Spark connector, we need kerberos authentication Vertica user to copy orc file from HDFS internally. but we will meet issue "Server not found in Kerberos database (7) - LOOKING_UP_SERVER" according to our SAMPLE CODE in document.
Here is the jira VER-65515 Spark connector can not save DataFrame to kerberized HDFS.
It looks like spark connector connects to Vertica with IP and can not use "KerberosHostname" or other properties of kerberos in options map.
To walk around this issue, we can set ("db" -> "vmart?KerberosHostname=v001.hadoop.com") in options map for com.vertica.spark.datasource.DefaultSource.
By the way, with additional ("debug" -> "true") in options map, you can see more detail info for troubleshooting.
Here is a same code:
// ...
val opts: Map[String, String] = Map(
"debug" -> "true",
"KerberosHostname" -> "v001.hadoop.com",
"table" -> "s2v",
"db" -> "vmart?KerberosHostname=v001.hadoop.com", // the magic
"user" -> "ktuser",
"password" -> "vertica",
"host" -> "v001.hadoop.com",
"port" -> "5433",
"hdfs_url" -> "hdfs://hdp001:25000/tmp"
)
val mode = SaveMode.Overwrite
df.write.format("com.vertica.spark.datasource.DefaultSource").options(opts).mode(mode).save()