Skip to main content

Problem when trying to write from spark into vertica

  • July 17, 2017
  • 2 replies
  • 14 views

DieterC
Forum|alt.badge.img+2

Hi forum,

a Vertica prospect has started testing Spark/Vertica connectivity.
They are encountering below error when trying to write from spark into vertica.

Please advise on how to resolve this.
Thank you
Dieter

Hi,
Can anyone help us, we are encountering below error, we are trying to write from spark into vertica, but its not working.
◄ import org.apache.spark.sql.SaveMode
val df = spark.read.orc("hdfs:///ameda/*.orc")
df.createOrReplaceTempView("tabel1")
val df2 = spark.sql("WITH tmp1 AS(SELECT ts, Value>5 AS cond, LAG(Value>5) OVER (PARTITION BY mea_id ORDER BY ts) AS prev_cond, LEAD(ts) OVER (PARTITION BY mea_id ORDER BY ts) AS next_ts, mea_id, Odo, fleet FROM tabel1), tmp2 AS (SELECT ts, LEAD(ts) OVER (PARTITION BY mea_id ORDER BY ts) AS next_ts, LEAD(Odo) OVER (PARTITION BY mea_id ORDER BY ts) AS next_Odo, cond, Odo, mea_id, fleet FROM tmp1 WHERE cond <> prev_cond OR next_ts IS NULL) SELECT ts AS start_ts, next_ts AS end_ts, (next_Odo - Odo) AS distance, mea_id, fleet FROM tmp2 WHERE cond AND next_ts IS NOT NULL");val table = "test_iav_controller.result_5"
val hdfs_url="hdfs:///user/livy/vertica/"
val web_hdfs_url="webhdfs:///user/livy/vertica"
val db = "iav"
val user = "dbadmin"
val password = "iav"
val host = "10.128.5.202"
val port = "3306";
val opt = Map("host" -> host, "table" -> table, "db" -> db, "port" -> port, "user" -> user, "password" -> password, "hdfs_url" -> hdfs_url, "web_hdfs_url" -> web_hdfs_url)
df2.write.format("com.vertica.spark.datasource.DefaultSource").options(opt).mode(SaveMode.Overwrite).save();
17:01:40

WARNUNG: Statement failed: Error: java.lang.Exception: S2V: FATAL ERROR for job S2V_job1628864483610041422. Job status information is available in the Vertica table public.S2V_JOB_STATUS_USER_DBADMIN. Unable to save intermediate orc files to HDFS path:hdfs:///user/livy/vertica/S2V_job1628864483610041422. Error message:org.apache.spark.SparkException: Job aborted.: [ at com.vertica.spark.s2v.S2V.do2Stage(S2V.scala:99)
, at com.vertica.spark.s2v.S2V.save(S2V.scala:389)
, at com.vertica.spark.datasource.DefaultSource.createRelation(VerticaSource.scala:88)
, at org.apache.spark.sql.execution.datasources.DataSource.write(DataSource.scala:518)
, at org.apache.spark.sql.DataFrameWriter.save(DataFrameWriter.scala:215)
, ... 47 elided]

2 replies

aurorak
Forum|alt.badge.img
  • Participating Frequently
  • July 17, 2017

I don't have any experience with the vertica spark connector but it looks like an issue about permissions. If I were to try and debug this I would give the hdfs path world write permissions just for test purposes


Connyrt-emp1
Forum|alt.badge.img
  • Participating Frequently
  • July 17, 2017

Hi Dieter,
Try writing directly to your HDFS path using Spark's API:
df2.write.orc(hdfs_path)
Resolve any issues and try again using Vertica's Spark connector.