Skip to main content

Spark to Vertica [w/o using HDFS]

  • July 4, 2017
  • 2 replies
  • 17 views

MAW
Forum|alt.badge.img+1
  • Participating Frequently

Colleagues,

Apologies if this sounds a rather dumb question, but if one can use the generic Spark JDBC connector to write directly to Vertica (accepting its limitations), I was wondering why do we have to “…have an HDFS cluster for an intermediate staging location [to save data from Spark to Vertica] when using our Vertica Connector for Apache Spark/JDBC client library?

Is it not possible to stage the data outside HDFS (when using our Connector / JDBC)?

Regards
Mark

2 replies

dcanadillas
Forum|alt.badge.img+1
  • Participating Frequently
  • July 4, 2017

Hi Mark,

Spark doesn't have a distributed storage system (it only has the compute distribution), so it usually relies on Hadoop (HDFS and Yarn) for this. I think you can use S3 instead of Hadoop when using Spark, but you don't have all the functionality.

I suppose that the reason to use HDFS when integrating Vertica and Spark is just because of the use of a DFS for taking advantadge of all Spark features. And it is the most common use of Spark.

I suppose that using the JDBC connector we would deliver the distributed system for the structured and semi-structured data resident in the DB, and not for the Spark Python or Scala custom development (But UDx's in Vertica are distributed like a DFS in the cluster).

Anyway, I think it is a really good question, because I am wondering if we could develop some complete specific integration for this.

Let's open the question!!

Regards,
David.


aurorak
Forum|alt.badge.img
  • Participating Frequently
  • July 6, 2017