Skip to main content

projection design consideration for load high volume of data in a cluster

  • July 12, 2017
  • 2 replies
  • 12 views

hoseiney
Forum|alt.badge.img

we have a system that load high volume of data every minute using Kafka and Vertica
what is the best practices and important consideration for designing projection for this case ?

2 replies

sKwa
Forum|alt.badge.img+1
  • Participating Frequently
  • July 12, 2017

Hi!

From my experience:

  • proper denormalization (not every denormalization is good)
  • manually created projections that gives to your queries required performance (use DBD only for encoding only)

PS
Sorry, "silver bullet" not exists(imho).


TomM
Forum|alt.badge.img
  • Participating Frequently
  • July 12, 2017

The general rule is to stand up your cluster, load a good amount of data, and then run the DBD on the entire dataset along with a representative set of queries. You can then edit the DDL as you see fit or just accept the entire recommendation. Projection design is an iterative process; it's something you'll want to re-visit over time as more data, more users, and new queries are added to your database.