Skip to main content
Question

Adding Node to existing node cluster

  • December 24, 2019
  • 6 replies
  • 16 views

TarunKumar

Hi-
We have 3 node cluster.We have added another node.It took 5 hour to sysnc everything. If I add another 2 nodes.How much time it will take to perform the same activity. 5 hour or 10 hours ? Is there way we can faster the whole process. Data is around 1 TB

6 replies

Bryan_H
Forum|alt.badge.img+2
  • Participating Frequently
  • December 24, 2019

What version of Vertica are you running? We strongly recommend upgrading to 9.2.1 or later to take advantage of faster catalog and Tuple Mover operations.
What is the size of the REFRESH resource pool? This pool is used to handle sync and update between nodes so we recommend to add memory to this pool to improve refresh performance. It's recommended to reset REFRESH pool to the original memory size after completing the refresh / sync operation.


hsaxena20
Forum|alt.badge.img+1
  • Participating Frequently
  • December 25, 2019

I have one more question:
My database has more segmented projection as compared to unsegmented projection. Is rebalancing will take more time?


Bryan_H
Forum|alt.badge.img+2
  • Participating Frequently
  • December 25, 2019

It will depend on a lot of factors, such as size and complexity of the segmented tables. However, segmented tables need to be rebalanced across all nodes, while unsegmented tables need only copy to new nodes without change (* if the DDL specifies UNSEGMENTED ALL NODES), so in general it will take longer to process segmented tables. It is important to ensure cluster is running latest Vertica version and also to increase REFRESH resource pool to allow more resources to be available to process segmented tables.


hsaxena20
Forum|alt.badge.img+1
  • Participating Frequently
  • December 25, 2019

Hi Bryan,
By adding the nodes to cluster, will query performance improve?


Bryan_H
Forum|alt.badge.img+2
  • Participating Frequently
  • December 25, 2019

Adding compute resources often helps, but other options should be considered.
Have you reviewed data types and encoding for tables and projections? For example, to reduce memory usage, VARCHAR should be the minimum width needed to hold all data, and consider using INTEGER type instead or NUMERIC or FLOAT since INTEGER can usually be processed in a single processor cycle where other types usually require additional FPU cycles. Also check encoding, which will reduce I/O wait times, see the blog post for example: https://www.vertica.com/blog/checking-and-improving-column-compression-and-encoding/
Have you reviewed the explain plan for any long-running query? Explain plan may show steps like resegment that can be improved by projection tuning, and can also identify complex queries that use views or CTE's that can be optimized. Please see the tuning guide with hints for specific operations such as JOIN, GROUP BY, ORDER BY at https://www.vertica.com/docs/9.2.x/HTML/Content/Authoring/AnalyzingData/Optimizations/OptimizingQueryPerformance.htm


Jim_Knicely
Forum|alt.badge.img+2
  • Participating Frequently
  • December 25, 2019