Skip to main content
Question

best practice for OS patching on-prem Vertica clusters?

  • April 19, 2024
  • 3 replies
  • 6 views

dgrumann
Forum|alt.badge.img+2

We have encountered situations with ITOM customers where they bring down one node of a large vertica cluster for patching the Linux OS but then it takes days for Vertica to recover that node when it comes back up. What do vertica experts advice for preparing and patching Linux Vertica nodes efficiently?

3 replies

mosheg
Forum|alt.badge.img+2
  • Participating Frequently
  • April 21, 2024
  1. Having fewer than 1024 ROS containers per projection allows Vertica to maintain optimal recovery performance.

  2. Try to hold or reduce ETL activity while the EE node is down.

  3. If you have a large database that contains a single large table with two projections, and with default settings,
    and recovery is taking too long, give recovery more memory to improve speed,
    by setting PLANNEDCONCURRENCY and MAXCONCURRENCY in the RECOVERY pool to 1,
    so recovery can take as much memory as possible from the GENERAL pool and run only one thread at once.
    See: https://www.vertica.com/docs/10.1.x/HTML/Content/Authoring/AdministratorsGuide/ResourceManager/Scenarios/TuningForRecovery.htm

  4. Check:
    tail -f vertica.log | egrep '(RECOVER_ERROR)|(incrCatchUpFailureCount)|(CatchUp)|(RecoverTable)'

If it shows the same table recovering all the time:
try copy_table so it can create a new one and then drop the old one.
Once complete restart recovery again on that node.
Or do recovery from scratch on that node.

  1. In general, prefer Vertica EON mode architecture to overcome the node recovery time issue.

dgrumann
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • April 22, 2024

Hi Moshe thanks for the response. I think segmentation of our ITOM tables has been carefully planned and tested we do not allow customers to do tuning (that I am aware of) which could reduce their ROS containers overall. There are hundreds of tables used by ITOM software. ETL (actually microbatch ingestion) is continuous and while there is some buffering, no customer would be willing to disable ingestion for days just to patch a vertica node (and of course all nodes in the cluster need patching). The problem with tuning memory for recovery is that memory may already be tight on nodes already and so making ingestion fall behind in favor of recovery would be a delicate balance. ITOM doesn't support Eon mode at this time. We have advised customers to execute select make_ahm_now() to reduce delete vectors needing recovery, but other than that perhaps our customers just need to live with the long recovery after node patching.


dgrumann
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • May 8, 2024

OK so to summarize - ITOM customers can not typically afford to stop ingestion for upgrades, and if they don't then the only "trick" we have to help them with patching recovery is to have them run select make_ahm_now() before they begin.

We also have customers who originally deployed ITOM with Vertica versions 9, 10, and 11 on CentOS 7.9 and they need to move to RHEL 8.x/9.x. We are assuming they can upgrade a cluster one node at a time, by first upgrading their vertica cluster to a version of Vertica that supports both CentOS 7.9 and RHEL 8.x, and then taking one node at a time out of the cluster, re-creating that VM/node on a new OS base like RHEL 8.8, and then adding it back to the Vertica cluster. Once that one node is added and the cluster rebalancing finishes, they would move on to the next node until all the VMs in the cluster are on RHEL 8.8.

Does anyone on this forum have experience with a customer who has done this successfully on a production-type on-prem enterprise environment? Is there a better way to safely move off CentOS 7.9?