Skip to main content
Question

How to limit number of concurrent very large grouping yearly mergeouts

  • January 3, 2023
  • 3 replies
  • 5 views

Sergey_Cherepan_1
Forum|alt.badge.img+2

Happy New Year!

And, this is the time of the year when Vertica is doing grouping for yearly partitions.

I found that in one of my smallish Vertica clusters I have a bunch of aborted mergeouts, due to running out of temp disk space.

Investigation shows, there are two largest tables roughly same size, and they undergo merging of 2021 data into single yearly partition grouping.

There is enough temp space, and it is on same large disk as data.

Problem here is that Vertica starts at least 3 very large grouping yearly mergeout in parallel on each node.
While it is enough disk space for single yearly mergeout, database cannot fit temp space for 3 yearly mergeouts running in parallel.

I can limit max size of merged ROS, and yearly mergeouts will go through, resulting into several ROS per year. That would work but would be undesired outcome.

Apart from this very large yearly mergeouts, tuple mover works perfectly well, never had problems on this cluster.

Database easily fit several monthly mergeouts concurrently on each node.

Question: How to limit number and total size of concurrent mergeouts per node for very large mergeouts on data that is outside of active partitions?

(I can suggest MaxTotalMergeoutSize, total size of mergeouts per node cannot go above that)

3 replies

Bryan_H
Forum|alt.badge.img+2
  • Participating Frequently
  • January 4, 2023

Do the mergeouts kick off at a specific time, e.g. MoveOutInterval / MergeOutInterval ? That is, when one fails, does it attempt to resume X seconds or minutes later in line with a scheduled interval?
If so, try to set the intervals to very long timeouts to pause automatic TM actions, then run TM manually on each projection:
SELECT DO_TM_TASK('mergeout'[, '[[database.]schema.]{table | projection} ]');


Sergey_Cherepan_1
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • January 5, 2023

Thanks for advice, that may be would work. But only on weekends when load to tables are not that active.
Reducing mergeout concurrency most likely would cause ROS backpressure, if done during business hours.
I am kinda settled with having more than one ROS per year, that is not very bad.
I believe Vertica need to do something to address issue properly. No more than one concurrent huge mergeout per node!


Sergey_Cherepan_1
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • January 9, 2023

@Hibiki ,

Let me know if I need to open service request. Yes, please proceed and open enhancement report.

When yearly mergeout happens, I can see several deleted temp files in DATA dir. They are growing, and finally fill out all disk space. Next, mergeout process on node get killed with error "cannot allocate 1MB on disk", and disk space being freed. This happens in infinite loop, I observed this behaviour for few days.