We run a three node cluster with Vertica 7.2.0-1, and since november 7th our Last Good Epoch / Ancient History Mark has not advanced.
Current AHM Time: 2016-11-07 15:36:48.442813+00
I followed the steps in this guide: https://my.vertica.com/hpe-vertica-troubleshooting-checklists/ancient-history-mark-not-advancing/ , but this did not have any effect.
Looking at the active events table (SELECT * FROM active_events;) mentions a stale checkpoint on one of the nodes from around the same time the AHM stopped progressing:
2016-11-07 17:40:50 2016-11-29 10:40:14 Stale Checkpoint Node v_node0002 has data 3841 seconds old which has not yet made it to disk. (LGE: 26549170 Last: 26549664)
From the logs I can find one error from just before the AHM stopped on november 7th:
2016-11-07 15:01:20.923 TM Moveout:0x7f6a5cd6cd90 <ERROR> @v_byk_dwh_node0002: {threadShim} 55V03/5157: Unavailable: [Txn 0xb000002e05f0a3] Moveout of mart_sales_assistant.snowplow_events_b0 - timeout error Timed out T locking Table:mart_sales_assistant.snowplow_events. X held by [user etl_user
From other posts (https://community.dev.hpe.com/t5/Vertica-Forum/Stale-checkpoint-and-too-many-ROS/td-p/221273, https://community.dev.hpe.com/t5/Vertica-Forum/get-last-good-epoch-not-advancing-so-can-t-drop-old-projections/td-p/210267) I've gathered the problem might be that data that is stuck in WOS which prevents the LGE from advancing, or there are too many ROS containers. We've fixed a lot of issues with ROS containers mainly by repartitioning tables. Running a Move Out manually still does not seem to solve anything.
Further details:
MoveOutInterval: 200
MergeOutInterval: 200
HistoryRetentionTime: 0
wosdata resource pool: memorysize: 2G, maxmemorysize: 2G
All help is appreciated.