Skip to main content
Question

Restart/repair when recovery epoch less than AHM

  • June 13, 2020
  • 1 reply
  • 5 views

Bryan_H
Forum|alt.badge.img+2

Hi, my personal cluster broke as a result of a power failure. I am stuck in a situation where start and restart / ASR fail because the computed recovery epoch is much less than AHM so a valid recovery epoch can't be found. Even starting vertica binary with -S flag fails. Here is what I get in startup.log, and hoping someone can point me to a fix, since even with catalog editor I don't see an issue with epochs:
{
"node" : "v_docker_node0001",
"stage" : "Plan Recovery",
"text" : "Clerk: Aborting recovery: recover epoch is too high (only 10 is possible)",
"timestamp" : "2020-06-13 18:39:49.630"
}
{
"node" : "v_docker_node0001",
"stage" : "Startup Failed, ASR Required",
"text" : "Node Dependencies:\n1 - cnt: 36310\n\n1 - name: v_docker_node0001\nNodes certainly in the cluster:\n\tNode 0(v_docker_node0001), epoch 10\nFilling more nodes to satisfy node dependencies:\nData dependencies fulfilled, remaining nodes LGEs don't matter:\n--",
"timestamp" : "2020-06-13 18:39:49.827"
}
{
"node" : "v_docker_node0001",
"stage" : "Plan Recovery",
"text" : "Waiting for more nodes; cannot ASR to 22166482 with current nodeset: Node Dependencies:\n1 - cnt: 36310\n\n1 - name: v_docker_node0001\nNodes certainly in the cluster:\n\t0 down nodes treated as current\tNode 0(v_docker_node0001), epoch 10\nFilling more nodes to satisfy node dependencies:\nData dependencies fulfilled, remaining nodes LGEs don't matter:\n--",
"timestamp" : "2020-06-13 18:39:52.598"
}

1 reply

Bryan_H
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • June 14, 2020

Follow on to this - I found and followed the directions on Confluence at http://10.10.10.242:8080/pages/viewpage.action?spaceKey=TechSupport&title=Steps+to+follow+when+recovery+in+ASR+mode+is+showing+as+2001-01-01+or+older
but recovery epoch is still bad so I am still not able to restart except in read-only unsafe (-U) mode. Are there any other steps I should try to reset the recovery epoch? Only other issue I see is in vertica.log about missing DFS files:
2020-06-14 08:57:17.482 Recover:0x7ff5dbbff700-a00000052cebd3 [Catalog] Found 46 missing DFS files