Hi, this is connected to SD 02335401 [ACTIVE] : STANDBY host took over for a failed Control Node, but not completely. The issue is that the control node of the fault group experienced a disk failure on a data volume so Vertica won't run but spread does. If they turn off the node, then the entire fault group does a safety shutdown. Customer wants to know, should there have been a failover of spread as well as vertica? And how should they plan to fix since it could take over 24 hours of downtime for data center staff to repair the affected hardware?
Here is the full customer question:
We have a Fault Group with these 4 nodes: node0037 thru node0040.
Friday, Nov 2, Vertica stopped on node0037, which is the control node for the group.
Node0040 is a STANDBY node. 12 hours later, node0040 automatically replaced node0037, began participating in queries.
BUT, spread continued to run on node0037. It was not affected by standby takeover.
On Tuesday, Nov 6 technicians shut down host/node0037 to perform diagnostics. At that, Vertica stopped on nodes 0038, 0039, 0040, safe shutdown because their "control node" was gone.
Only after the host was restarted, and spread restarted, were we able to restart the other 3 hosts.
Host0037 still has hardware issues so that Vertica will not run on it -- it will have to be repaired. But, it seems that we can't stop Linux (ie, spread) on it without causing the other 3 hosts to shut down Vertica.
We need to understand::
- why the cluster was allowed to reach this state (ie, standby took over, but left spread running on the taken-over host?
- what is our path to bringing this fault group/the cluster back to complete health, starting from where we are now.