Stopping a single node using stop_node or kill_node should not cause a full cluster outage, especially in a large cluster (~166 nodes). If the entire cluster goes down, it is likely due to a script sequencing or timing issue during S3 token renewal rather than a Vertica fault.
Environment
Product: Vertica Analytics Platform
Mode: Eon (communal storage on S3)
Situation
Customer observed that stopping one node as part of a scripted token renewal process caused the entire cluster to crash.
Expected behavior: cluster remains operational when one node is stopped.
Cause
Likely a script sequencing/timing issue:
- The script may proceed before the node has fully restarted and rejoined the cluster.
- Fixed timers instead of status checks can lead to premature operations, destabilizing the cluster.