Do we have a script or best practice to determine failed node root cause?
If not, please help to improve the following..
1. Is there a core file? --> cat /proc/sys/kernel/core_pattern
2. If core file found,
a. Determine what RPM was being used? --> strings your_corefile | grep BrandId
b. Get current value for LD_PRELOAD from vertica.log on all nodes.
c. Generate GDB full backtrace from the core file.
To generate the backtrace, gdb tool is needed --> gdb /opt/vertica/bin/vertica /data/abrt/core_file_name
Once you are in the gdb prompt run the below command
bt full
Copy the stack generated to a file and upload to Support or RnD
bt --> copy the stack generated to a file and upload
exit or quit
3. Get your vertica.log path
/opt/vertica/bin/admintools -t list_db -d admintools -t list_allnodes | grep ' UP ' | awk '{print $9}' | grep "Database Log" | awk '{print $4}'
4. How many restarts happened?
The first occurrence of “INFO New log” most likely generated during a log rotation.
Start the count after that.
grep -n "INFO New log" vertica.log
2:2018-09-16 03:05:06.167 INFO New log # Generated during a log rotation
3680379:2018-09-16 07:58:02.230 INFO New log # 1
3714589:2018-09-16 08:31:06.205 INFO New log # 2
3746948:2018-09-16 08:53:40.280 INFO New log # 3
7019632:2018-09-16 12:54:20.519 INFO New log # 4
7026132:2018-09-16 13:05:11.785 INFO New log # 5
7032455:2018-09-16 13:11:33.510 INFO New log # 6
7038737:2018-09-16 13:22:26.543 INFO New log # 7
7442661:2018-09-16 14:55:56.079 INFO New log # 8 restarts happened between 2018-09-16 03:05:06 and 14:55:56
nano +7442661 vertica.log # use the nano or vi text editor to examine what happened
during last restart. Instruct nano to go to line 7442661
Why was the node(s) down?
Scan vertica.log for messages with high severities, like: , and severities.
Also, the strings “Shutdown” and “shutting down” will point you towards parts of the log you are interested in.
grep -n -i -e 'FATAL' -e '' -e '' -e 'signal' -e 'select shutdown' vertica.logWhich node left first?
grep -n -e "left the cluster" -e 'Starting up Vertica Analytic Database' vertica.log | grep -A 2 -e 'Starting up Vertica Analytic Database'