We are testing multiple backup and restore scenarios.
In one of the scenario, we are testing to see if the recovery is automatic when we have a 3 node cluster and a datafile gets deleted on one node.
Here is what we did:
1) created a table and noted the timestamp on the datafiles that got created on one node (say node2)
dbadmin=> create table testing_df_delete (n number);
CREATE TABLE
2) removed the files on node 2 which were recently created (4 .fdb and .idx files)3) checked that the table existed on node 1
dbadmin=> select * from testing_df_delete;
n
---
1
1
4) Checked that the table is not there in node 2dbadmin=> select * from testing_df_delete;
ERROR 3413: FileColumnReader: unable to open position index /data/VERLABQA01/v_verlabqa01_node0002_data/201/49539595901177201/49539595901177201_0.pidx: No such file or directory
5) Brought down the node2 by doing a ps -ef |grep vertica and killing the session6) Tried restarting node 2 but it doesnt come up with the error as shown below
[dbadmin@genalblabdb07n2 v_verlabqa01_node0002_catalog]$ tail -f vertica.log
2014-01-13 18:38:27.831 nameless:0x5ee14a0 [Catalog] <WARNING> Error getting size of file [/data/VERLABQA01/v_verlabqa01_node0002_data/201/49539595901125201/49539595901125201_0.fdb]: No such file or directory
2014-01-13 18:38:27.831 nameless:0x5ee14a0 [Catalog] <WARNING> Error getting size of file [/data/VERLABQA01/v_verlabqa01_node0002_data/205/49539595901125205/49539595901125205_0.fdb]: No such file or directory
2014-01-13 18:38:27.833 nameless:0x5ee1960 [Catalog] <WARNING> Error getting size of file [/data/VERLABQA01/v_verlabqa01_node0002_data/201/49539595901177201/49539595901177201_0.fdb]: No such file or directory
2014-01-13 18:38:27.833 nameless:0x5ee1960 [Catalog] <WARNING> Error getting size of file [/data/VERLABQA01/v_verlabqa01_node0002_data/205/49539595901177205/49539595901177205_0.fdb]: No such file or directory
2014-01-13 18:38:27.833 Main:0x5b456d0 [Recover] <INFO> Loading UDx libraries
2014-01-13 18:38:27.833 Main:0x5b456d0 [Recover] <INFO> Setting up UDx pointers
2014-01-13 18:38:27.834 Main:0x5b456d0 <PANIC> @v_verlabqa01_node0002: VX001/2973: Data consistency problems found; startup aborted
HINT: Check that all file systems are properly mounted. Also, the --force option can be used to delete corrupted data and recover from the cluster
LOCATION: mainEntryPoint, /scratch_a/release/vbuild/vertica/Basics/vertica.cpp:1166
2014-01-13 18:38:27.907 Main:0x5b456d0 [Main] <PANIC> Wrote backtrace to ErrorReport.txt
My question is: As this is a 3 node cluster wouldnt the files be created automatically on restart as a recovery process?
Is the solution only to restore the latest backup on node2 and then see if it recovers?
In such cases how do we bring up the cluster again to its working state?
Thanks
SAumya