Need help in debugging an issue with Vertica on AWS.We have set up a 3 instance cluster on AWS with t2 large across all nodes while trying to partition the table facing an issue that shuts down Vertica database abruptly, there are no significant messages recorded on /var/log/messages file.
messages file around the shutdown on
Node-1
Nov 9 17:21:04 ip-10.0.0.1 su: (to dbadmin) ec2-user on pts/1
Nov 9 17:21:45 ip- 10.0.0.1 dhclient[3217]: XMT: Solicit on eth0, interval 108330ms.
Nov 9 17:23:34 ip- 10.0.0.1 dhclient[3217]: XMT: Solicit on eth0, interval 120900ms.
Nov 9 17:25:35 ip- 10.0.0.1 dhclient[3217]: XMT: Solicit on eth0, interval 113820ms.
Nov 9 17:27:29 ip- 10.0.0.1 dhclient[3217]: XMT: Solicit on eth0, interval 117800ms.
Nov 9 17:29:27 ip- 10.0.0.1 dhclient[3217]: XMT: Solicit on eth0, interval 119340ms.
Nov 9 17:30:01 ip- 10.0.0.1 systemd: Created slice User Slice of root.
Node-2
Nov 9 17:20:01 ip-10.0.0.2 systemd: Created slice User Slice of root.
Nov 9 17:20:01 ip-10.0.0.2 systemd: Starting User Slice of root.
Nov 9 17:20:01 ip-10.0.0.2 systemd: Started Session 51 of user root.
Nov 9 17:20:01 ip-10.0.0.2 systemd: Starting Session 51 of user root.
Nov 9 17:20:01 ip-10.0.0.2 systemd: Removed slice User Slice of root.
Nov 9 17:20:01 ip-10.0.0.2 systemd: Stopping User Slice of root.
Nov 9 17:20:20 ip-10.0.0.2 dhclient[3189]: XMT: Solicit on eth0, interval 110860ms.
Nov 9 17:22:03 ip-10.0.0.2 systemd-logind: Removed session 40.
Node-3
Nov 9 17:20:01 ip-10.0.0.3 systemd: Stopping User Slice of root.
Nov 9 17:21:14 ip-10.0.0.3 systemd-logind: Removed session 40.
Nov 9 17:21:14 ip-10.0.0.3 systemd: Removed slice User Slice of dbadmin.
Nov 9 17:21:14 ip-10.0.0.3 systemd: Stopping User Slice of dbadmin.
Instance Type Diskspace Swap Size ulimit -n
Node-1(Initiator) t2 large 130 GB 10 GB 65536
Node-2 t2 large 30 GB 10 GB 65536
Node-3 t2 large 30 GB 10 GB 65536
The error is similar for 1GB of data and 10 GB while partitioning any table
records from DC_ERRORS table for that session
/SELECT event_timestamp,node_name,user_name,session_id,error_level,error_code,message,hint
FROM error_messages
where session_id = 'v_tpcds_db2_node0001-31783:0x40b';/
event_timestamp node_name user_name session_id error_level error_code message hint
09-11-2020 17:21 v_tpcds_db2_node0001 dbadmin v_tpcds_db2_node0001-31783:0x40b NOTICE 0 The new partitioning scheme will produce partitions in 72 physical storage containers per projection
09-11-2020 17:21 v_tpcds_db2_node0001 dbadmin v_tpcds_db2_node0001-31783:0x40b WARNING 64 Queries using table "web_returns" may not perform optimally since the data may not be repartitioned in accordance with the new partition expression Use "ALTER TABLE tpc1gb.web_returns REORGANIZE;" to repartition the data.
In the comment down will add the message from the vertica.log file on Node-1 during table partition .