This guide provides a series of simple, copy-paste-friendly tests to help you find and fix network problems.
Common Symptoms of Network Trouble
You might have a network issue if you see errors like these in your logs:
-
Log Files to Check:
-
/catalog-path/database-name/node-name_catalog/vertica.log -
/catalog-path/database-name/dbLog -
/catalog-path/database-name/YourNodeName_catalog/startup.log
-
-
Common Error Messages:
-
SP_connect: unable to connect via UNIX socket... Error: Connection refused -
Waiting for cluster invite -
SP_connect: unable to send connect handshake: Transport endpoint is not connected
-
Let’s get testing!
Prerequisite Test: Verify Passwordless SSH
Vertica superuser (dbadmin) rely on passwordless SSH to communicate between nodes. This is the first thing you should check.
-
From one node, try to SSH to another node in the cluster and run a simple command like
hostname.ssh node2_private_ip 'hostname' -
Success: The command runs immediately and prints the other node's hostname without asking for a password.
-
Failure: You are prompted for a password, or the connection times out.
Test 1: Check Basic Port Connectivity with nc
This test uses the netcat (nc) tool to see if a direct communication channel can be opened between two nodes on a specific port.
Note: Vertica and its spread daemon must be shut down for this test. Otherwise, the ports will already be in use, and the test will fail.
-
On one Vertica node (the "listener"), open a terminal and run this command:
nc -l 5433Note: Ensure that your firewall (or cloud security group) allows inbound TCP traffic on that port from the connecting node’s IP.
If blocked, the test will fail with a “Connection timed out” message even if the network is otherwise functional. -
On a different node in the cluster (the "connector"), run this command, using the private IP of the listening node:
nc <The_listening_node_private_IP> 5433 -
Success: Both terminals will appear to freeze. This indicates the connection is open. You can type a message in one window, press Enter, and it will appear in the other.
-
Failure: The connecting node will immediately return an error like "Connection refused" or "Connection timed out."
Repeat this test for all key Vertica ports, such as 5433, 4803, and 4804.
Test 2: Check Network Integrity with ping
This tests basic reachability using a larger-than-normal packet size, which can help detect issues with network hardware or MTU settings.
-
Run this command between all nodes in the cluster (e.g., from Node 1 to Node 2, Node 1 to Node 3, etc.).
ping -s 4096 <private_network_node_ip>
Test 3: Measure Network Speed with vnetperf
Vertica provides its own tool to measure the actual data transfer speed between nodes. For good performance, you want to see speeds above 100 MB/s on your private network.
-
Run this command, listing the private IPs of all your cluster nodes:
/opt/vertica/bin/vnetperf --hosts node1_ip,node2_ip,node3_ip
Test 4: Inspect Firewall and Switch Settings
External factors like firewalls and network switch settings are common culprits.
Firewalls / Security Groups: Check your cloud security groups (AWS, GCP, Azure) or on-premise firewalls (iptables). Ensure that TCP and UDP traffic is allowed between all cluster nodes on all necessary Vertica ports.
Switch Configuration (Flow Control): Flow control is a network switch feature that can harm Vertica's performance. It must be OFF.
-
Connect to your network switch's command line.
-
Run the command to show interface status (the exact command may vary by vendor).
show interfaces brief -
Look for the Flow Ctrl column. It must say off.
-
GOOD Example:
Port Type Enabled Status Mode Flow Ctrl ------- --------- ------- ------ ------- --------- A1-Trk6 10GbE-SR Yes Up 10GigFD off A2-Trk6 10GbE-SR Yes Up 10GigFD off
-
-
If flow control is enabled, disable it on all cluster-facing ports.
Test 5: Ensure Network Interfaces Connect Automatically
If a node reboots and its network card doesn't come online automatically, it won't be able to join the cluster.
-
Check the status of network devices:
nmcli d -
To fix this, use the simple text-based menu tool as root:
sudo nmtui -
In the menu, navigate to Edit a connection. Select your private network interface, and make sure the [ ] Automatically connect checkbox is checked (use the spacebar to check it). Select OK to save.
Test 6: Check for Latency Spikes with fping
Sometimes, the network isn't slow all the time, but has brief "spikes" of high latency that can disrupt the cluster. fping is great for finding these.
How to Install fping (if needed): If the fping command isn't found, you may need to install it.
# For RHEL/CentOS systems sudo yum groupinstall "Development Tools" -y sudo yum install wget -y
# Download and compile wget https://fping.org/dist/fping-5.1.tar.gz tar xzf fping-5.1.tar.gz cd fping-5.1 ./configure make sudo make install
Running the Latency Test: Run this from one node (as root) to all other private IPs in the cluster. It will continuously check the latency and add a timestamp, making it easy to see when a spike occurred. Press Ctrl+C to stop it.
sudo fping -D -l -b 1472 host1_ip host2_ip host3_ip
