Skip to main content

Data Loading and configuration parameters

  • December 1, 2014
  • 1 reply
  • 11 views

Adam_1
Forum|alt.badge.img
  • Participating Frequently
I'd like to load large files (from 1GB up to 100GB) and I'm wondering if there are any configuration parameters that can accelerate the data loading. Do you have any recommendations?

For example, I found that MaxAutoSegColumns parameter is set to 32 by default and don't really know if I can change it because 0 indicates to use all columns in the hash segmentation expression.

1 reply

Prasanta_Pal
Forum|alt.badge.img
  • Participating Frequently
  • December 2, 2014
I think 32 columns are enough to create a unique record and good segmentation. More columns you do include in the segmentation, the more time required for data loading unless it is not require. Segmentation is required to evenly distribute the data across all the nodes.

You may refer to the below link for data loading a faster way.
Using Parallel Load Streams
http://my.vertica.com/docs/7.1.x/HTML/index.htm#Authoring/AdministratorsGuide/BulkLoadCOPY/UsingPara...