Skip to main content
Question

How Data colelctor data is purged

  • December 21, 2018
  • 5 replies
  • 18 views

mkheir
Forum|alt.badge.img+1
  • Participating Frequently

Hello,
What is the vertica process that is responsible of data colector tables purge, is the purge job scheduled to run on intervals?

Many thanks.

5 replies

Jim_Knicely
Forum|alt.badge.img+2
  • Participating Frequently
  • December 21, 2018

mkheir
Forum|alt.badge.img+1
  • Author
  • Participating Frequently
  • January 3, 2019

Happy new year to all,

My question was probably not very clear.

Given that:

  • I can specify a limit of dc data stored in memory
  • I can specify a limit of dc data stored in disk
  • Load can be different on different nodes.
  • DC data files (log files under /DataCollector/ like: ..log ) are stored in multiple files per component, in my case I have 4 files per components.

What I’m not sure about is:

  • When the memory limit is reached: what happen? Is data moved to file on disk?
  • When does the DC data purge trigger: When the disk limit is reached? on any node? Or is there a job that check periodically file size.
  • When a purge must happen, are all the 4 files deleted or is it a purge that delete
    only part of a file as necessary.

  • In case of different load across nodes, can it happen that DC data in a file of a node cover a different period than files in another node (assuming that each node store its own data locally).

Many thanks.


DaveT
Forum|alt.badge.img
  • Participating Frequently
  • January 3, 2019

I haven't tested every aspect but he memory log is circular so oldest just starts getting overlaid as new data arrives. The data is written to disk quickly but I believe it is asynchronous so possible but highly unlikely that data is lost to disk unless your memory is significantly undersized. The bigger issue would likely be whether you are allocating enough disk space to get the data copied elsewhere. The data files are handled per file as you were saying so only the oldest file is replaced when the newest files is filled. Yes, files could be different per node and it does depend on activity per node so in some areas you could retain more data on some nodes than others. Load balancing is helpful here.


praveshbhardwaj
Forum|alt.badge.img+1
  • Participating Frequently
  • January 4, 2019

I've had the situation once when Execution_engine_profile data was partially purged for a given transaction, statement. Possibly because profiling data for this transaction was spread across more than one files and one of them were purged.


mkheir
Forum|alt.badge.img+1
  • Author
  • Participating Frequently
  • January 9, 2019

Many thanks for you answers, very helpful.