Skip to main content
Question

Vertica 30 second DR over multiple data centers

  • July 30, 2021
  • 8 replies
  • 11 views

dgrumann
Forum|alt.badge.img+2

Customer giving a requirement that ITOM-OPTIC Vertica DB be protected by having nodes in 2 geographically dispersed data centers connected via "dark fibre" - low latency high speed network. They require that the the non active site to take over support of the live itomdb within 30 seconds of an issue and then to revert back to the primary in 30 seconds of recovery. I assume other customers have posed similar cross-DC fast disaster recovery questions - what is the Micro Focus recommendation?

8 replies

Jim_Knicely
Forum|alt.badge.img+2
  • Participating Frequently
  • July 30, 2021

"..having nodes in 2 geographically dispersed data centers..."

Bad idea... Atleast for now.


dgrumann
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • July 30, 2021

Jim I am sure you know it is difficult to tell a customer who is buying our product "bad idea". They gave us a requirement of 30 second recovery from a DC failure and we must either say "Yes, Vertica can support that and here is how..." or we must say "I am sorry Vertica can not support 30 second recovery, but here is what we recommend for a Vertica architecture to support the fastest Disaster Recovery possible..." Either way I need some help from a Vertica expert who knows - thanks! And - feel free to contact me outside this forum if you wish.


Jim_Knicely
Forum|alt.badge.img+2
  • Participating Frequently
  • July 30, 2021

I disagree. We are supposed to be trusted advisors. We can deploy a Vertica Cluster anywhere, but the nodes should be network near one another...

Can you have the a Vertica DB nodes be anywhere? Sure.

We have a "smart' client that deployed Vertica across data centers.

Allot of support cases followed…

Bad idea.


dgrumann
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • July 30, 2021

OK Jim I guess you are leaving me with "I am sorry Vertica can not support 30 second recovery across DCs" and will tell them the best DR alternative that ITOM can provide is to use vbr as in https://www.vertica.com/docs/9.2.x/HTML/Content/Authoring/AdministratorsGuide/BackupRestore/CopyingTheDatabaseToAnotherCluster.htm


Jim_Knicely
Forum|alt.badge.img+2
  • Participating Frequently
  • July 30, 2021

You can have millisecond fail over with dual load... :D What database offers 30 second failover? Oh, the one that is already running a DR for you and charging you for it behind the scenes? Vertica can do that with replication for free (well less, infrastructure).


dgrumann
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • July 30, 2021

Jim if you are aware that ITOM OPTIC supports dual-load across multiple Verticas, PLEASE do let me know there is no documentation on that. The OPTIC/COSO uses a microbatch loader with attached pulsar UDX library.

Aside from that, I am confused by your "replication for free" comment can you please explain? Please keep in mind I am not arguing with you, nor am I going to argue with the customer's stated requirement - I just need to know how to best reply to them. I am trying to learn from you and represent Vertica in the best possible light.


Jim_Knicely
Forum|alt.badge.img+2
  • Participating Frequently
  • August 2, 2021

Check out:

Replicating Objects to an Alternate Cluster

I am not sure if "ITOM OPTIC supports dual-load across multiple Verticas". But if using replication, that wouldn't be an issue as the Database itself is "kind of" doing the dual load for you.


dgrumann
Forum|alt.badge.img+2
  • Author
  • Participating Frequently
  • August 6, 2021

FYI I have confirmed that the DR process for ITOM (documented for OpsB at https://docs.microfocus.com/itom/Containerized_Operations_Bridge:2021.05/ConfigureDisasterRecovery ) involves the customer implementing some scripts that end up executing /opt/vertica/bin/vbr.py --task replicate via crontab on a regular basis. The OPTIC ingestion requires many other steps to switch over in the case of DR, but the underlying standby DB sync does use the built-in functions.