Skip to main content
Question

Open Database Connectivity (ODBC) error occurred. state: 'HY000'. Native Error Code: 50310

  • March 7, 2019
  • 3 replies
  • 16 views

mosheg
Forum|alt.badge.img+2
  • Participating Frequently

One of our customers is getting the following error when trying to load data from same table to same table in Vertica using SSIS:
“: Open Database Connectivity (ODBC) error occurred. state: 'HY000'. Native Error Code: 50310. [Vertica][Support] (50310) Unrecognized ICU conversion error.”

In other databases I saw in google that this was a problem with the data itself where some field values had problem characters in them, such as ñ and ’ (not ').
Do you think changing the default locale for Vertica will help here?

See also:
https://stackoverflow.com/questions/52966407/verticasupport-50310-unrecognized-icu-conversion-error
https://forum.vertica.com/discussion/220737/unrecognized-icu-conversion-error-using-pyodbc

3 replies

marcothesane
Forum|alt.badge.img+1
  • Participating Frequently
  • March 7, 2019

Hi Moshe -

On ICU conversion, you can get more info here: ...

http://userguide.icu-project.org/conversion

ICU conversion is conversion from Unicode to another encoding, and vice versa.

Changing the default locale can help, but could also just kick off the problem in another place in the client stack. Even between a screen field and the database, for example, although this is highly unprobable in the case of Vertica.

In over 95% of the cases I ran into such a one, the problem was this:

Imagine you want to store the name of the Swedish city Malmö into a VARCHAR field.

Now, imagine you create a VARCHAR(5) to store Malmö - which is 5 letters long.
But ö , in UTF-8, requires 2 bytes:0xC3 and 0xB6. In case of an insert with string data right truncation, only the first byte , 0xC3, ends up at the end of the VARCHAR(5). And 0xC3 is always recognised as the first part of a multibyte UTF-8 character. If nothing follows it, it can't be converted. period.

Make the target field longer. How much longer, really depends, on the language, on the alphabet even more.

To be absolutely sure, I usually double the target string lengths, fill the table, and finally profile the filled table for maximum string lengths, and whether the string fields are single-byte ASCII or not.

Not a stroll in the park, I know ....


s_crossman
Forum|alt.badge.img
  • Participating Frequently
  • March 7, 2019

Hi Moshe,

If Marco's details don't help, note that there is also an open Jira related to ICU Conversion errors. In the jira use case there was a specific table load failing, and by process of elimination I was able to narrow it to the row and column, then parsing the content narrowed to the character (an acute accent 'o'). I was able to create a reproducible case for it and dev has determined for that particular reproducer the issue is in the Simba layer, so they've opened a ticket with Simba. Not sure if your issue is the same or not.


Connyrt-emp1
Forum|alt.badge.img
  • Participating Frequently
  • March 7, 2019

Also, I have seen that error when data is not encoded in UTF8 format, Vertica requires that all data is loaded in UTF8 format. So the table you are trying to move has to be a UTF8 encoded table. If the table is non-utf8 then it should be converted before storing it into the new table.

https://www.vertica.com/docs/8.1.x/HTML/index.htm#Authoring/ConnectingToVertica/ClientODBC/SettingTheLocaleForODBCSessions.htm?Highlight=UCS-2