Did anybody try to integrate Lucene library with Vertica as shown here?
https://github.com/luceneplusplus/LucenePlusPlus
Did anybody try the following with Vertica?
https://dzone.com/articles/lucene-database-oracle-five
We have a table name rawdata.forecasting_system with 493,900,272 rows.
Each row include a field name pv_userSegmentsInfo_uddIds varchar(10000) with about 200 integers separated with a comma.
All together we have about 50K unique numbers, with returns the total is more than 100 billion appearances as shown here:
SELECT count(words) from
(SELECT v_txtindex.StringTokenizerDelim(pv_userSegmentsInfo_uddIds,',')
OVER (PARTITION BY pv_pageViewKey_pageViewUniqueId ORDER BY pv_pageViewKey_pageViewUniqueId)
FROM rawdata.forecasting_system) foo;
count
--------------
100666594134
(1 row)
Vertica current text index is not fast enough, and its inherent sort restrictions break other fields tuning requirements. Also manual index nor regexp search like the following is not fast
like MEMSQL with Lucene.
select xyz from forecast.forecasting_system
where pv_countryCode = 'US' AND pv_userSegmentsInfo_uddIds like '%,158354,%';
as part of several levels of GBY