How do you verify that an eland-imported tree model matches the original?

Hi,

Laura Trotta at Elastic suggested I post this here.

When eland imports an XGBoost, LightGBM or scikit-learn model into Elasticsearch, what Elasticsearch runs is a converted copy of the trained model. I am curious how people check that the copy gives the same predictions as the original. A comparison on a test set can miss differences that only show up near a split threshold, at exactly zero, or when a feature is missing. I have seen that kind of silent mismatch in other converters, for example XGBoost missing value handling in the OpenSearch learning to rank plugin.

Has anyone run into eland imports that did not match the original model? And are there known differences in how missing values or thresholds are handled compared with XGBoost or LightGBM?

For disclosure, I am building a tool called Leafparity that checks a tree model against its converted copy for every possible input. Today it works with ONNX only, not with Elasticsearch, so it does not help here yet. I would only build support for eland models if people in this forum actually need it.

Milivoje

OpenSearch/OpenDistro are AWS run products and differ from the original Elasticsearch and Kibana products that Elastic builds and maintains. You may need to contact them directly for further assistance. See What is OpenSearch and the OpenSearch Dashboard? | Elastic for more details.

(This is an automated response from your friendly Elastic bot. Please report this post if you have any suggestions or concerns :elasticheart: )