# Query's getTermsEnum not executed

**URL:** <https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501>\
**Category:** Elasticsearch\
**Created:** [September 22, 2017, 2:15pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501 "2017-09-22T14:15:36Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![dominik.safaric](https://avatars.discourse-cdn.com/v4/letter/d/b38774/32.png) [@dominik.safaric](https://discuss.elastic.co/u/dominik.safaric)\
**Post date:** [September 22, 2017, 2:15pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/1 "2017-09-22T14:15:36Z")

</div>

I've implemented an ElasticSearch custom Java plugin running an `CustomScoreQuery` instance. The purpose of the plugin is to search for documents whose long field value is less or equal in terms of Hamming distance then the value provided in the query. In turn, the `CustomScoreQuery` instance is instantiated with an `MultiTermQuery` instance. The `MultiTermQuery` subclass searches matching documents using a `long` field. Interestedly, the `protected TermsEnum getTermsEnum(Terms terms, AttributeSource atts) throws IOException` function is never executed. On the other hand, the terms iterator is returned if the property of the field is defined as `text`. In addition, the function is executed when using Solr, which leads to an assumption that the problem might be in the index document type mapping. The function is paramount to searching the index because it implements methods classifying documents as hit or not based on the given criteria.

Based on this, I would appreciate if anyone could explain why the function never gets executes when the field type is long? The field has been set to `not_analyzed` and `index=true`.

Below you may find fragments of the source code.

```
    public final class SimilarityCustomScoreQuery extends CustomScoreQuery {

    private final String queryField;
    private final Inference.Response response;
    private final String scoreField;

    public SimilarityCustomScoreQuery(String queryField, String scoreField, Inference.Response response, int maxDistance) {
        super(new SimilarityQuery(queryField, response, maxDistance));
        this.queryField = queryField;
        this.scoreField = scoreField;
        this.response = response;
    }

    @Override
    protected CustomScoreProvider getCustomScoreProvider(LeafReaderContext context) throws IOException {
        return new SimilarityCustomScoreProvider(context, this.scoreField, this.response);
    }
}

```

The `SimilarityQuery` implementation is as follows:

```
public final class SimilarityQuery extends MultiTermQuery {

    interface Params {
        String MAX_DISTANCE = "max_distance";
    }

    /**
     * The default maximum Hamming distance. 
    public static int MAX_DISTANCE_DEFAULT = 13;

    /**
     * The maximum allowed distance a document is said to be accepted
     * as a search hit or not.
     */
    private final int maxDistance;

    private final long value;

    public SimilarityQuery(String field, Inference.Response response, Integer maxDistance) {
        super(field);
        this.maxDistance = Objects.isNull(maxDistance) ? MAX_DISTANCE_DEFAULT : maxDistance;
        this.value = SimilarityQuery.value(response);
    }

    @Override
    // NOTE: the getTermsEnum is never executed when the field type within the index mapping is long.
    protected TermsEnum getTermsEnum(Terms terms, AttributeSource atts) throws IOException {
        return new SimilarityTermsEnum(terms.iterator(), this.value, this.maxDistance);
    }

    @Override
    public String toString(String field) {
        return String.format("%s:%d", this.field, this.value);
    }

    @Override
    public int hashCode() {
        final int prime = 31;
        int hashCode = prime * super.hashCode() + this.maxDistance;
        if (this.field != null) {
            hashCode = prime * hashCode + this.field.hashCode();
        }
        return hashCode;
    }
}
```

---

<div class="post-metadata">

**Author:** ![jimczi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jimczi/32/47985_2.png) [@jimczi](https://discuss.elastic.co/u/jimczi)\
**Post date:** [September 29, 2017, 9:40am UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/2 "2017-09-29T09:40:44Z")

</div>

The `long` type in Elasticsearch is mapped with the Point datatype in Lucene. This datatype does not create Terms and cannot be used in a MultiTermQuery. The fact that it works on Solr is linked to the fact that the version of Solr that you use still maps numbers to terms. This has changed in es starting in v5 and should also be true in Solr for the latest version (v7).  
If you want to do hamming distance on strings you should define your field as a keyword, in this case the MultiTermQuery will be able to access the TermsEnum for the field. For numbers, hamming distance is not applicable and you should use a range query instead.

---

<div class="post-metadata">

**Author:** ![dominik.safaric](https://avatars.discourse-cdn.com/v4/letter/d/b38774/32.png) [@dominik.safaric](https://discuss.elastic.co/u/dominik.safaric)\
**Post date:** [September 29, 2017, 10:00am UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/3 "2017-09-29T10:00:59Z")

</div>

Thanks for the reply. Hamming distance refers in this case to bit distance and it is calculated using a bitwise XOR. What query class should I implement in order to achieve this?

---

<div class="post-metadata">

**Author:** ![jimczi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jimczi/32/47985_2.png) [@jimczi](https://discuss.elastic.co/u/jimczi)\
**Post date:** [September 29, 2017, 10:04am UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/4 "2017-09-29T10:04:33Z")

</div>

You can look at `PointRangeQuery` which is the abstract impl for range queries on points. Though you'll have to enumerate all points (numbers in your case) and test them all to get the matching candidates, this can be slow if you have a lot of long values to index.

---

<div class="post-metadata">

**Author:** ![dominik.safaric](https://avatars.discourse-cdn.com/v4/letter/d/b38774/32.png) [@dominik.safaric](https://discuss.elastic.co/u/dominik.safaric)\
**Post date:** [September 29, 2017, 1:52pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/5 "2017-09-29T13:52:27Z")

</div>

Thanks for the clarification. How would you then go along with calculating custom scores by implementing the `Scorer` interface? Because the custom query is expected to return a custom score based on a point multi value long field.

---

<div class="post-metadata">

**Author:** ![jimczi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jimczi/32/47985_2.png) [@jimczi](https://discuss.elastic.co/u/jimczi)\
**Post date:** [September 29, 2017, 2:28pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/6 "2017-09-29T14:28:25Z")

</div>

Sorry I don't understand the question. Why are you using a CustomScoreQuery ? It seems unrelated to your problem.  
You only need to extend `Query` or `PointRangeQuery` and implement the logic that you want on this point field.  
The `PointRangeQuery` is a good starting point because it shows how you can iterate the values on a point field.

---

<div class="post-metadata">

**Author:** ![dominik.safaric](https://avatars.discourse-cdn.com/v4/letter/d/b38774/32.png) [@dominik.safaric](https://discuss.elastic.co/u/dominik.safaric)\
**Post date:** [September 29, 2017, 2:33pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/7 "2017-09-29T14:33:53Z")

</div>

I apologize for the misunderstanding. The CustomScoreQuery was used before, whereas now that I've subclassed the PointRangeQuery I'm not using the CustomScoreQuery any longer.

The PointRangeQuery uses the ConstantScoreScoret, whereas I need a custom Scorer in order to compute the score for each matching document using another long point field. That is, for querying the X long point field is used, whereas scores need to be calculated using another point long field of each matching document.

---

<div class="post-metadata">

**Author:** ![jimczi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jimczi/32/47985_2.png) [@jimczi](https://discuss.elastic.co/u/jimczi)\
**Post date:** [September 29, 2017, 6:47pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/8 "2017-09-29T18:47:11Z")

</div>

Ok I understand. Then you need to write you own scorer that takes the distance in account. Though this query as I said before will likely be very slow since it needs to iterate **all** values even when they are indexed as points. You should maybe try another approach that does not require such cost.

---

<div class="post-metadata">

**Author:** ![dominik.safaric](https://avatars.discourse-cdn.com/v4/letter/d/b38774/32.png) [@dominik.safaric](https://discuss.elastic.co/u/dominik.safaric)\
**Post date:** [September 29, 2017, 8:00pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/9 "2017-09-29T20:00:59Z")

</div>

Thanks for the clarification. Is the performance impact due to using multi point long values or using any data type would have the same impact? For example, for the current `Scorer` implementation retrieves an instance of a `Document` and its associated `IndexedField` that is used for scoring. This value is a multi point value, i.e. an array of longs. Would using another data type, such as a `keyword` for example boost performance because there ElasticSearch treats arrays as individual values or not?

---

<div class="post-metadata">

**Author:** ![jimczi](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jimczi/32/47985_2.png) [@jimczi](https://discuss.elastic.co/u/jimczi)\
**Post date:** [October 5, 2017, 3:13pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/10 "2017-10-05T15:13:35Z")

</div>

> Is the performance impact due to using multi point long values or using any data type would have the same impact?

Any data type would be slow if you need to iterate all possible values to find matching candidates.

Regarding the Scorer implementation you should not retrieve a `Document`, that is far too costly to do it for all documents so you should rely on the indexed field or the doc\_values. You can check  
`SortedNumericDocValues.newSlowRangeQuery` for an example of query that retrieves numeric values from doc\_values to match specific documents.  
This discussion is more about Lucene than Elasticsearch so you'd have better advice if you ask the Lucene mailing list instead:  
[java-user@lucene.apache.org](mailto:java-user@lucene.apache.org)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 2, 2017, 3:13pm UTC](https://discuss.elastic.co/t/querys-gettermsenum-not-executed/101501/11 "2017-11-02T15:13:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
