# Elasticsearch simple scripted similarity performance issues

**URL:** <https://discuss.elastic.co/t/elasticsearch-simple-scripted-similarity-performance-issues/240134>\
**Category:** Elasticsearch\
**Created:** [July 7, 2020, 9:30am UTC](https://discuss.elastic.co/t/elasticsearch-simple-scripted-similarity-performance-issues/240134 "2020-07-07T09:30:39Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![kamils468](https://avatars.discourse-cdn.com/v4/letter/k/f4b2a3/32.png) [@kamils468](https://discuss.elastic.co/u/kamils468)\
**Post date:** [July 7, 2020, 9:30am UTC](https://discuss.elastic.co/t/elasticsearch-simple-scripted-similarity-performance-issues/240134/1 "2020-07-07T09:30:39Z")

</div>

I have created simple similarity which all it does is returning `doc.freq` .

```
{
  "similarity": {
    "custom_similarity_score": {
      "type": "scripted",
      "script": {
        "source": "return doc.freq;"
      }
    }
  }
}

```

There are also +- 500k documents in index `foo-bar` with structure (Most of them contains term `test`):

```
{
  "mappings": {
    "properties": {
      "name": {
        "type": "keyword",
        "fields": {
          "my": {
            "type": "text",
            "similarity": "simple_similarity"
          },
          "bm25": {
            "type": "text"
          }
        }
      }
    }
  }
}

```

And the query I am using is, eg.:

```auto
{
    "query": {
        "bool": {
            "should": {
                "match": {
                    "name.my": "Test Foo"
                }
            }
        }
    }
}

```

The problem is performance.

For standard `BM25` similarity algorithm, query time takes about up to 10ms (which I test by replacing a query part `name.my` with `name.bm25`).  
However for my `simple_similarity` algorithm, query time takes about 50ms which is weird because it is much simpler than `BM25` . It does not even have any math operations.

What is more...

`Profile` API shows that `score_count` for my simple similarity script of term `test` equals to 501 232, which is the same as `term.docFreq` (The number of documents that contain the current term in the index.) However, `score_count` for `BM25` equals to 10201.

Similar difference is for `advance` in a `profile` api.

Does anybody have any idea why the difference is so huge?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 4, 2020, 9:30am UTC](https://discuss.elastic.co/t/elasticsearch-simple-scripted-similarity-performance-issues/240134/2 "2020-08-04T09:30:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
