# A question around to get relevant content By using TF-IDF algorithm

**URL:** https://discuss.elastic.co/t/a-question-around-to-get-relevant-content-by-using-tf-idf-algorithm/286440
**Category:** Elasticsearch
**Created:** [October 12, 2021, 7:03am UTC](https://discuss.elastic.co/t/a-question-around-to-get-relevant-content-by-using-tf-idf-algorithm/286440 "2021-10-12T07:03:59Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![Lahu\_Gosavi](https://avatars.discourse-cdn.com/v4/letter/l/48db29/32.png) [@Lahu\_Gosavi](https://discuss.elastic.co/u/Lahu_Gosavi)
#### Post date: [October 12, 2021, 7:03am UTC](https://discuss.elastic.co/t/a-question-around-to-get-relevant-content-by-using-tf-idf-algorithm/286440/1 "2021-10-12T07:03:59Z")

</div>

Hi everyone

I am doing a POC on best match document should rank higher  
basically, we are using the TF-IDF algorithm to rank the documents

we are using a multi\_match query to find a document

here is the query:

```auto
{
  "query": {
    "function_score": {
      "query": {
        "bool": {
          "should": [
            {
              "multi_match": {
                "query": "Leadership Development",
                "fields": [
                  "title^35",
                  "description^15",
                  "tags^55"
                ],
                "type": "phrase"
              }
            },
            {
              "multi_match": {
                "query": "Leadership Development",
                "fields": [
                  "title^25",
                  "description^5",
                  "tags^45"
                ]
              }
            }
          ]
        }
      },
      "functions": [
      ]
    }
  },
  "from": 0,
  "size": 100
}

```

Mappings:

```auto
{
  "tags": {
    "analyzer": "standard",
    "type": "text",
    "fields": {
      "keyword": {
        "normalizer": "lcase_keyword",
        "type": "keyword"
      }
    }
  },
  "title": {
    "analyzer": "standard",
    "store": true,
    "type": "text"
  },
  "description": {
    "analyzer": "standard",
    "store": true,
    "type": "text"
  },
  "normalizer": {
    "lcase_keyword": {
      "filter": [
        "lowercase"
      ],
      "type": "custom",
      "char_filter": []
    }
  }
}

```

we have tagging concept where the user can add n number of tags to a document

we support partial as well as phrase match.

TF is calculate based on the length of the field as our use case is that a document can have n number of tags. because of the high number of tags present for a document, it gives lower TF and becasue of lower TF the overall doc score is also low

so because this relevant doc is shown at the end

Is there any way we can avoid this?

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [November 9, 2021, 7:04am UTC](https://discuss.elastic.co/t/a-question-around-to-get-relevant-content-by-using-tf-idf-algorithm/286440/2 "2021-11-09T07:04:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
