# Storing the html stripped version of a document in elasticsearch

**URL:** <https://discuss.elastic.co/t/storing-the-html-stripped-version-of-a-document-in-elasticsearch/98519>\
**Category:** Elasticsearch\
**Created:** [August 28, 2017, 8:55am UTC](https://discuss.elastic.co/t/storing-the-html-stripped-version-of-a-document-in-elasticsearch/98519 "2017-08-28T08:55:56Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Melvyn\_Peignon](https://avatars.discourse-cdn.com/v4/letter/m/a88e57/32.png) [@Melvyn\_Peignon](https://discuss.elastic.co/u/Melvyn_Peignon)\
**Post date:** [August 28, 2017, 8:55am UTC](https://discuss.elastic.co/t/storing-the-html-stripped-version-of-a-document-in-elasticsearch/98519/1 "2017-08-28T08:55:57Z")

</div>

Hi!

I have a piece of text, let's say:

`"<p><br/>台灣人，奧運直播，使用PPStream，(PPS網路電視)，觀看同步奧運實況</b>!"`

I succeed to index this text without the html tags, using "strip\_html". Now I'm trying to **store** this text without the HTML tags:

```
PUT test
{
  "settings" : {
    "index" : {
        "number_of_shards" : 1, 
        "number_of_replicas" : 0
    },
    "analysis": {
      "analyzer": {
        "ch_analyzer": {
          "tokenizer": "icu_tokenizer",
          "char_filter": ["html_strip"]
        }
      }
    }
  },
  "mappings": {
    "qa": {
      "properties": {
        "comment_desc": {
          "type": "text",
          "analyzer": "ch_analyzer"
        },
        "article_title": {
          "type": "text",
          "analyzer": "ch_analyzer"
        },
        "article_desc": {
          "type": "text",
          "analyzer": "ch_analyzer"
        }
      }
    }, 
    "sport": {
      "properties": {
        "title": {
          "type": "text",
          "analyzer": "ch_analyzer"
        },
        "content": {
          "type": "text",
          "analyzer": "ch_analyzer"
        }
      }
    }
  }
}

```

This is my mappings and settings. What should I change to be able to store this text without its HTML tags? Is that possible? Should I preprocess my document beforehand?

---

<div class="post-metadata">

**Author:** ![Mike.Barretta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike.barretta/32/16688_2.png) [@Mike.Barretta](https://discuss.elastic.co/u/Mike.Barretta)\
**Post date:** [August 28, 2017, 11:36pm UTC](https://discuss.elastic.co/t/storing-the-html-stripped-version-of-a-document-in-elasticsearch/98519/2 "2017-08-28T23:36:16Z")

</div>

@Melvyn_Peignon as far as I know, the `html_strip` char filter is indeed removing those tags prior to indexing, so that you are _not_ storing them.

Do you see those HTML tags coming back from a query?

---

<div class="post-metadata">

**Author:** ![Melvyn\_Peignon](https://avatars.discourse-cdn.com/v4/letter/m/a88e57/32.png) [@Melvyn\_Peignon](https://discuss.elastic.co/u/Melvyn_Peignon)\
**Post date:** [August 29, 2017, 5:58am UTC](https://discuss.elastic.co/t/storing-the-html-stripped-version-of-a-document-in-elasticsearch/98519/3 "2017-08-29T05:58:26Z")

</div>

@Mike.Barretta from what I understood I am _not Indexing_ the HTML tags in my inverted index, but I _am storing_ documents with the HTML tags. I made a little example that is easily reproducible:

```
PUT test
{
  "settings" : {
    "index" : {
        "number_of_shards" : 1, 
        "number_of_replicas" : 0
    },
    "analysis": {
      "analyzer": {
        "ch_analyzer": {
          "tokenizer": "standard",
          "char_filter": ["html_strip"]
        }
      }
    }
  },
  "mappings": {
    "sport": {
      "properties": {
        "title": {
          "type": "text",
          "analyzer": "ch_analyzer",
          "store": true
        }
      }
    }
  }
}

```

Now you can add a document that contain some HTML tags

```
PUT test/sport/0
{
  "title": "<div>A little test just for testing </div>"
}

```

I cannot search this document using HTML tags (which is great):

```
GET test/sport/_search
{
  "query": {
    "match": {
      "title": "div"
    }
  }
}

```

Return me:

```
{
  "took": 27,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "failed": 0
  },
  "hits": {
    "total": 0,
    "max_score": null,
    "hits": []
  }
}

```

But I cannot retrieve my document without the HTML:

```
GET test/sport/_search
{
  "stored_fields": "title" 
}

```

Which return me:

```
{
  "took": 1,
  "timed_out": false,
  "_shards": {
    "total": 1,
    "successful": 1,
    "failed": 0
  },
  "hits": {
    "total": 1,
    "max_score": 1,
    "hits": [
      {
        "_index": "test",
        "_type": "sport",
        "_id": "0",
        "_score": 1,
        "fields": {
          "title": [
            "<div>A little test just for testing </div>"
          ]
        }
      }
    ]
  }
}

```

You can see that the document still contains the HTML tags.

---

<div class="post-metadata">

**Author:** ![Mike.Barretta](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mike.barretta/32/16688_2.png) [@Mike.Barretta](https://discuss.elastic.co/u/Mike.Barretta)\
**Post date:** [August 29, 2017, 1:16pm UTC](https://discuss.elastic.co/t/storing-the-html-stripped-version-of-a-document-in-elasticsearch/98519/4 "2017-08-29T13:16:24Z")

</div>

@Melvyn_Peignon I see, I'm sorry - missed the distinction you were making.

So no, if you elect to store the field (`stored:true` or by default in `_source` if not disabled), Elasticsearch stores the "raw" value, not the value post-analyzer. If you want the raw stored value to not include HTML tags, you'll need to remove them before you put them into Elasticsearch. That said, you could probably hack together a [scripted field](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-request-script-fields.html) (which can be stored) that removes the tags, but I wouldn't recommend it.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 26, 2017, 1:16pm UTC](https://discuss.elastic.co/t/storing-the-html-stripped-version-of-a-document-in-elasticsearch/98519/5 "2017-09-26T13:16:29Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
