# How to do 'Stemming' in - attachment / encoded content

**URL:** <https://discuss.elastic.co/t/how-to-do-stemming-in-attachment-encoded-content/66952>\
**Category:** Elasticsearch\
**Created:** [November 23, 2016, 7:51am UTC](https://discuss.elastic.co/t/how-to-do-stemming-in-attachment-encoded-content/66952 "2016-11-23T07:51:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![radsan84](https://avatars.discourse-cdn.com/v4/letter/r/edb3f5/32.png) [@radsan84](https://discuss.elastic.co/u/radsan84)\
**Post date:** [November 23, 2016, 7:51am UTC](https://discuss.elastic.co/t/how-to-do-stemming-in-attachment-encoded-content/66952/1 "2016-11-23T07:51:56Z")

</div>

I installed elasticsearch 5.0.1 and ingest attachment plugin.  
I have indexed pdf document using ingest attachment processor.  
Now i want to do 'Stemming' in the content of attachement. I tried as below,

1. Created index and set the analyzer for the same.

> ```
> curl -XGET 'http://localhost:9200/idx_analyser?pretty'
> {
> > ` "idx_analyser" : {`
> > "aliases" : { },
> > "mappings" : {
> > "test" : {
> > "properties" : {
> > "attachment" : {
> > "properties" : {
> > "content" : {
> > "type" : "text",
> > "fields" : {
> > "keyword" : {
> > "type" : "keyword",
> > "ignore_above" : 256
> > }
> > }
> > },
> > "content_length" : {
> > "type" : "long"
> > },
> > "content_type" : {
> > "type" : "text",
> > "fields" : {
> > "keyword" : {
> > "type" : "keyword",
> > "ignore_above" : 256
> > }
> > }
> > },
> > "language" : {
> > "type" : "text",
> > "fields" : {
> > "keyword" : {
> > "type" : "keyword",
> > "ignore_above" : 256
> > }
> > }
> > }
> > }
> > },
> > "data" : {
> > "type" : "text",
> > "fields" : {
> > "keyword" : {
> > "type" : "keyword",
> > "ignore_above" : 256
> > }
> > }
> > },
> > "text" : {
> > "type" : "text",
> > "analyzer" : "custom_lowercase_stemmed"
> > }
> > }
> > }
> > },
> > "settings" : {
> > "index" : {
> > "number_of_shards" : "5",
> > "provided_name" : "idx_analyser",
> > "creation_date" : "1479885039440",
> > "analysis" : {
> > "filter" : {
> > "custom_english_stemmer" : {
> > "name" : "english",
> > "type" : "stemmer"
> > }
> > },
> > "analyzer" : {
> > "custom_lowercase_stemmed" : {
> > "filter" : [
> > "lowercase",
> > "custom_english_stemmer"
> > ],
> > "tokenizer" : "standard"
> > }
> > }
> > },
> > "number_of_replicas" : "1",
> > "uuid" : "FrJEtt-BSgq2ROka2PZ4CA",
> > "version" : {
> > "created" : "5000199"
> > }
> > }
> > }
> > }
> > }
> 
> ```

1. Indexed base64content using "pipeline = attachment" processor

> ```
> curl -XPUT 'http://localhost:9200/idx_analyser/test/1?pipeline=attachment&pretty' -d'
> {
> "text": "VGhpcyBpbmRleCBoYXZpbmcgaW5mb3JtYXRpb24="
> }'
> 
> ```

```
   > `{

```

> "\_index" : "idx\_analyser",  
> "\_type" : "test",  
> "\_id" : "1",  
> "\_version" : 2,  
> "found" : true,  
> "\_source" : {  
> "data" : "VGhpcyBpbmRleCBoYXZpbmcgaW5mb3JtYXRpb24=",  
> "attachment" : {  
> "content\_type" : "text/plain; charset=ISO-8859-1",  
> "language" : "en",  
> "content" : "This index having information",  
> "content\_length" : 30  
> }  
> }  
> }  
> `

Searching the content 'having' returns the expected result

```
 curl -XGET 'http://localhost:9200/idx_analyser/_search?q=attachment.content=having'

```

Where as i want to get the same result if i search for 'have' (shown below) . this is not coming !!

```
 curl -XGET 'http://localhost:9200/idx_analyser/_search?q=attachment.content=have'

```

Am i doing anything wrong here ? Please help to resolve this ...

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 23, 2016, 8:22am UTC](https://discuss.elastic.co/t/how-to-do-stemming-in-attachment-encoded-content/66952/2 "2016-11-23T08:22:25Z")

</div>

Please format your code using `</>` icon. It will make your post more readable.

Here you applied your analyzer to `text` field but ingest is writing the extracted content to `attachment.content`.  
Apply your analyzer to `attachment.content` instead.

BTW I doubt that adding a subfield `keyword` to your `attachment.content` field is a good idea.

---

<div class="post-metadata">

**Author:** ![radsan84](https://avatars.discourse-cdn.com/v4/letter/r/edb3f5/32.png) [@radsan84](https://discuss.elastic.co/u/radsan84)\
**Post date:** [November 23, 2016, 9:32am UTC](https://discuss.elastic.co/t/how-to-do-stemming-in-attachment-encoded-content/66952/3 "2016-11-23T09:32:06Z")

</div>

Many Thanks David...  
That filed attachment (with subfield keyword) was not created by me initially ! it has been created automatically by the ingest attachment processor when indexing the pdf encoded content i guess.  
However now i,  
--\> created a new index  
--\> created a mapping type field "attachment" . "content" with 'stem' analyzer on it  
--\> indexed base64 encoded content  
--\> Then did the search ...  
Its worked perfectly ......  
Thank you again....

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 21, 2016, 9:32am UTC](https://discuss.elastic.co/t/how-to-do-stemming-in-attachment-encoded-content/66952/4 "2016-12-21T09:32:27Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
