# Ingest-attachment: case insensitive search

**URL:** <https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737>\
**Category:** Elasticsearch\
**Created:** [November 15, 2017, 11:56am UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737 "2017-11-15T11:56:35Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vikentyi](https://avatars.discourse-cdn.com/v4/letter/v/41988e/32.png) [@Vikentyi](https://discuss.elastic.co/u/Vikentyi)\
**Post date:** [November 15, 2017, 11:56am UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737/1 "2017-11-15T11:56:35Z")

</div>

Hello. I have a requirement to perform search in file content. Everything is ok. But is it possible to implement case insensitive search in file content?  
I've read a lot and did't find an answer.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 15, 2017, 12:13pm UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737/2 "2017-11-15T12:13:08Z")

</div>

It’s a question of analyzer that you apply to the field you extract data to with ingest-attachment.  
Not related to ingest-attachment itself.

By default text fields don’t care about case.

What is your problem ?

---

<div class="post-metadata">

**Author:** ![Vikentyi](https://avatars.discourse-cdn.com/v4/letter/v/41988e/32.png) [@Vikentyi](https://discuss.elastic.co/u/Vikentyi)\
**Post date:** [December 4, 2017, 3:30pm UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737/3 "2017-12-04T15:30:08Z")

</div>

ok, I have the following index configuration settings:

```
 {
	"index": {
		"analysis": {
			"filter": {
				"swedish_stop": {
					"type": "stop",
					"stopwords": "_none_"
				},
				"swedish_stemmer": {
					"type": "stemmer",
					"language": "swedish"
				}
			},
			"analyzer": {
				"any": {
					"type": "custom",
					"tokenizer": "standard"
				},
				"any_lowercase": {
					"type": "custom",
					"tokenizer": "standard",
					"filter": [
						"lowercase"
					]
				},
				"swedish": {
					"type": "custom",
					"tokenizer": "standard",
					"filter": [
						"swedish_stop",
						"swedish_stemmer"
					]
				},
				"swedish_lowercase": {
					"type": "custom",
					"tokenizer": "standard",
					"filter": [
						"lowercase",
						"swedish_stop",
						"swedish_stemmer"
					]
				}
			},
			"normalizer": {
				"lowercase_normalizer": {
					"type": "custom",
					"char_filter": [],
					"filter": [
						"lowercase"
					]
				}
			}
		}
	}
}

```

Ingest attachment:

```
 {
    "attachment": {
        "description": "Extract attachment information",
        "processors": [
            {
                "attachment": {
                    "field": "payload",
                    "indexed_chars": "-1",
                    "properties": [
                        "content",
                        "content_type",
                        "content_length",
                        "title",
                        "language"
                    ]
                }
            }
        ]
    }
}

```

For my index I am using the following mapping:

> ```
> {
> "test": {
> "mappings": {
> "_type": {
> "properties": {
> "attachment": {
> "properties": {
> "content": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore_above": 256
> }
> }
> },
> "content_length": {
> "type": "long"
> },
> "content_type": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore_above": 256
> }
> }
> },
> "language": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore_above": 256
> }
> }
> },
> "title": {
> "type": "text",
> "fields": {
> "keyword": {
> "type": "keyword",
> "ignore_above": 256
> }
> }
> }
> }
> },
> "payload": {
> "type": "text",
> "fields": {
> "any_lowercase": {
> "type": "text",
> "analyzer": "any_lowercase"
> }
> },
> "analyzer": "any"
> }
> }
> }
> }
> }
> }
> 
> ```

'payload' is a base64 encoded binary. Is it possible to store only this encoded binary without content? Or vice versa?

And what is the maximum file size for indexing?

---

<div class="post-metadata">

**Author:** ![Vikentyi](https://avatars.discourse-cdn.com/v4/letter/v/41988e/32.png) [@Vikentyi](https://discuss.elastic.co/u/Vikentyi)\
**Post date:** [December 6, 2017, 11:44am UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737/4 "2017-12-06T11:44:45Z")

</div>

Also I have one question more: if I don't want to have encoded content file and remove my field 'payload' from processor - so how can I use analyzer for lowercase? What should I send as request?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 6, 2017, 4:00pm UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737/5 "2017-12-06T16:00:43Z")

</div>

> Is it possible to store only this encoded binary without content?

Yes. But you won't be able to search for the content then.  
And it's not recommended to store blobs in elasticsearch.

> Or vice versa?

Yes. You can add a `remove` processor in your ingest pipeline.

> And what is the maximum file size for indexing?

IIRC by default attachment processor only extracts the first 10000 characters.  
That being said it might be a bad idea to send as a JSON document a very huge file, like a mp4 video file of 10gb if you only want to index some metadata like the filename or what have you. It will overload your node memory.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 6, 2017, 4:01pm UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737/6 "2017-12-06T16:01:56Z")

</div>

> Also I have one question more: if I don't want to have encoded content file and remove my field 'payload' from processor - so how can I use analyzer for lowercase? What should I send as request?

You apply the analyzer not on the `payload` field but on `attachment.content` field where the text is actually extracted to.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 3, 2018, 4:02pm UTC](https://discuss.elastic.co/t/ingest-attachment-case-insensitive-search/107737/7 "2018-01-03T16:02:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
