# \_analyze on multiple documents

**URL:** <https://discuss.elastic.co/t/analyze-on-multiple-documents/108487>\
**Category:** Elasticsearch\
**Created:** [November 21, 2017, 2:00am UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487 "2017-11-21T02:00:04Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![seanlyu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seanlyu/32/24492_2.png) [@seanlyu](https://discuss.elastic.co/u/seanlyu)\
**Post date:** [November 21, 2017, 2:00am UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/1 "2017-11-21T02:00:04Z")

</div>

Sorry for the Korean values that could confuse you. But it wouldn't be that hard to understand when you are reading this.

I am using the index named sliced\_data and it his milions of documents in it.  
I am using Kibana and have set Mecab\_Ko(Korean tokenizer, analyzer) as analyzer.  
Analyzer is working fine so when I run the command below

```auto
POST /sliced_data/_analyze
{
  "analyzer": "korean",
  "text": "꽃을든남자"
}

```

**These are the results**

```auto
{
"tokens": [
{
"token": "꽃을",
"start_offset": 0,
"end_offset": 2,
"type": "EOJEOL",
"position": 0
},
{
"token": "꽃",
"start_offset": 0,
"end_offset": 1,
"type": "NNG",
"position": 0
},
{
"token": "든",
"start_offset": 2,
"end_offset": 3,
"type": "INFLECT",
"position": 1
},
{
"token": "들/VV",
"start_offset": 2,
"end_offset": 3,
"type": "VV",
"position": 1
},
{
"token": "남자",
"start_offset": 3,
"end_offset": 5,
"type": "NNG",
"position": 2
}
]
}

```

I want to collect the tokens that have "NNG" as the value of "type".

So this means that I have to \_analyze milions of texts to get the result i want.

It would take a long time to run the query million times.

Is there anyway that elasticsearch provides to \_analyze multiple documents?

I found a way to analyze an array of text like below, but it would be hard to paste all the texts since I have millions of data.

```auto
POST /sliced_data/_analyze
{
"analyzer": "korean",
"text": ["꽃을 든 남자", "초보개발자", "blashhs", "blahblah", "blahblahblahblah"]
}

```

Is there a good solution??

Thank you.

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [November 21, 2017, 1:50pm UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/2 "2017-11-21T13:50:59Z")

</div>

There is no easier way to do it, you will have to run batches of documents through the analyze API by passing an array to `text` like you did in your last example.

---

<div class="post-metadata">

**Author:** ![seanlyu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seanlyu/32/24492_2.png) [@seanlyu](https://discuss.elastic.co/u/seanlyu)\
**Post date:** [November 22, 2017, 12:41am UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/3 "2017-11-22T00:41:32Z")

</div>

Is there any way to pass an array to "text" without typing it??  
For now I made an array that has all the strings inside it. And I am analyzing it one by one.  
I have to run \_analyze million times if i have million documents to analyze it.  
Is there a way to pass an array to "text"?

Thanks.

---

<div class="post-metadata">

**Author:** ![johtani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johtani/32/44956_2.png) [@johtani](https://discuss.elastic.co/u/johtani)\
**Post date:** [November 22, 2017, 6:42am UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/4 "2017-11-22T06:42:21Z")

</div>

Are you only collect "NNG" takens in your collection?  
How about using [`keep_type`](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-keep-types-tokenfilter.html) filter with reindex?  
It is only keep term that has types you specified. So, you can get terms from the field with keep\_type.

---

<div class="post-metadata">

**Author:** ![seanlyu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seanlyu/32/24492_2.png) [@seanlyu](https://discuss.elastic.co/u/seanlyu)\
**Post date:** [November 22, 2017, 7:35am UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/5 "2017-11-22T07:35:16Z")

</div>

I tried using keep\_type.

PUT /extras  
{  
"settings" : {  
"analysis":{  
"analyzer":{  
"korean":{  
"type":"custom",  
"tokenizer":"mecab\_ko\_standard\_tokenizer",  
"filter" : ["erase\_noise"]  
}  
},  
"tokenizer": "mecab\_ko\_standard\_tokenizer",  
"filter" : {  
"erase\_noise" : {  
"type" : "keep\_types",  
"types" : ["NNG"]  
}  
}  
}  
},  
"mappings": {  
"product\_details": {  
"properties": {  
"message": {  
"type": "text",  
"analyzer": "korean",  
"search\_analyzer": "korean"  
}  
}  
}  
}  
}  
This is what I tried.  
I think that if I can get the list of all the tokens in the index there will be no problem.  
But I don't know how.  
Is there a method that I can see all the tokens in the index???

Thanks.

---

<div class="post-metadata">

**Author:** ![seanlyu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/seanlyu/32/24492_2.png) [@seanlyu](https://discuss.elastic.co/u/seanlyu)\
**Post date:** [November 23, 2017, 12:16pm UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/6 "2017-11-23T12:16:33Z")

</div>

Is it simillar to this?

POST \_reindex  
{  
"source": {  
"index": "sliced\_data"  
},  
"query":{  
"filter" : {  
"erase\_noise" : {  
"type" : "keep\_types",  
"types" : ["NNG"]  
}  
}  
},  
"dest": {  
"index": "all\_nng"  
}  
}

please give me an advice.

Thanks.

---

<div class="post-metadata">

**Author:** ![johtani](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/johtani/32/44956_2.png) [@johtani](https://discuss.elastic.co/u/johtani)\
**Post date:** [November 26, 2017, 12:34pm UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/7 "2017-11-26T12:34:14Z")

</div>

You can get terms with [terms aggregation](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html).  
And you can get all terms using [https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#\_filtering\_values\_with\_partitions](https://www.elastic.co/guide/en/elasticsearch/reference/current/search-aggregations-bucket-terms-aggregation.html#_filtering_values_with_partitions)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 24, 2017, 12:34pm UTC](https://discuss.elastic.co/t/analyze-on-multiple-documents/108487/8 "2017-12-24T12:34:15Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
