# Searching after Indexing in ElasticSearch Problem

**URL:** https://discuss.elastic.co/t/searching-after-indexing-in-elasticsearch-problem/51294
**Category:** Elasticsearch
**Created:** [May 30, 2016, 9:13am UTC](https://discuss.elastic.co/t/searching-after-indexing-in-elasticsearch-problem/51294 "2016-05-30T09:13:36Z")
**Posts on this page:** 4
**Page:** 1

<div class="post-metadata">

### Author: ![ghasem1992](https://avatars.discourse-cdn.com/v4/letter/g/ba8739/32.png) [@ghasem1992](https://discuss.elastic.co/u/ghasem1992)
#### Post date: [May 30, 2016, 9:13am UTC](https://discuss.elastic.co/t/searching-after-indexing-in-elasticsearch-problem/51294/1 "2016-05-30T09:13:36Z")

</div>

I want to index 1 billion records. each record has 2 attributes (attribute1 and attribute2).  
each record that has same value in attribute1 must be merge. for example, I have two record

attribute1 attribute2  
1 4  
1 6

my elastic document must be

```
{
    “attribute1”: 1
    “attribute2”: 4,6
}

```

due to huge amount of data, I must to read a bulk (about 1000 records) and merge them based on the above rule (in memory) and then search them in ElasticSearch and merge them with search result and then index/reindex them.  
In summary I have to Search and Index per bulk respectively.  
I implemented this rule but in some cases Elastic does not return all results and some documents have been indexed duplicately.  
after each Index I Refresh ElasticSearch so that it be ready for next search. but in some case it doesn’t work.  
my index setting is followed as:

```
{
"test_index": {
    "settings": {
        "index": {
            "refresh_interval": "-1",
            "translog": {
                "flush_threshold_size": "1g"
            },
            "max_result_window": "1000000",
            "creation_date": "1464577964635",
            "store": {
                "throttle": {
                    "type": "merge"
                }
            }
        },
        "number_of_replicas": "0",
        "uuid": "TZOse2tLRqGk-vHRMGc2GQ",
        "version": {
            "created": "2030199"
        },
        "warmer": {
            "enabled": "false"
        },
        "indices": {
            "memory": {
                "index_buffer_size": "40%"
            }
        },
        "number_of_shards": "5",
        "merge": {
            "policy": {
                "max_merge_size": "2g"
            }
        }
    }
}

```

how can I resolve this problem?  
Is there any other setting to handle this situation?

---

<div class="post-metadata">

### Author: ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)
#### Post date: [May 30, 2016, 9:53am UTC](https://discuss.elastic.co/t/searching-after-indexing-in-elasticsearch-problem/51294/2 "2016-05-30T09:53:00Z")

</div>

> [@ghasem1992](#):
>
> merge them based on the above rule (in memory)

How does this part work exactly?

---

<div class="post-metadata">

### Author: ![ghasem1992](https://avatars.discourse-cdn.com/v4/letter/g/ba8739/32.png) [@ghasem1992](https://discuss.elastic.co/u/ghasem1992)
#### Post date: [May 30, 2016, 10:21am UTC](https://discuss.elastic.co/t/searching-after-indexing-in-elasticsearch-problem/51294/3 "2016-05-30T10:21:07Z")

</div>

I used a hashmap (dictionary) and check attribute1 in all records of bulk (1000) and based on value of attribute1 grouped them. After this processing step, I will have a smaller set records (e.g. 600) that Is distinct based on attribute1.  
Then I want to search these 600 records based on attributed1 (Term Query) in ElasticSearch.

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [July 5, 2017, 10:48pm UTC](https://discuss.elastic.co/t/searching-after-indexing-in-elasticsearch-problem/51294/4 "2017-07-05T22:48:06Z")

</div>


