# ElasticSearch 1.7.3 vs 2.0 vs 2.1 missing data

**URL:** <https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502>\
**Category:** Elasticsearch\
**Created:** [December 7, 2015, 8:56am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502 "2015-12-07T08:56:25Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Adam\_Wrobel](https://avatars.discourse-cdn.com/v4/letter/a/8baadc/32.png) [@Adam\_Wrobel](https://discuss.elastic.co/u/Adam_Wrobel)\
**Post date:** [December 7, 2015, 8:56am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/1 "2015-12-07T08:56:25Z")

</div>

Hi.  
I have ES cluster running on 1.7.3 storing logs parsed by logstash.  
I want to upgrade to ES 2.x, so I ran migration plugin to check what I needed to change.  
I prepare new logstash template compatible with ES 2.x and run new separated cluster with version 2.0.1 and another one separated with 2.1. I'm using logstash 2.1.0

Logs are send to 3 clusters with this part of code:

> output {  
> elasticsearch {  
> hosts =\> "cluster17.es.service.consul"  
> template =\> "/etc/logstash/template\_api.json"  
> index =\> "logstash-%{[@context][\_index]}-%{+YYYY.MM.dd}"  
> template\_overwrite =\> true  
> flush\_size =\> 2000  
> retry\_max\_interval =\> 15  
> max\_retries =\> 6  
> }  
> elasticsearch {  
> hosts =\> "cluster20.es.service.consul"  
> template =\> "/etc/logstash/template\_api2.json"  
> index =\> "logstash-%{[@context][\_index]}-%{+YYYY.MM.dd}"  
> template\_overwrite =\> true  
> flush\_size =\> 2000  
> retry\_max\_interval =\> 15  
> max\_retries =\> 6  
> }  
> elasticsearch {  
> hosts =\> "cluster21.es.service.consul"  
> template =\> "/etc/logstash/template\_api2.json"  
> index =\> "logstash-%{[@context][\_index]}-%{+YYYY.MM.dd}"  
> template\_overwrite =\> true  
> flush\_size =\> 2000  
> retry\_max\_interval =\> 15  
> max\_retries =\> 6  
> }  
> }

And I had weird problem.  
In 1.7 cluster index from 1 day had:  
logstash-api-2015.12.06 items: 11,473,555 size: 5.3GB  
In 2.0.1 cluster:  
logstash-api-2015.12.06 items: 9,609,880 size: 4.7GB  
In 2.1 cluster:  
logstash-api-2015.12.06 items: 9,608,696 size: 4.6GB

Difference between 1.7 and 2.x is huge. And for each full daily indexes 2.x had 15-18% less data.

I tested ES 2.X on different hardware hosts/vms to exclude hardware problems. Also there was no errors in logs.  
I wrote script to compare indexes from 1.7 and 2.x and check what type of message is missing. But for each missing message I can POST it directly using curl to each cluster and everything saved without problems.  
How to debug this issue ?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 7, 2015, 9:37am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/2 "2015-12-07T09:37:51Z")

</div>

It sounds like that a shard is not back. Assuming you have 5 shards per index, that could represent almost 20%.

Did you look at the pending tasks?

---

<div class="post-metadata">

**Author:** ![Adam\_Wrobel](https://avatars.discourse-cdn.com/v4/letter/a/8baadc/32.png) [@Adam\_Wrobel](https://discuss.elastic.co/u/Adam_Wrobel)\
**Post date:** [December 7, 2015, 10:04am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/3 "2015-12-07T10:04:51Z")

</div>

@dadoonet all shards are alocated:

es 2.0.1:  
{"cluster\_name":"logstash20","status":"green","timed\_out":false,"number\_of\_nodes":1,"number\_of\_data\_nodes":1,"active\_primary\_shards":21,"active\_shards":21,"relocating\_shards":0,"initializing\_shards":0,"unassigned\_shards":0,"delayed\_unassigned\_shards":0,"number\_of\_pending\_tasks":0,"number\_of\_in\_flight\_fetch":0,"task\_max\_waiting\_in\_queue\_millis":0,"active\_shards\_percent\_as\_number":100.0}

es 2.1:  
{"cluster\_name":"logstash21","status":"green","timed\_out":false,"number\_of\_nodes":1,"number\_of\_data\_nodes":1,"active\_primary\_shards":20,"active\_shards":20,"relocating\_shards":0,"initializing\_shards":0,"unassigned\_shards":0,"delayed\_unassigned\_shards":0,"number\_of\_pending\_tasks":0,"number\_of\_in\_flight\_fetch":0,"task\_max\_waiting\_in\_queue\_millis":0,"active\_shards\_percent\_as\_number":100.0}

and pending\_tasks on both clusters shows:  
{"tasks":[]}

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 7, 2015, 10:19am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/4 "2015-12-07T10:19:00Z")

</div>

Ha! I totally missed the logstash part.

So you are sending your data to 3 clusters at the same time. But you have a different amount of logs.  
1.7 and 2.0 have the same number of docs. But 2.1 has less.

When you said "Also there was no errors in logs.", did you mean logstash logs or elasticsearch logs?

---

<div class="post-metadata">

**Author:** ![Adam\_Wrobel](https://avatars.discourse-cdn.com/v4/letter/a/8baadc/32.png) [@Adam\_Wrobel](https://discuss.elastic.co/u/Adam_Wrobel)\
**Post date:** [December 7, 2015, 10:37am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/5 "2015-12-07T10:37:48Z")

</div>

@dadoonet: I'm sending data from logstash instance to 3 different ES cluster. And there is no errors in elasticsearch and logstash logs. Everything looks normal.  
And when I post manually messages missed from 2.x clusters and existing in 1.7 I get confirmation with new inserted document \_id.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 7, 2015, 10:59am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/6 "2015-12-07T10:59:08Z")

</div>

I have no explanation.  
May be you could open this thread in logstash forum so experts there might explain or trace things?

---

<div class="post-metadata">

**Author:** ![Adam\_Wrobel](https://avatars.discourse-cdn.com/v4/letter/a/8baadc/32.png) [@Adam\_Wrobel](https://discuss.elastic.co/u/Adam_Wrobel)\
**Post date:** [December 7, 2015, 11:16am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/7 "2015-12-07T11:16:42Z")

</div>

Sure, I write this also on logstash forum. Thanks for trying 🙂

---

<div class="post-metadata">

**Author:** ![Adam\_Wrobel](https://avatars.discourse-cdn.com/v4/letter/a/8baadc/32.png) [@Adam\_Wrobel](https://discuss.elastic.co/u/Adam_Wrobel)\
**Post date:** [December 8, 2015, 9:04am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/8 "2015-12-08T09:04:46Z")

</div>

@warkolm So rest of logstash config I used:

> input {  
> redis {  
> host =\> "10.8.34.27"  
> port =\> 6379  
> data\_type =\> "list"  
> key =\> "logstash"  
> }  
> redis {  
> host =\> "10.8.34.34"  
> port =\> 6379  
> data\_type =\> "list"  
> key =\> "logstash"  
> }  
> redis {  
> host =\> "10.8.34.36"  
> port =\> 6379  
> data\_type =\> "list"  
> key =\> "logstash"  
> }  
> redis {  
> host =\> "10.8.38.42"  
> port =\> 6379  
> data\_type =\> "list"  
> key =\> "logstash"  
> }  
> }  
> filter {
> 
> # drop messages bigger than 16kb
> 
> range {  
> ranges =\> ["message", 16384, 99999999, "drop"]  
> }  
> if [type] == "syslog" {  
> grok {  
> patterns\_dir =\> "/opt/logstash/patterns"  
> match =\> [  
> # apache/nginx access logs  
> "message", "%{APACHE\_ACCESS\_COMBINED\_VHOST}",  
> "message", "%{NGINX\_ACCESS\_LOGS}",  
> # json-formatted messages  
> "message", "%{JSON\_MESSAGES}",  
> # logs from border load balancers  
> "message", "%{BORDER\_LB\_LOG}",  
> # bind logs  
> "message", "%{NAMED\_LOG}",  
> # logs from fastly  
> "message", "%{EDGE\_CACHE\_LOG}",  
> "message", "%{FASTLY\_DEBUG\_LOG}",  
> "message", "%{FASTLY\_RESTARTS\_LOG}",  
> "message", "%{FASTLY\_HOTLINKS}",  
> "message", "%{S\_MAX\_DEBUG}",  
> "message", "%{CB\_DEBUG\_LOG}",  
> # pt-kill  
> "message", "%{PT\_KILL}",  
> # powerconnect  
> "message", "%{POWERCONNECT\_LOG}",  
> "message", "%{VARNISHNCSA}",  
> # normal syslog messages  
> "message", "%{SYSLOG\_STANDARD}"  
> ]  
> }  
> if "\_grokparsefailure" in [tags] {  
> grok {  
> match =\> {  
> "message" =\> "%{GREEDYDATA:syslog\_message}"  
> }  
> }  
> } else {  
> if [jsonMessage] != [null] {  
> json {  
> source =\> "jsonMessage"  
> }  
> mutate {  
> remove\_field =\> ["jsonMessage"]  
> add\_tag =\> ["json"]  
> }  
> } else if [httpversion] != [null] {  
> mutate {  
> add\_tag =\> ["apache\_access\_log"]  
> }  
> if [request] =~ /^/api/v1/ {  
> mutate {  
> add\_field =\> ["[@context][\_index]","api"]  
> }  
> }  
> } else {  
> mutate {  
> rename =\> ["syslog\_message", "@message"]  
> add\_tag =\> ["message"]  
> }  
> }  
> syslog\_pri {}  
> mutate {  
> rename =\> ["syslog\_hostname", "@source\_host"]  
> rename =\> ["syslog\_severity", "severity"]  
> rename =\> ["syslog\_facility", "facility"]  
> rename =\> ["syslog\_pri", "priority"]  
> rename =\> ["syslog\_program", "program"]  
> remove\_field =\> [  
> "@version",  
> "host",  
> "message",  
> "syslog\_facility\_code",  
> "syslog\_severity\_code",  
> "syslog\_timestamp",  
> "type"  
> ]  
> }  
> }  
> }  
> #access log from fastly syslog  
> if [tags] == "edge-cache-requestmessage" {  
> mutate {  
> add\_field =\> ["[@context][\_index]","api"]  
> }  
> }
> 
> # database queries killed
> 
> if [program] == "pt-kill" {  
> date {  
> match =\> ["timestamp", "ISO8601"]  
> }  
> mutate {  
> remove\_field =\> ["timestamp"]  
> }  
> }  
> mutate{  
> lowercase =\> ["@context", "\_index"]  
> }  
> }

---

<div class="post-metadata">

**Author:** ![Adam\_Wrobel](https://avatars.discourse-cdn.com/v4/letter/a/8baadc/32.png) [@Adam\_Wrobel](https://discuss.elastic.co/u/Adam_Wrobel)\
**Post date:** [December 16, 2015, 7:39am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/9 "2015-12-16T07:39:40Z")

</div>

It was actually logstash fault. After upgrading to 2.1.1 each cluster had this same amount of data in each index.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 16, 2015, 8:31am UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/10 "2015-12-16T08:31:30Z")

</div>

Thanks for the update.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:30pm UTC](https://discuss.elastic.co/t/elasticsearch-1-7-3-vs-2-0-vs-2-1-missing-data/36502/11 "2017-07-05T23:30:43Z")

</div>


