# Kibana hits are different

**URL:** <https://discuss.elastic.co/t/kibana-hits-are-different/103033>\
**Category:** Kibana\
**Created:** [October 6, 2017, 2:49pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033 "2017-10-06T14:49:40Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 6, 2017, 2:49pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/1 "2017-10-06T14:49:40Z")

</div>

Hi Team,

Thanks a lot for all the advice and help its been a wonderful forum to ask and learn stuffs.  
The issue is that the hit counts in discovery tab are varying when compared to the actual data because of which i am not able to get the right visualizations

[http://sandbox.com:9200/fsimage-2017.10.03/\_count](http://sandbox.com:9200/fsimage-2017.10.03/_count)  
output: {"count":26304276,"\_shards":{"total":5,"successful":5,"skipped":0,"failed":0}}

[http://sandox.com:9200/fsimage-2017.10.04/\_count](http://sandox.com:9200/fsimage-2017.10.04/_count)  
{"count":36942343,"\_shards":{"total":5,"successful":5,"skipped":0,"failed":0}}

[http://sandbox.com:9200/\_cat/indices?v](http://sandbox.com:9200/_cat/indices?v)

yellow open fsimage-2017.10.04 BzplL5UMQFadIVScJapsUQ 5 1 36942343 0 32.4GB 32.4gb  
yellow open fsimage-2017.10.03 b-4oEnRYQMK44trBWIGhBQ 5 1 26304276 0 22.2gb 22.2gb

fsimage-2017.10.04 is the correct one..the rest ones are way off limit and in one entire week except for the Oct 4th rest other days my kibana hits are less..

Please advise how to rectify this

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 6, 2017, 2:57pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/2 "2017-10-06T14:57:23Z")

</div>

Hi @sudi_2611,

thank you for the nice words. 🙂

Could you maybe go into more detail about what the queries in the discover tab or in the visualizations are, maybe with some comparisons of expected vs actual results?

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 6, 2017, 3:20pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/3 "2017-10-06T15:20:09Z")

</div>

welcome!!!

In the discover tab without querying anything i should see approx 36,942,343 or more hits whereas i see approx 26,304,276 hits...  
Expected hits are 36 million or more

during indexing I see the following error in logstash:  
[2017-10-05T17:22:10,166][WARN][logstash.outputs.elasticsearch] UNEXPECTED POOL ERROR {:e=\>#\<LogStash::Outputs::ElasticSearch::HttpClient::Pool::NoConnectionAvailableError: No Available connections\>}  
[2017-10-05T17:22:10,166][ERROR][logstash.outputs.elasticsearch] Attempted to send a bulk request to elasticsearch, but no there are no living connections in the con:

There are no errors in elastic search and kibana logs..

there were 2 index pattern and one had issues which has been removed from logstash file could it be because of that?

But elasticsearch was working fine....  
I am not sure why kibana hits are way off the correct value

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 9, 2017, 10:56am UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/4 "2017-10-09T10:56:09Z")

</div>

The first thing that comes to mind is that there are a few significant differences between the queries Discover uses and the `_count` queries you showed. The `_count` queries just count all documents within the index regardless of their timestamp while the Discover query filters by the date range selected in the time picker. That means that if the index `fsimage-2017.10.03` includes documents that have no timestamp or whose timestamp is not on that day, they will be counted in the `_count` query, but not in the Discover query that uses the timefilter for the day `2017-10-03`.

To check what the range of timestamp values in an index is, you could use something like the following (with the timestamp field name replaced):

```
GET fsimage-2017.10.03/_search
{ "aggs": { "timestamp_stats": { "stats": { "field": "@timestamp" } } }, "size": 0 }

```

To get the number of documents that have the timestamp field:

```
GET fsimage-2017.10.03/_count
{ "query": { "exists": { "field": "@timestamp" } } }

```

The request that Kibana sends to Elasticsearch can be inspected using the small arrow icon beneath the histogram at the top.

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 10, 2017, 12:26pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/5 "2017-10-10T12:26:29Z")

</div>

Hi,  
Thanks a lot for the suggestions.

when i execute the above mentioned query i get the following output:  
The below is for GET fsimage-2017.10.03/\_search  
{  
"took": 0,  
"timed\_out": false,  
"\_shards": {  
"total": 5,  
"successful": 5,  
"skipped": 0,  
"failed": 0  
},  
"hits": {  
"total": 26304276,  
"max\_score": 0,  
"hits": []  
},  
"aggregations": {  
"timestamp\_stats": {  
"count": 26304276,  
"min": 1507035305954,  
"max": 1507039205420,  
"avg": 1507037136233.399,  
"sum": 39641520773732925000,  
"min\_as\_string": "2017-10-03T12:55:05.954Z",  
"max\_as\_string": "2017-10-03T14:00:05.420Z",  
"avg\_as\_string": "2017-10-03T13:25:36.233Z",  
"sum\_as\_string": "292278994-08-17T07:12:55.807Z"  
}  
}  
}

this is for fsimage-2017.10.04/\_search  
{  
"took": 329,  
"timed\_out": false,  
"\_shards": {  
"total": 5,  
"successful": 5,  
"skipped": 0,  
"failed": 0  
},  
"hits": {  
"total": 36942343,  
"max\_score": 0,  
"hits": []  
},  
"aggregations": {  
"timestamp\_stats": {  
"count": 36942343,  
"min": 1507121781038,  
"max": 1507126610368,  
"avg": 1507123908634.2017,  
"sum": 55676688376265335000,  
"min\_as\_string": "2017-10-04T12:56:21.038Z",  
"max\_as\_string": "2017-10-04T14:16:50.368Z",  
"avg\_as\_string": "2017-10-04T13:31:48.634Z",  
"sum\_as\_string": "292278994-08-17T07:12:55.807Z"  
}  
}  
}

The actual count of documents should be more than 36 million but expect for one instance rest all days i get the count somewhere close to 27 million...i am not sure why the rest of data is not showing.

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 11, 2017, 2:54am UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/6 "2017-10-11T02:54:23Z")

</div>

hi,

I also find these errors when trying to output Logstash to Elasticsearch  
[2017-10-10T16:52:31,992][INFO][logstash.outputs.elasticsearch] retrying failed action with response code: 429 ({"type"=\>"es\_rejected\_execution\_exception", "reason"=\>"rejected execution of org.elasticsearch.transport.TransportService$7@2a875324 on EsThreadPoolExecutor[bulk, queue capacity = 200, org.elasticsearch.common.util.concurrent.EsThreadPoolExecutor@113cc928[Running, pool size = 32, active threads = 32, queued tasks = 200, completed tasks = 16304171]]"})

[2017-10-10T16:56:56,983][WARN][logstash.outputs.elasticsearch] UNEXPECTED POOL ERROR {:e=\>#\<LogStash::Outputs::ElasticSearch::HttpClient::Pool::NoConnectionAvailableError: No Available connections\>}

[2017-10-10T16:56:56,983][ERROR][logstash.outputs.elasticsearch] Attempted to send a bulk request to elasticsearch, but no there are no living connections in the connection pool. Perhaps Elasticsearch is unreachable or down? {:error\_message=\>"No Available connections", :class=\>"LogStash::Outputs::ElasticSearch::HttpClient::Pool::NoConnectionAvailableError", :will\_retry\_in\_seconds=\>4}

could it be because of this that all data is not being output to elasticsearch hence the data variance.  
If so how do i rectify that?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [October 11, 2017, 6:12am UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/7 "2017-10-11T06:12:08Z")

</div>

Which version of Elasticsearch and Logstash are you using?

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 11, 2017, 12:11pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/8 "2017-10-11T12:11:19Z")

</div>

Hi,  
logstash 5.6.1 and ES is also 5.6.1

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 12, 2017, 3:24pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/9 "2017-10-12T15:24:51Z")

</div>

Hi All,  
Please provide advice for below issue as its impacting production

In logstash logs i see the following errors:

[2017-10-11T16:51:50,619][INFO][logstash.outputs.elasticsearch] retrying failed action with response code: 500 ({"type"=\>"class\_cast\_exception", "reason"=\>"org.elasticsearch.index.mapper.TextFieldMapper cannot be cast to org.elasticsearch.index.mapper.DateFieldMapper"})

[2017-10-11T16:51:50,487][INFO][logstash.outputs.elasticsearch] retrying failed action with response code: 429 ({"type"=\>"es\_rejected\_execution\_exception", "reason"=\>"rejected execution of org.elasticsearch.transport.TransportService$7@5dd8577d on EsThreadPoolExecutor[bulk, queue capacity = 200, org.elasticsearch.common.util.concurrent.EsThreadPoolExecutor@113cc928[Running, pool size = 32, active threads = 32, queued tasks = 200, completed tasks = 19473750]]"})

[2017-10-11T16:54:48,871][ERROR][logstash.outputs.elasticsearch] Attempted to send a bulk request to elasticsearch, but no there are no living connections in the connection pool. Perhaps Elasticsearch is unreachable or down? {:error\_message=\>"No Available connections", :class=\>"LogStash::Outputs::ElasticSearch::HttpClient::Pool::NoConnectionAvailableError", :will\_retry\_in\_seconds=\>4}

because of the above errors elasticsearch is not able to receive accurate data.  
ELK is running on a single node cluster.

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 13, 2017, 8:56am UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/10 "2017-10-13T08:56:59Z")

</div>

@sudi_2611,

let me preface this by recommending to contact our support through [https://support.elastic.co](https://support.elastic.co) if you have a support contract and let them know you have a problem in your production environment. That way you will receive priority support that will probably be more timely than this best-effort forum. 😉

With that out of the way, the error messages show two kinds of problems:

- The first message indicates that you are trying to index documents with fields that can not be parsed according to the Elasticsearch mapping, e.g. a string can not be parsed as a date. The Elasticsearch logs should contain corresponding error messages for the rejected actions.
- The second indicates that the connectivity of your Logstash instances to your Elasticsearch cluster is interrupted from time to time. That can have various reasons specific to your deployment environment, e.g. a lossy network connection, DNS problems, etc. The Elasticsearch logs might show corresponding entries as well. Otherwise I would suggest searching the system logs and system monitoring tools you use for any indication of networking problems.

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 13, 2017, 6:40pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/11 "2017-10-13T18:40:50Z")

</div>

Hi,

thanks a lot for the inputs...  
my input is as follows:

filter {  
if [type] == "fsimage" {  
csv {  
separator =\> "|"  
columns =\> ["HDFSPath", "replication", "ModificationTime", "AccessTime", "PreferredBlockSize", "BlocksCount", "FileSize", "NSQUOTA", "DSQUOTA", "permission", "user", "group"]  
convert =\> {  
'replication' =\> 'integer'  
'PreferredBlockSize' =\> 'integer'  
'BlocksCount' =\> 'integer'  
'FileSize' =\> 'integer'  
'NSQUOTA' =\> 'integer'  
'DSQUOTA' =\> 'integer'  
}  
}

date {  
match =\> ['ModificationTime', 'YYYY-MM-ddHH:mm']  
target =\> "modifyTime"  
remove\_field =\> ['ModificationTime']  
}

```
    date {
            match => ['AccessTime', 'YYYY-MM-ddHH:mm']
            }

    date {
            match => ['AccessTime', 'YYYY-MM-ddHH:mm']
            target => "accessTime"
            remove_field => ['AccessTime']
            }

```

There are no errors for rejections in elasticsearch as well..

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 25, 2017, 12:18pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/12 "2017-10-25T12:18:59Z")

</div>

Hi @sudi_2611,

did you make any progress with your problem? The "rejected execution" error message also indicates your cluster is overloaded. Maybe expanding it to a 3-node cluster could help.

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 25, 2017, 1:16pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/13 "2017-10-25T13:16:09Z")

</div>

Hi,

Thanks a lot for getting back..still my elasticsearch is failing as its not able to cope up with incoming records..logstash is outputting close to 47 million records but its failing with above mentioned error in previous posts...

I have ELK setup in single node cluster which is of 256 Gb RAM and 48 cpu cores, i am not sure where the issue is...will increasing the queue capacity, heap size help?? my heap size for elasticsearch is 25 GB ..any inputs would be helpful...

I have tried all possible solutions, split the output file but still same errors.

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 25, 2017, 1:37pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/14 "2017-10-25T13:37:57Z")

</div>

The heap size sounds ok. (Due to limitations of the JVM, allocating 31GB of RAM (-Xmx32600m) to a node's heap is the recommended maximum.)

You might be able to improve your situation by running a three node Elasticsearch cluster on three machines with 64 GB each. That way half of each machine's memory can be used for the JVM heap and the other half for the OS filesystem buffer. Distributing the load on three machines should increase your indexing throughput due to improved parallelism. (In addition to all the other advantages of a multi-node cluster such as rolling updates and high availability.)

Increasing the queue size could help you if you want to compensate for temporary spikes, but a constant overload will still fill it up. Scaling the cluster by adding more nodes would probably be your best bet to increase the indexing rate.

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 25, 2017, 1:45pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/15 "2017-10-25T13:45:14Z")

</div>

thanks for the suggestions will look into...i am extracting the fsimage and delimiting the file into a csv will apache hive with ES hadoop connector work?? i have see the documentation my only question is can apache hive solve the problem??

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 25, 2017, 1:58pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/16 "2017-10-25T13:58:58Z")

</div>

I don't know much about Apache Hive, but I don't immediately see how using the ES Hadoop connector would help you. It could help if you had problems with retaining a long history of data, but if the ingest rate remains high, the single-node cluster will remain overloaded.

---

<div class="post-metadata">

**Author:** ![sudi\_2611](https://avatars.discourse-cdn.com/v4/letter/s/47e85d/32.png) [@sudi\_2611](https://discuss.elastic.co/u/sudi_2611)\
**Post date:** [October 25, 2017, 2:07pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/17 "2017-10-25T14:07:58Z")

</div>

Oh k.. What if i split the 47million record file and store the split file in a staging directory and only output 100,000 records in intervals of time..will that work?? how much bulk request can elasticsearch take at a single point in time?

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 25, 2017, 2:15pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/18 "2017-10-25T14:15:12Z")

</div>

That depends on the performance of your cluster, which depends on the performance of the system it is running on. Batching up the records could reduce the overhead a bit, but logstash already does some batching itself, I think. In the end, indexing documents just takes time. Scaling the cluster horizontally is a good way to keep up with the rate at which documents are produced.

---

<div class="post-metadata">

**Author:** ![weltenwort](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/weltenwort/32/53885_2.png) [@weltenwort](https://discuss.elastic.co/u/weltenwort)\
**Post date:** [October 25, 2017, 2:22pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/19 "2017-10-25T14:22:20Z")

</div>

If your data are read from disk and not produced at a constant high rate, you might have some success using [logstash's sleep filter](https://www.elastic.co/guide/en/logstash/current/plugins-filters-sleep.html) to limit the rate at which logstash sends documents to Elasticsearch. That obviously introduces another significant delay into the processing until the data are available for search. If the rate of events is constant over the whole day scaling the cluster really seems to be the only option.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 22, 2017, 2:22pm UTC](https://discuss.elastic.co/t/kibana-hits-are-different/103033/20 "2017-11-22T14:22:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
