# Logstash 2.3.1 and Kafka offset issue

**URL:** <https://discuss.elastic.co/t/logstash-2-3-1-and-kafka-offset-issue/48331>\
**Category:** Logstash\
**Created:** [April 25, 2016, 2:19pm UTC](https://discuss.elastic.co/t/logstash-2-3-1-and-kafka-offset-issue/48331 "2016-04-25T14:19:18Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![crazyphil](https://avatars.discourse-cdn.com/v4/letter/c/9f8e36/32.png) [@crazyphil](https://discuss.elastic.co/u/crazyphil)\
**Post date:** [April 25, 2016, 2:19pm UTC](https://discuss.elastic.co/t/logstash-2-3-1-and-kafka-offset-issue/48331/1 "2016-04-25T14:19:18Z")

</div>

I have finally built a proper pipeline for logstash to pull data from Kafka, and insert into elasticsearch.

However I seem to be having an issue, or possibly self induced configuration problem:

I need to pull from 8 topics, and here is the general config:

```
kafka {
        auto_offset_reset => "smallest"
        reset_beginning => "true"
        consumer_id => "logf003"
        consumer_threads => "6"
        group_id => "kafka_DataMetrics"
        topic_id => "DataMetrics_SharedDB"
        type => "kafka-datametrics"
        zk_connect => "zookeeper.service.consul:2181"
}

```

So my issue is this - if I use auto\_offset\_reset smallest, my 3 logstash forwarders if they crash would always go to offset #1 in a topic? Would this not cause a lot of data to be resubmitted? Given I have over 2 billion offsets to put into elasticsearch, wouldn't a crash and restart really set my processing back?

If I use the defaults, my logstash forwarders never pull all the data from the start of my kafka topics.

What should I use in this case so that I can:

1. Load all data up to current from my Kafka topics
2. Ensure that should my logstash instances crash, they will restart in an appropriate place, and not pull/send duplicate entries to elasticsearch?

I will also add that it would appear that all my logstash instances are pulling the exact same data at the same time because their message rates as tracked by metrics are all exactly the same.

Thanks for any advice.

---

<div class="post-metadata">

**Author:** ![Joe\_Lawson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joe_lawson/32/3390_2.png) [@Joe\_Lawson](https://discuss.elastic.co/u/Joe_Lawson)\
**Post date:** [April 27, 2016, 2:16pm UTC](https://discuss.elastic.co/t/logstash-2-3-1-and-kafka-offset-issue/48331/2 "2016-04-27T14:16:25Z")

</div>

Take out the offset reset as it is destructive for consumer groups. The  
toggle says, delete my offsets for my group so if a crash occurred you will  
lose all progress. That should fix your problem.

---

<div class="post-metadata">

**Author:** ![Joe\_Lawson](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joe_lawson/32/3390_2.png) [@Joe\_Lawson](https://discuss.elastic.co/u/Joe_Lawson)\
**Post date:** [April 27, 2016, 2:18pm UTC](https://discuss.elastic.co/t/logstash-2-3-1-and-kafka-offset-issue/48331/3 "2016-04-27T14:18:09Z")

</div>

Reset beginning is only applicable when either a consumer group doesn't  
exist or the consumer group's topic offset is out of bounds with the  
current topic.

---

<div class="post-metadata">

**Author:** ![crazyphil](https://avatars.discourse-cdn.com/v4/letter/c/9f8e36/32.png) [@crazyphil](https://discuss.elastic.co/u/crazyphil)\
**Post date:** [May 2, 2016, 6:14pm UTC](https://discuss.elastic.co/t/logstash-2-3-1-and-kafka-offset-issue/48331/4 "2016-05-02T18:14:55Z")

</div>

Ok, that corrected my understanding of the two settings.

I now run with "smallest" offset as the default, and things seem to be moving along quite well.

Thanks!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:59am UTC](https://discuss.elastic.co/t/logstash-2-3-1-and-kafka-offset-issue/48331/5 "2017-07-06T04:59:41Z")

</div>


