# How to create unique id in logstash/elastic search for apache logs in a distributed server environment

**URL:** <https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241>\
**Category:** Logstash\
**Created:** [December 18, 2018, 5:41am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241 "2018-12-18T05:41:14Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![elk\_dni\_1](https://avatars.discourse-cdn.com/v4/letter/e/82dd89/32.png) [@elk\_dni\_1](https://discuss.elastic.co/u/elk_dni_1)\
**Post date:** [December 18, 2018, 5:41am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/1 "2018-12-18T05:41:15Z")

</div>

How to create unique id in logstash/elastic search for apache logs in a distributed server environment  
so that when you reupload apache logs, logstash/es will update them instead of creating duplicate records.

---

<div class="post-metadata">

**Author:** ![Wolfram\_Haussig](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wolfram_haussig/32/70528_2.png) [@Wolfram\_Haussig](https://discuss.elastic.co/u/Wolfram_Haussig)\
**Post date:** [December 18, 2018, 6:44am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/2 "2018-12-18T06:44:07Z")

</div>

Hello,

My guess would be to use the [fingerprint plugin](https://www.elastic.co/guide/en/logstash/current/plugins-filters-fingerprint.html) of logstash to generate a hash based on the input line. This value can then be used as document\_id. At last you would have to set the output of logstash to [update](https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-action). Be aware though that I didn't try that out.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 18, 2018, 7:27am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/3 "2018-12-18T07:27:57Z")

</div>

Have a look at the following blogs on the topic:

> **[Efficient Duplicate Prevention for Event-Based Data in Elasticsearch
	  	 |...](https://www.elastic.co/blog/efficient-duplicate-prevention-for-event-based-data-in-elasticsearch)**
>
> Need to prevent duplicates in the Elastic Stack with minimal performance impact? In this blog we look at different options in Elasticsearch and provide some practical guidelines.

> **[Little Logstash Lessons: Handling Duplicates](https://www.elastic.co/blog/logstash-lessons-handling-duplicates)**
>
> Approaches for de-duplicating data in Elasticsearch using Logstash. We also go into examples of how you can use IDs in Elasticsearch Output.

---

<div class="post-metadata">

**Author:** ![elk\_dni\_1](https://avatars.discourse-cdn.com/v4/letter/e/82dd89/32.png) [@elk\_dni\_1](https://discuss.elastic.co/u/elk_dni_1)\
**Post date:** [December 18, 2018, 8:11am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/4 "2018-12-18T08:11:36Z")

</div>

The problem is say if there are hard hitting users. Same request log line can be present more than once in the log(At same time, same user made multiple same requests from same machine).So if we hash the data it will come same, and valid entries also be treated as duplicates.

---

<div class="post-metadata">

**Author:** ![elk\_dni\_1](https://avatars.discourse-cdn.com/v4/letter/e/82dd89/32.png) [@elk\_dni\_1](https://discuss.elastic.co/u/elk_dni_1)\
**Post date:** [December 18, 2018, 8:13am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/5 "2018-12-18T08:13:56Z")

</div>

If we use fingerprint :

The problem is say if there are hard hitting users. Same request log line can be present more than once in the log(At same time, same user made multiple same requests from same machine).So if we hash the data it will come same, and valid entries also be treated as duplicates.

If we use uuid -

Reuploading the log file is creating duplicate entries.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 18, 2018, 8:16am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/6 "2018-12-18T08:16:52Z")

</div>

If you use UUID, you need to do so at the source and then not reprocess the same data. If you calculate a hash and can have identical messages, make sure you include information that sets them apart when you perform your hash calculation, e.g. file name and offset. Another option is to add a UUID to the data before it is first written to file.

---

<div class="post-metadata">

**Author:** ![elk\_dni\_1](https://avatars.discourse-cdn.com/v4/letter/e/82dd89/32.png) [@elk\_dni\_1](https://discuss.elastic.co/u/elk_dni_1)\
**Post date:** [December 18, 2018, 9:45am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/8 "2018-12-18T09:45:09Z")

</div>

Hi Christian,

How to do that in logstash... Can u share one sample/example to understand more about the use cases u shared.

Thanks

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 18, 2018, 10:07am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/9 "2018-12-18T10:07:34Z")

</div>

Did you read the blog posts I linked to? They should contain the information you need.

---

<div class="post-metadata">

**Author:** ![elk\_dni\_1](https://avatars.discourse-cdn.com/v4/letter/e/82dd89/32.png) [@elk\_dni\_1](https://discuss.elastic.co/u/elk_dni_1)\
**Post date:** [December 20, 2018, 11:57am UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/11 "2018-12-20T11:57:34Z")

</div>

Hi Christian Thanks a lot for the blogs.

Can you please check : [Logstash to parse only part of apache log file](https://discuss.elastic.co/t/logstash-to-parse-only-part-of-apache-log-file/161682)

Just want to know using logstash if we can parse only part of apache log file. More details are in the post.

Thanks

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [December 20, 2018, 5:56pm UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/12 "2018-12-20T17:56:29Z")

</div>

without reading too much in to your post here is what I did for unique id

I pay around with fingerprint and did manage to create unique id. but didn't like the idea about it.  
here created my own unique id

this was data coming from jdbc connection with number of row. and when I combine these three field it is unique.

For example  
Where projectname, systemtypeid and username combination will be unique forever.

filter {  
mutate {  
add\_field =\> {  
"doc\_id" =\> "%{projectname}%{systemtypeid}%{username}"  
}  
}

This give me exact duplicate of what I receive from database via jdbc.

Now again I was reading same database table with jdbc but wanted to save this data. i.e read a table once a day and save it. again read a data and put it in ES with second day. each document present size of project by username.

to do that I created another document id which can be unique per day

doc\_id =\> "%{projectname}%{systemtypeid}%{username}%{+dd-MM-YYYY}"

and use this in to output section  
document\_id =\> "%{doc\_id}"

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 17, 2019, 5:56pm UTC](https://discuss.elastic.co/t/how-to-create-unique-id-in-logstash-elastic-search-for-apache-logs-in-a-distributed-server-environment/161241/13 "2019-01-17T17:56:44Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
