# Use case question - Can Elasticsearch be used as a log de-duplication solution?

**URL:** <https://discuss.elastic.co/t/use-case-question-can-elasticsearch-be-used-as-a-log-de-duplication-solution/22467>\
**Category:** Elasticsearch\
**Created:** [March 2, 2015, 2:45pm UTC](https://discuss.elastic.co/t/use-case-question-can-elasticsearch-be-used-as-a-log-de-duplication-solution/22467 "2015-03-02T14:45:18Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mihai\_Lucaciu](https://avatars.discourse-cdn.com/v4/letter/m/ebca7d/32.png) [@Mihai\_Lucaciu](https://discuss.elastic.co/u/Mihai_Lucaciu)\
**Post date:** [March 2, 2015, 2:45pm UTC](https://discuss.elastic.co/t/use-case-question-can-elasticsearch-be-used-as-a-log-de-duplication-solution/22467/1 "2015-03-02T14:45:18Z")

</div>

Hi,

I am new to Elasticsearch which I understand can do much more than this...  
but could it be used just for that ?

I am storing 100GB of log files daily. The data scientists require this log  
data to not contain duplicate log lines. Duplicates may come within the  
same log file, with two sequential log files - it's better to expect any  
possible scenario.

What I would like to achieve is to use Elasticsearch to detect & remove the  
duplicate log lines from all logs in an HDFS directory. Can this be done ?

Thank you,  
Mihai

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/a2dd6d7e-698f-4e03-908f-17358c902f6c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a2dd6d7e-698f-4e03-908f-17358c902f6c%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [March 2, 2015, 2:56pm UTC](https://discuss.elastic.co/t/use-case-question-can-elasticsearch-be-used-as-a-log-de-duplication-solution/22467/2 "2015-03-02T14:56:23Z")

</div>

You _could_ do it by using the hash of the unique bits of the log as the  
id. But most systems would support this.

The trouble is that most operations in Elasticsearch are async and  
non-atomic. Operations on ID are atomic and synchronous.

All and all, its not a horrible choice but its not the first tool I'd reach  
for. If your researchers want the data in Elasticsearch in the end I'd go  
with the hash hack. If not I'd investigate some more log processing tools.

It sounds like a fun problem that'd be fun to put together a solution for  
but I'm reasonably sure someone has already done this though. OTOH some  
fun mostly right system using hashing to push data to the right node and  
time bucketed bloom filters would be fun to build and could probably be  
tuned to be pretty good. But I'm sure smarter people than me have already  
solved this problem though. And open sourced the solution.

Nik

On Mon, Mar 2, 2015 at 9:45 AM, Mihai Lucaciu [mlucaciu@gmail.com](mailto:mlucaciu@gmail.com) wrote:

> Hi,
> 
> I am new to Elasticsearch which I understand can do much more than this...  
> but could it be used just for that ?
> 
> I am storing 100GB of log files daily. The data scientists require this  
> log data to not contain duplicate log lines. Duplicates may come within the  
> same log file, with two sequential log files - it's better to expect any  
> possible scenario.
> 
> What I would like to achieve is to use Elasticsearch to detect & remove  
> the duplicate log lines from all logs in an HDFS directory. Can this be  
> done ?
> 
> Thank you,  
> Mihai
> 
> --  
> You received this message because you are subscribed to the Google Groups  
> "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an  
> email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> To view this discussion on the web visit  
> [https://groups.google.com/d/msgid/elasticsearch/a2dd6d7e-698f-4e03-908f-17358c902f6c%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/a2dd6d7e-698f-4e03-908f-17358c902f6c%40googlegroups.com)  
> [https://groups.google.com/d/msgid/elasticsearch/a2dd6d7e-698f-4e03-908f-17358c902f6c%40googlegroups.com?utm\_medium=email&utm\_source=footer](https://groups.google.com/d/msgid/elasticsearch/a2dd6d7e-698f-4e03-908f-17358c902f6c%40googlegroups.com?utm_medium=email&utm_source=footer)  
> .  
> For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/CAPmjWd2Ck-C%2BW84VfTSiUrrm-h35\_chFU80v%2BuGNzvvgxzpKtg%40mail.gmail.com](https://groups.google.com/d/msgid/elasticsearch/CAPmjWd2Ck-C%2BW84VfTSiUrrm-h35_chFU80v%2BuGNzvvgxzpKtg%40mail.gmail.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![Joshua\_Rich1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/joshua_rich1/32/814_2.png) [@Joshua\_Rich1](https://discuss.elastic.co/u/Joshua_Rich1)\
**Post date:** [March 3, 2015, 10:42am UTC](https://discuss.elastic.co/t/use-case-question-can-elasticsearch-be-used-as-a-log-de-duplication-solution/22467/3 "2015-03-03T10:42:24Z")

</div>

Hi Mihai,

This sounds like something you could do in a pre-processing pipeline before  
indexing with Elasticsearch. Have you heard of Logstash? It is designed  
to slurp up logs, filter them (including detecting and removing duplicates)  
then insert them in Elasticsearch (or elsewhere). It can handle live  
streaming of logs or can be run on existing log files. Definitely check it  
out, but I'd imagine if not Logstash, some other kind of pre-processing is  
going to be your best bet.

Regards,

Joshua

On Tuesday, 3 March 2015 01:45:18 UTC+11, Mihai Lucaciu wrote:

> Hi,
> 
> I am new to Elasticsearch which I understand can do much more than this...  
> but could it be used just for that ?
> 
> I am storing 100GB of log files daily. The data scientists require this  
> log data to not contain duplicate log lines. Duplicates may come within the  
> same log file, with two sequential log files - it's better to expect any  
> possible scenario.
> 
> What I would like to achieve is to use Elasticsearch to detect & remove  
> the duplicate log lines from all logs in an HDFS directory. Can this be  
> done ?
> 
> Thank you,  
> Mihai

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/b47ca8d1-6ed2-4165-a021-820b26aa44de%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/b47ca8d1-6ed2-4165-a021-820b26aa44de%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 12:28am UTC](https://discuss.elastic.co/t/use-case-question-can-elasticsearch-be-used-as-a-log-de-duplication-solution/22467/4 "2017-07-06T00:28:58Z")

</div>


