# Ruby filter counting error Workers

**URL:** <https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330>\
**Category:** Logstash\
**Created:** [February 23, 2023, 1:39pm UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330 "2023-02-23T13:39:57Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![puched](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/puched/32/117661_2.png) [@puched](https://discuss.elastic.co/u/puched)\
**Post date:** [February 23, 2023, 1:39pm UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/1 "2023-02-23T13:39:57Z")

</div>

I want to store the line number as the document\_id, but i saw that sometimes its not storing in the right way the numbers in order, i guess its because of the multiple workers, the question is Is there a way to store numbers in the right line without losing a lot of performance?  
input {  
file {  
path =\> "c:/path/path/_/_"  
start\_position =\> "beginning"  
sincedb\_path =\>"NUL"  
file\_completed\_action =\> "log\_and\_delete"  
file\_completed\_log\_path =\> "c:/path/log/log.log"  
file\_sort\_by =\> "path"  
mode =\> "read"  
}  
}  
filter{   
ruby { init =\> '@number = 0'  
code =\> '  
@number += 1  
event.set("numLines", @number)' }  
}

output {  
elasticsearch {  
hosts =\> ["localhost:9200"]  
index =\> "index"  
ssl =\> false  
ilm\_enabled =\> false  
user =\>'user'  
document\_id =\> "%{numLines}"  
password=\>'password'  
}

stdout {  
codec =\> line {format =\> "%{numLines}"}  
}

}

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 23, 2023, 1:57pm UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/2 "2023-02-23T13:57:20Z")

</div>

> [@puched](#):
>
> the question is Is there a way to store numbers in the right line without losing a lot of performance?

No, not possible, to get the number of line you would need the file to be processed sequentially and to do that you need to run the pipeline with just one worker, this may or may note impact the performance of this specific pipeline.

Depending on what you are consuming you may preprocess your files to add the line number on the line, this way you would have the line number on each line and could parse your message to get it.

---

<div class="post-metadata">

**Author:** ![puched](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/puched/32/117661_2.png) [@puched](https://discuss.elastic.co/u/puched)\
**Post date:** [February 24, 2023, 7:19am UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/3 "2023-02-24T07:19:28Z")

</div>

Thats the only 2 options I have right?, I read everyday a file that may have repeated logs from other previous days, so I want to store that number as document\_Id so in the case its repeated it wont be added again.

Could be an option to run 2 logstash instances, 1 for that type of file running with 1 worker, and the other one with x workers?, that would made me lose a lot of performance right?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [February 24, 2023, 1:17pm UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/4 "2023-02-24T13:17:16Z")

</div>

> [@puched](#):
>
> I read everyday a file that may have repeated logs from other previous days, so I want to store that number as document\_Id so in the case its repeated it wont be added again.

If this is the only reason to have the line number you actually do not need it, you can use the [`fingerprint`](https://www.elastic.co/guide/en/logstash/current/plugins-filters-fingerprint.html) filter on some field to create a unique ID and then use this unique ID as the document id of the document.

Check this [blog post](https://www.elastic.co/blog/logstash-lessons-handling-duplicates) with some examples.

> [@puched](#):
>
> Could be an option to run 2 logstash instances, 1 for that type of file running with 1 worker, and the other one with x workers?, that would made me lose a lot of performance right?

Not sure how this would work, would you store in different indices? If not, how would this help the duplication case? Also, you are reading the file with `log_and_delete`.

---

<div class="post-metadata">

**Author:** ![puched](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/puched/32/117661_2.png) [@puched](https://discuss.elastic.co/u/puched)\
**Post date:** [March 2, 2023, 7:52am UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/5 "2023-03-02T07:52:59Z")

</div>

I'm really grateful for your answers, sadly I think that the solutions doesn´t works in this case. By the way maybe I explained myself bad, the thing is that the log files may be repetied but with new inserted lines until it gets the maximum size to create another log, so we want to just store the new data, not the past data that was already stored, the fingerprint would be nice if it was different message but the message is the same than the previous twin.

Our Indexes works taking info from the line we reading, not by the document name we reading.

---

<div class="post-metadata">

**Author:** ![Rios](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/rios/32/95745_2.png) [@Rios](https://discuss.elastic.co/u/Rios)\
**Post date:** [March 2, 2023, 10:35am UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/6 "2023-03-02T10:35:52Z")

</div>

puched, ES is not a relational database as MySQL , the focus is on inverted indices, there is no an internal autoincrement id.  
Can you show how does your data look like?

---

<div class="post-metadata">

**Author:** ![leandrojmp](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/leandrojmp/32/107231_2.png) [@leandrojmp](https://discuss.elastic.co/u/leandrojmp)\
**Post date:** [March 2, 2023, 12:51pm UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/7 "2023-03-02T12:51:42Z")

</div>

> [@puched](#):
>
> the fingerprint would be nice if it was different message but the message is the same than the previous twin.

Not sure if I got what the issue is, the `fingerprint` filter is used when you have the **same** message or id and want to store the most recent, you can create a fingerprint based on the entire `message` field, which in your case would be the line that you are reading.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 30, 2023, 12:52pm UTC](https://discuss.elastic.co/t/ruby-filter-counting-error-workers/326330/8 "2023-03-30T12:52:28Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
