# File Input from directory with 100K files

**URL:** <https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195>\
**Category:** Logstash\
**Created:** [November 4, 2020, 2:03am UTC](https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195 "2020-11-04T02:03:11Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![surphix](https://avatars.discourse-cdn.com/v4/letter/s/65b543/32.png) [@surphix](https://discuss.elastic.co/u/surphix)\
**Post date:** [November 4, 2020, 2:03am UTC](https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195/1 "2020-11-04T02:03:11Z")

</div>

Hey first time posting here and looking for some understandings.

I have a logstash config that is trying to read a directory containing over 100,000 files. I've ran trace logs and even with sincedb\_path set to /dev/null none of the files get processed. After every file I see sincedbcollection - associate: unmatched.

Is there a limit to how much logstash can handle without sincedb?

_unfortunately I cannot share any configs or logs as it is work related_

---

<div class="post-metadata">

**Author:** ![Wolfram\_Haussig](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wolfram_haussig/32/70528_2.png) [@Wolfram\_Haussig](https://discuss.elastic.co/u/Wolfram_Haussig)\
**Post date:** [November 4, 2020, 5:56am UTC](https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195/2 "2020-11-04T05:56:45Z")

</div>

Hi,

If you cannot share the configs it is hard to help you: Can you anonymize the fields which contain paths and other data which might point to a company or are otherwise confidential?  
Are the files written completely or are they still written to?

You might want to check the following settings:

- [mode](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-mode)
- [start\_position](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-start_position) should be set to `beginning` if `mode` is either unset or explicitly set to `tail` to read all data instead of only ingesting new data.
- [ignore\_older](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-ignore_older) - maybe this setting forces LogStash to ignore your files?
- [path](https://www.elastic.co/guide/en/logstash/current/plugins-inputs-file.html#plugins-inputs-file-path) - have you checked that the path is correct and LogStash is able to read the files?

Best regards  
Wolfram

---

<div class="post-metadata">

**Author:** ![surphix](https://avatars.discourse-cdn.com/v4/letter/s/65b543/32.png) [@surphix](https://discuss.elastic.co/u/surphix)\
**Post date:** [November 4, 2020, 7:05am UTC](https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195/3 "2020-11-04T07:05:33Z")

</div>

Thanks for the reply. The current mode is set to read, start\_position is beginning, I've tried ignore\_older, but it does not work. I know the path is correct as with a smaller subset of data it works with no issues. I am using an XML filter, not sure if this would be the bottleneck?

---

<div class="post-metadata">

**Author:** ![Wolfram\_Haussig](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/wolfram_haussig/32/70528_2.png) [@Wolfram\_Haussig](https://discuss.elastic.co/u/Wolfram_Haussig)\
**Post date:** [November 4, 2020, 7:11am UTC](https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195/4 "2020-11-04T07:11:16Z")

</div>

Do you have [monitoring](https://www.elastic.co/guide/en/logstash/current/logstash-pipeline-viewer.html) enabled for logStash? In this case you can check the throughput in Kibana under Stack Monitoring-\>Pipelines-\>`your pipeline` which looks like this:

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/6/b/6ba11faf594f6efb68cd0550f499ab661eb51feb.png)

If a filter limits the performance you would see that the input plugin would have more events emitted per second than the XML filter.

---

<div class="post-metadata">

**Author:** ![surphix](https://avatars.discourse-cdn.com/v4/letter/s/65b543/32.png) [@surphix](https://discuss.elastic.co/u/surphix)\
**Post date:** [November 4, 2020, 2:40pm UTC](https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195/5 "2020-11-04T14:40:19Z")

</div>

Unfortunately I don't have stack monitoring on the pipelines.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 2, 2020, 2:40pm UTC](https://discuss.elastic.co/t/file-input-from-directory-with-100k-files/254195/6 "2020-12-02T14:40:20Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
