# Can we index incremental data for files using FSCrawler?

**URL:** <https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031>\
**Category:** Elasticsearch\
**Created:** [July 31, 2019, 4:34am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031 "2019-07-31T04:34:47Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [July 31, 2019, 4:34am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/1 "2019-07-31T04:34:47Z")

</div>

Hello,

Currently i am using url parameter from .settings.yaml file. in url i have mentioned the path of the drive from which files are getting indexed.  
if any new file is added into the file system and if i want to index only that file and add into the old file index. is that possible using FSCrawler?

Kindly guide.

Regards,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 31, 2019, 8:42am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/2 "2019-07-31T08:42:10Z")

</div>

That's the default behavior of FSCrawler.

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [July 31, 2019, 8:59am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/3 "2019-07-31T08:59:44Z")

</div>

Hello @dadoonet,

Thanks for your reply!!!!  
That means if any new file is added to drive and i have run the fSCrawler job, then it will index only that file. It will not index all the other files and will create duplicate entries. Am i right? correct me if i am wrong.

Reagrds,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 31, 2019, 9:15am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/4 "2019-07-31T09:15:56Z")

</div>

That's correct.

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [July 31, 2019, 10:10am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/5 "2019-07-31T10:10:25Z")

</div>

Hello @dadoonet,

Yes, thanks for your help!!!  
I have tried this solution and it works.  
another question, can we schedule FSCrawler job which currently i am running it manually?  
and can we provide more than one file URL paths to create index?

Regards,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 31, 2019, 10:22am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/6 "2019-07-31T10:22:36Z")

</div>

Once started it runs every 15 minutes by default. You can change this with [https://fscrawler.readthedocs.io/en/latest/admin/fs/local-fs.html](https://fscrawler.readthedocs.io/en/latest/admin/fs/local-fs.html)

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [July 31, 2019, 11:35am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/7 "2019-07-31T11:35:39Z")

</div>

Hello @dadoonet,

I can see that my job settings file has Update rate as 15m by default.  
If any new changes are there it will run atomatically after 15m? or we have to run the FSCrawler job every time using cmd? anyways running through cmd will index the newly added document.  
clear me on this.

Regards,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [July 31, 2019, 12:07pm UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/8 "2019-07-31T12:07:24Z")

</div>

It should detect any new change every 15 minutes.

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [August 1, 2019, 4:40am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/9 "2019-08-01T04:40:37Z")

</div>

Hello @dadoonet,

Yes it has run successfully. I was using --loop 1 while running FSCrawler job.  
Thanks for your help!!

Regards,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 1, 2019, 4:43pm UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/10 "2019-08-01T16:43:18Z")

</div>

Yeah. `loop 1` means it runs once and exits.

---

<div class="post-metadata">

**Author:** ![pyerunka](https://avatars.discourse-cdn.com/v4/letter/p/d6d6ee/32.png) [@pyerunka](https://discuss.elastic.co/u/pyerunka)\
**Post date:** [August 2, 2019, 4:54am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/11 "2019-08-02T04:54:14Z")

</div>

Hello @dadoonet,

Thanks! one more question, is there any way we can give more than one file system URL using FSCrawler job? so that we can index files from another file systems also.

Regards,  
Priyanka

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 2, 2019, 8:10am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/12 "2019-08-02T08:10:19Z")

</div>

It needs to be another job for now.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 30, 2019, 9:22am UTC](https://discuss.elastic.co/t/can-we-index-incremental-data-for-files-using-fscrawler/193031/14 "2019-08-30T09:22:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
