# \[ANNOUNCEMENT\] - fscrawler 2.5 released

**URL:** <https://discuss.elastic.co/t/announcement-fscrawler-2-5-released/143024>\
**Category:** Community Ecosystem\
**Created:** [August 4, 2018, 3:41pm UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-5-released/143024 "2018-08-04T15:41:05Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 4, 2018, 3:41pm UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-5-released/143024/1 "2018-08-04T15:41:06Z")

</div>

The FSCrawler team is pleased to announce the **FSCrawler 2.5** release!

# FSCrawler

FS Crawler offers a simple way to index binary files into elasticsearch.

## Usage

Download [FSCrawler 2.5](https://repo1.maven.org/maven2/fr/pilato/elasticsearch/crawler/fscrawler/2.5/fscrawler-2.5.zip):

```auto
wget https://repo1.maven.org/maven2/fr/pilato/elasticsearch/crawler/fscrawler/2.5/fscrawler-2.5.zip

```

Start FS crawler with:

```auto
bin/fscrawler job_name

```

FS crawler will read a local file (default to `~/.fscrawler/{job_name}/_settings.json`).  
If the file does not exist, FS crawler will propose to create your first job.

```auto
$ bin/fscrawler job_name
18:28:58,174 WARN [f.p.e.c.f.FsCrawler] job [job_name] does not exist
18:28:58,177 INFO [f.p.e.c.f.FsCrawler] Do you want to create it (Y/N)?
y
18:29:05,711 INFO [f.p.e.c.f.FsCrawler] Settings have been created in [~/.fscrawler/job_name/_settings.json]. Please review and edit before relaunch

```

Create a directory named `/tmp/es` or `c:\tmp\es`, add some files you want to index in it and start again:

```auto
$ bin/fscrawler job_name
18:30:34,330 INFO [f.p.e.c.f.FsCrawlerImpl] Starting FS crawler
18:30:34,332 INFO [f.p.e.c.f.FsCrawlerImpl] FS crawler started in watch mode. It will run unless you stop it with CTRL+C.
18:30:34,682 INFO [f.p.e.c.f.FsCrawlerImpl] FS crawler started for [job_name] for [/tmp/es] every [15m]

```

More details in the [documentation](https://fscrawler.readthedocs.io/).

## Some of the new features

- [#585](https://github.com/dadoonet/fscrawler//issues/585): Add a filter by content option . Thanks to dadoonet.
- [#584](https://github.com/dadoonet/fscrawler//issues/584): Ignore files bigger than X . Thanks to dadoonet.
- [#583](https://github.com/dadoonet/fscrawler//issues/583): Add `hocr` option for Tesseract-based OCR . Thanks to dadoonet.
- [#582](https://github.com/dadoonet/fscrawler//issues/582): Allow path partial matching . Thanks to dadoonet.
- [#580](https://github.com/dadoonet/fscrawler//issues/580): Add support for Last Accessed date and Created date . Thanks to dadoonet.
- [#577](https://github.com/dadoonet/fscrawler//issues/577): Add support for cloud id . Thanks to dadoonet.
- [#567](https://github.com/dadoonet/fscrawler//issues/567): Add File Permissions to generated documents . Thanks to dadoonet.
- [#564](https://github.com/dadoonet/fscrawler//issues/564): Add custom tags to documents . Thanks to gpcmol.
- [#563](https://github.com/dadoonet/fscrawler//issues/563): Add support for bulk size in bytes with unit . Thanks to dadoonet.
- [#520](https://github.com/dadoonet/fscrawler//issues/520): Allow setting Tesseract path to executable and data . Thanks to dadoonet.

## Some of the fixed Bugs

- [#579](https://github.com/dadoonet/fscrawler//issues/579): Fix wrong detection of removed settings . Thanks to dadoonet.
- [#553](https://github.com/dadoonet/fscrawler//issues/553): excludes doesn't appear to work with subdirectories/paths . Thanks to a344254.
- [#547](https://github.com/dadoonet/fscrawler//issues/547): fscrawler throws error when using flag --loop 1 . Thanks to jeanp413.
- [#544](https://github.com/dadoonet/fscrawler//issues/544): Allow using `store_source` without indexing content . Thanks to dadoonet.
- [#526](https://github.com/dadoonet/fscrawler//issues/526): Raw fields should be considered as text/keyword . Thanks to dadoonet.
- [#490](https://github.com/dadoonet/fscrawler//issues/490): Missing ES pipeline shows up in fscrawler logs but REST API returns JSON with `"ok": True` . Thanks to shadiakiki1986.
- [#486](https://github.com/dadoonet/fscrawler//issues/486): Includes and Excludes should not be case sensitive . Thanks to dadoonet.
- [#475](https://github.com/dadoonet/fscrawler//issues/475): add setPipeline call when using REST . Thanks to shadiakiki1986.
- [#461](https://github.com/dadoonet/fscrawler//issues/461): ES Pipeline is not working in Rest API . Thanks to suresh-nataraj.
- [#448](https://github.com/dadoonet/fscrawler//issues/448): Fscrawler missing the field file.extension when indexing through Rest API . Thanks to suresh-nataraj.
- [#444](https://github.com/dadoonet/fscrawler//issues/444): Tesseract not detected on Windows . Thanks to HBKarlHolzinger.
- [#439](https://github.com/dadoonet/fscrawler//issues/439): ES Documents missing - Date Mapping issue in RAW field . Thanks to suresh-nataraj.
- [#409](https://github.com/dadoonet/fscrawler//issues/409): Indexed document is not deleted . Thanks to faizalpribadi.
- [#327](https://github.com/dadoonet/fscrawler//issues/327): Indexing Json document via bulk indexing folders also . Thanks to Spandana-Sai.

## Some of the changes

- [#588](https://github.com/dadoonet/fscrawler//issues/588): Update Maven plugins and Libs . Thanks to dadoonet.
- [#569](https://github.com/dadoonet/fscrawler//issues/569): Update to elasticsearch 6.3.2 . Thanks to dadoonet.
- [#554](https://github.com/dadoonet/fscrawler//issues/554): Use \_doc doc type instead of doc . Thanks to dadoonet.
- [#542](https://github.com/dadoonet/fscrawler//issues/542): Update to Tika 1.18 . Thanks to dadoonet.
- [#457](https://github.com/dadoonet/fscrawler//issues/457): Add more info in case of bulk failures . Thanks to dadoonet.

Have fun!  
-FSCrawler team

---

<div class="post-metadata">

**Author:** ![Technical\_Stuffer\_S](https://avatars.discourse-cdn.com/v4/letter/t/41988e/32.png) [@Technical\_Stuffer\_S](https://discuss.elastic.co/u/Technical_Stuffer_S)\
**Post date:** [October 27, 2018, 12:57pm UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-5-released/143024/2 "2018-10-27T12:57:58Z")

</div>

@dadoonet sir i want to import my pdfs into elastic search instance....can you pls provide me the steps for this

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 27, 2018, 2:34pm UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-5-released/143024/3 "2018-10-27T14:34:42Z")

</div>

@Technical_Stuffer_S The documentation explains all that. If you don't understand the documentation please open a new question in #elasticsearch forum with all what you did. I'll be happy to help.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [June 23, 2020, 6:03am UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-5-released/143024/5 "2020-06-23T06:03:54Z")

</div>


