# \[ANNOUNCEMENT\] - FSCrawler 2.7 released

**URL:** <https://discuss.elastic.co/t/announcement-fscrawler-2-7-released/280525>\
**Category:** Community Ecosystem\
**Created:** [August 5, 2021, 11:16am UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-7-released/280525 "2021-08-05T11:16:23Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 5, 2021, 11:16am UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-7-released/280525/1 "2021-08-05T11:16:24Z")

</div>

The FSCrawler team is pleased to announce the **FSCrawler 2.7** release!

# FSCrawler

FS Crawler offers a simple way to index binary files into elasticsearch.

## Usage

Download [FSCrawler 2.7](https://repo1.maven.org/maven2/fr/pilato/elasticsearch/crawler/fscrawler-es7/2.7/fscrawler-es7-2.7.zip):

```sh
wget https://repo1.maven.org/maven2/fr/pilato/elasticsearch/crawler/fscrawler-es7/2.7/fscrawler-es7-2.7.zip

```

Start FS crawler with:

```sh
bin/fscrawler job_name

```

FS crawler will read a local file (default to `~/.fscrawler/{job_name}/_settings.json`).  
If the file does not exist, FS crawler will propose to create your first job.

```sh
$ bin/fscrawler job_name
18:28:58,174 WARN [f.p.e.c.f.FsCrawler] job [job_name] does not exist
18:28:58,177 INFO [f.p.e.c.f.FsCrawler] Do you want to create it (Y/N)?
y
18:29:05,711 INFO [f.p.e.c.f.FsCrawler] Settings have been created in [~/.fscrawler/job_name/_settings.json]. Please review and edit before relaunch

```

Create a directory named `/tmp/es` or `c:\tmp\es`, add some files you want to index in it and start again:

```sh
$ bin/fscrawler job_name
18:30:34,330 INFO [f.p.e.c.f.FsCrawlerImpl] Starting FS crawler
18:30:34,332 INFO [f.p.e.c.f.FsCrawlerImpl] FS crawler started in watch mode. It will run unless you stop it with CTRL+C.
18:30:34,682 INFO [f.p.e.c.f.FsCrawlerImpl] FS crawler started for [job_name] for [/tmp/es] every [15m]

```

More details in the [documentation](https://fscrawler.readthedocs.io/en/fscrawler-2.7/).

## New features

- [#991](https://github.com/dadoonet/fscrawler//issues/991): Add Workplace Search connector.
- [#1203](https://github.com/dadoonet/fscrawler//issues/1203): Add FTP crawler. By helsonxiao.
- [#1211](https://github.com/dadoonet/fscrawler//issues/1211): Add `file.content_type` field on folders.
- [#1210](https://github.com/dadoonet/fscrawler//issues/1210): Add `file.filename` field on folders.
- [#1179](https://github.com/dadoonet/fscrawler//issues/1179): Automatically create Custom Sources.
- [#1037](https://github.com/dadoonet/fscrawler//issues/1037): Split console logs and actual logs and add a banner :).
- [#1036](https://github.com/dadoonet/fscrawler//issues/1036): Support ssl verification configurable. By TommyLike.
- [#1035](https://github.com/dadoonet/fscrawler//issues/1035): Log index errors in documents.log.
- [#1031](https://github.com/dadoonet/fscrawler//issues/1031): Add an external Log4J2 configuration file.
- [#907](https://github.com/dadoonet/fscrawler//issues/907): Add `path_prefix` option.
- [#820](https://github.com/dadoonet/fscrawler//issues/820): Generate FSCrawler docker images. By toto1310.
- [#776](https://github.com/dadoonet/fscrawler//issues/776): Report HEAP size at startup.
- [#752](https://github.com/dadoonet/fscrawler//issues/752): Add option to ignore symlinks. By budachst.
- [#715](https://github.com/dadoonet/fscrawler//issues/715): Allow custom index name in the REST API. By kikkauz.
- [#698](https://github.com/dadoonet/fscrawler//issues/698): Add Cross-Origin Resource Sharing (CORS) headers to RestServer. By isaac-ipl.
- [#692](https://github.com/dadoonet/fscrawler//issues/692): Allow running OCR but not on PDF files.
- [#673](https://github.com/dadoonet/fscrawler//issues/673): Add support for YAML configuration.
- [#663](https://github.com/dadoonet/fscrawler//issues/663): Add Patterns table to includes and excludes. By wrathagom.

## Fixed Bugs

- [#1224](https://github.com/dadoonet/fscrawler//issues/1224): Fix NPE in Console when running with Docker.
- [#1217](https://github.com/dadoonet/fscrawler//issues/1217): Check if date is null when formatting it to RFC3339.
- [#1204](https://github.com/dadoonet/fscrawler//issues/1204): Split build and deploy phases for Docker images.
- [#1201](https://github.com/dadoonet/fscrawler//issues/1201): 2.7 - Docker image broken. By agrantdeakin.
- [#1194](https://github.com/dadoonet/fscrawler//issues/1194): Elasticsearch node settings should not be null by default.
- [#1193](https://github.com/dadoonet/fscrawler//issues/1193): Corrupt PDF can lead to a StackOverflow.
- [#1137](https://github.com/dadoonet/fscrawler//issues/1137): Ignore errors when parsing a 0 byte file.
- [#1085](https://github.com/dadoonet/fscrawler//issues/1085): fscrawler.bat added a CD to move to the appropriate directory. By CircuitGuy.
- [#1084](https://github.com/dadoonet/fscrawler//issues/1084): InputStream must have \> 0 bytes. By yuanzhian.
- [#1066](https://github.com/dadoonet/fscrawler//issues/1066): Start fscrawler instead of internal services.
- [#1041](https://github.com/dadoonet/fscrawler//issues/1041): Fixed an issue that caused an error when running in a windows environment. By muraken720.
- [#1006](https://github.com/dadoonet/fscrawler//issues/1006): Running fscrawler with no argument now lists existing jobs. By janhoy.
- [#1005](https://github.com/dadoonet/fscrawler//issues/1005): Fix ENTRYPOINT in Dockerfile to allow variable substitution. By Maijin.
- [#994](https://github.com/dadoonet/fscrawler//issues/994): Using cloud id gives "invalid IPv6 Address". By tdaroly.
- [#973](https://github.com/dadoonet/fscrawler//issues/973): Fix SSH crawling from Windows machine.
- [#899](https://github.com/dadoonet/fscrawler//issues/899): FSCrawler can't index .doc or .docx elements. By LaaKii.
- [#895](https://github.com/dadoonet/fscrawler//issues/895): java.lang.NoSuchMethodError: parsing some Word files. By mwaltersbmc.
- [#860](https://github.com/dadoonet/fscrawler//issues/860): Bug Syntax error in fscrawler file, to init fscrawler. By CarlosRCDev.
- [#847](https://github.com/dadoonet/fscrawler//issues/847): sun.jnu.encoding=UTF-8 added in .bat and .sh both. By shahariaazam.
- [#834](https://github.com/dadoonet/fscrawler//issues/834): FS Crawler freezes when crawling a 0 byte TXT file. By dansfelix.
- [#819](https://github.com/dadoonet/fscrawler//issues/819): Fix Percentage computation.
- [#760](https://github.com/dadoonet/fscrawler//issues/760): Allow passing test parameters to Maven CLI.
- [#714](https://github.com/dadoonet/fscrawler//issues/714): fix release-drafter. By jetersen.
- [#701](https://github.com/dadoonet/fscrawler//issues/701): Change log level and display logs only if filters on content.
- [#691](https://github.com/dadoonet/fscrawler//issues/691): OCR without pdf\_ocr. By Newmski.
- [#686](https://github.com/dadoonet/fscrawler//issues/686): Wait for healthy index when creating the index.
- [#681](https://github.com/dadoonet/fscrawler//issues/681): SSH dirs should be seen as dirs and not files.
- [#680](https://github.com/dadoonet/fscrawler//issues/680): trying to index remote files with ssh - files seen as folder. By sblanc0054.
- [#660](https://github.com/dadoonet/fscrawler//issues/660): Fix authentication when sending announcement email.

## Main changes

- [#1213](https://github.com/dadoonet/fscrawler//issues/1213): Switch back to Java 11.
- [#1049](https://github.com/dadoonet/fscrawler//issues/1049): Update Dockerfile to use JDK14. By mario-89.
- [#1212](https://github.com/dadoonet/fscrawler//issues/1212): Let's use JsonPath.
- [#1207](https://github.com/dadoonet/fscrawler//issues/1207): Generate only 2 docker images.
- [#1206](https://github.com/dadoonet/fscrawler//issues/1206): Detect when fscrawler runs in foreground and adapt logs.
- [#1205](https://github.com/dadoonet/fscrawler//issues/1205): Add logs to the console when running a Docker instance.
- [#1172](https://github.com/dadoonet/fscrawler//issues/1172): Move CI from Travis to GitHub actions.
- [#872](https://github.com/dadoonet/fscrawler//issues/872): Add more information to the \_simulate API.
- [#700](https://github.com/dadoonet/fscrawler//issues/700): Add dependency convergence checks.
- [#695](https://github.com/dadoonet/fscrawler//issues/695): Exclude the PDFParser from the DefaultParser.
- [#694](https://github.com/dadoonet/fscrawler//issues/694): Display full names when catching parsing errors.
- [#693](https://github.com/dadoonet/fscrawler//issues/693): Move `fs.pdf_ocr` setting to `fs.ocr.pdf_strategy`.
- [#675](https://github.com/dadoonet/fscrawler//issues/675): Warn in case of Tika error.
- [#1219](https://github.com/dadoonet/fscrawler//issues/1219): Update to Elasticsearch 7.14.0 and 6.8.18.
- [#1180](https://github.com/dadoonet/fscrawler//issues/1180): Bump tika.version from 1.26 to 1.27.

## Removed

- [#978](https://github.com/dadoonet/fscrawler//issues/978): files lost. By bluebell1990.

Have fun!  
-FSCrawler team

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 2, 2021, 11:17am UTC](https://discuss.elastic.co/t/announcement-fscrawler-2-7-released/280525/2 "2021-09-02T11:17:02Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
