# FScrawler does not scans the \`/tmp/es\`

**URL:** <https://discuss.elastic.co/t/fscrawler-does-not-scans-the-tmp-es/290417>\
**Category:** Elasticsearch\
**Created:** [November 29, 2021, 10:04am UTC](https://discuss.elastic.co/t/fscrawler-does-not-scans-the-tmp-es/290417 "2021-11-29T10:04:48Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ehsan\_kabiri\_33](https://avatars.discourse-cdn.com/v4/letter/e/b77776/32.png) [@ehsan\_kabiri\_33](https://discuss.elastic.co/u/ehsan_kabiri_33)\
**Post date:** [November 29, 2021, 10:04am UTC](https://discuss.elastic.co/t/fscrawler-does-not-scans-the-tmp-es/290417/1 "2021-11-29T10:04:48Z")

</div>

Using Debian 10, Elasticsearch7,Java jdk 11 and FScrawler, when I run the crawler, it only index the files in `/tmp/es` directory at first lunch after first setup of `_settings.yaml`.  
At first initialize, it seems good cause it index all `.pdf` files in the `url` truely.

But after first lunch (which creates indices in Elasticsearch) adding more files to `url` directory, is not added/seen by the crawler. Even stoppnig and restarting the fscrawler, does not results in adding/indexing new files , unless I run `./fscrawler resumes --restart` that results in indexing recently added files to the `url`

This is `_settings.yaml`

```auto
---
name: "resumes"
fs:
  url: "/tmp/es"
  update_rate: "3m"
  excludes:
  - "*/~*"
  json_support: false
  filename_as_id: false
  add_filesize: true
  remove_deleted: true
  add_as_inner_object: false
  store_source: false
  index_content: true
  attributes_support: false
  raw_metadata: false
  xml_support: false
  index_folders: true
  lang_detect: false
  continue_on_error: false
  ocr:
    language: "eng"
    enabled: true
    pdf_strategy: "ocr_and_text"
  follow_symlinks: false
elasticsearch:
  nodes:
  - url: "http://192.168.225.129:9200"
  bulk_size: 100
  flush_interval: "5s"
  byte_size: "10mb"
  ssl_verification: true

```

fscrawler.log:

```auto
03:47:22,328 e[32mINFO e[m [f.p.e.c.f.c.BootstrapChecks] Memory [Free/Total=Percent]: HEAP [13.6mb/494mb=2.77%], RAM [178mb/1.9gb=9.03%], Swap [524.2mb/974.9mb=53.77%].
... Starting FS crawler
... FS crawler started in watch mode. It will run unless you stop it with CTRL+C.

//Many warnings about security ...

03:47:23,188 e[33mWARN e[m [o.e.c.RestClient] request [GET http://192.168.225.129:9200/] returned 1 warnings: [299 Elasticsearch-7.15.2-... "Elasticsearch built-in security features are not enabled. Without authentication, your cluster could be accessible to anyone. See https://www.elastic.co/guide/en/elasticsearch/reference/7.15/security-minimal-setup.html to enable security."]
...

...

03:47:23,605 e[32mINFO e[m [f.p.e.c.f.FsParserAbstract] FS crawler started for [resumes] for [/home/pdf] every [10s]

...

```

Is there any `config` which I have to make?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 30, 2021, 4:01pm UTC](https://discuss.elastic.co/t/fscrawler-does-not-scans-the-tmp-es/290417/2 "2021-11-30T16:01:32Z")

</div>

The current implementation of FSCrawler is not ideal.  
It uses date comparaison to check if something changed.

Depending on the OS, if you move a file for example, then the file does not appear "as new" so FSCrawler is unable to detect it.

The `--restart` option basically does not care about the dates and reindex everything.

There are multiple things I'd like to support in the future:

- Change the implementation: [Use a WatchService implementation · Issue #399 · dadoonet/fscrawler · GitHub](https://github.com/dadoonet/fscrawler/issues/399)
- Trigger manually a file using the REST interface: [Read from any FS Provider using the REST Service · Issue #1247 · dadoonet/fscrawler · GitHub](https://github.com/dadoonet/fscrawler/issues/1247)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 28, 2021, 4:01pm UTC](https://discuss.elastic.co/t/fscrawler-does-not-scans-the-tmp-es/290417/3 "2021-12-28T16:01:50Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
