# Ingesting documents (pdf, word, .txt) to elasticsearch

**URL:** <https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128>\
**Category:** Elasticsearch\
**Created:** [February 15, 2017, 3:17am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128 "2017-02-15T03:17:12Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 15, 2017, 3:17am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/1 "2017-02-15T03:17:12Z")

</div>

Hi I want to ask about how can this be done, so here's the scenario my group and I had a usecase about logs which we technically use logstash. The goal of the "logs usecase" is to ingest and analyse all logs . Every machine dumping logs on a single data center which is hadoop (MapR distribution using MapR-FS) while logstash continuously read this inputs and send them to elasticsearch.

Now on the next use-case I know logstash is not possible as a candidate to be used on Ingesting documents and make the elasticsearch search within the documents and eventually return where the document address is (specifically like a hyperlink where the document resides like "/maprfs/documents/thisdocument.pdf").

Clients continuously dumping new documents (pdf,word,text or whatsoever) and also elasticsearch is continuously ingesting these documents and when a client search a word elasticsearch will return what document has those words while giving a hyperlink where the document resides.

Im quite puzzled on what to use or is this even possible?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 15, 2017, 6:11am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/2 "2017-02-15T06:11:36Z")

</div>

Did you look at FSCrawler project?

Might be what you are looking for.

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 15, 2017, 6:32am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/3 "2017-02-15T06:32:34Z")

</div>

Is this one ? [https://github.com/dadoonet/fscrawler](https://github.com/dadoonet/fscrawler)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 15, 2017, 6:51am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/4 "2017-02-15T06:51:33Z")

</div>

Yes it is

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 15, 2017, 9:13am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/5 "2017-02-15T09:13:13Z")

</div>

Currently looking at it right now, i did make an index which is the job name and I tried it on a single pdf. Tried to search it to kibana and elasticsearch but the index search gives me an error.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 15, 2017, 10:10am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/6 "2017-02-15T10:10:57Z")

</div>

What kind of error?

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 16, 2017, 12:33am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/8 "2017-02-16T00:33:18Z")

</div>

@dadoonet

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 16, 2017, 12:45am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/9 "2017-02-16T00:45:50Z")

</div>

it isn't error I guess but when I try to ingest a pdf and do a search query here is the result.

 ![](https://us1.discourse-cdn.com/elastic/original/2X/7/793da9f135ff752fcf413a3bef762ddf6027fea4.png)

Heres my \_settings.json

 ![](https://us1.discourse-cdn.com/elastic/original/2X/1/1ab50d1a78e3f5e2ba34eefc2236076a752ad04a.png)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 16, 2017, 4:58am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/10 "2017-02-16T04:58:07Z")

</div>

Can you share the exact commands you used, FSCrawler logs and also use --debug option to get even more details?

And please don't share screenshots but formatted text with:

````
```
CODE
```
````

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 17, 2017, 12:39am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/11 "2017-02-17T00:39:35Z")

</div>

@dadoonet  
Hi ! I managed to make it work, but now I have a new problem. Every 15 minutes fscrawler search any new documents right?. I tried to adjust it to 1minute and start the fscrawler **./fscrawler job5** then after starting it i waited about some more minutes and add another documents inside the url folder but unfortunately those new documents are not indexed.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 17, 2017, 6:52am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/12 "2017-02-17T06:52:16Z")

</div>

Please run it with debug option and share your logs and config file.

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 17, 2017, 7:34am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/13 "2017-02-17T07:34:41Z")

</div>

Were can I see those logs ?

or should i just copy paste the logs on the screen ??

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 17, 2017, 8:08am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/14 "2017-02-17T08:08:42Z")

</div>

@dadoonet first log without putting any new documents

```auto
16:07:47,322 DEBUG [f.p.e.c.f.FsCrawlerImpl] Fs crawler is now waking up again...
16:07:47,323 DEBUG [f.p.e.c.f.FsCrawlerImpl] Fs crawler thread [jobs] is now running. Run #2...
16:07:47,336 DEBUG [f.p.e.c.f.FsCrawlerImpl] indexing [/mapr/my.cluster.com/vm1] content
16:07:47,336 DEBUG [f.p.e.c.f.f.FileAbstractor] Listing local files from /mapr/my.cluster.com/vm1
16:07:47,338 DEBUG [f.p.e.c.f.f.FileAbstractor] 1 local files found
16:07:47,338 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null], excludes = [[~*]]
16:07:47,339 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], excludes = [[~*]]
16:07:47,339 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null]
16:07:47,339 DEBUG [f.p.e.c.f.FsCrawlerImpl] [changelog.txt] can be indexed: [true]
16:07:47,339 DEBUG [f.p.e.c.f.FsCrawlerImpl] - file: changelog.txt
16:07:47,339 DEBUG [f.p.e.c.f.FsCrawlerImpl] - not modified: creation date 2016-12-12T13:48:24 , file date 2016-12-12T13:48:24, last scan date 2017-02-17T16:06:44.904
16:07:47,340 DEBUG [f.p.e.c.f.FsCrawlerImpl] Looking for removed files in [/mapr/my.cluster.com/vm1]...
16:07:47,340 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[doc], request [SearchRequest{query=path.encoded:66e1f91ce6a0761b736dbb8117e542e, fields=[_source, file.filename], size=10000}]
16:07:47,348 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null], excludes = [[~*]]
16:07:47,348 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], excludes = [[~*]]
16:07:47,349 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null]
16:07:47,349 DEBUG [f.p.e.c.f.FsCrawlerImpl] Looking for removed directories in [/mapr/my.cluster.com/vm1]...
16:07:47,349 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[folder], request [SearchRequest{query=encoded:66e1f91ce6a0761b736dbb8117e542e, fields=[], size=10000}]
16:07:47,352 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [/mapr/my.cluster.com/vm1], includes = [null], excludes = [[~*]]
16:07:47,353 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [/mapr/my.cluster.com/vm1], excludes = [[~*]]
16:07:47,353 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [/mapr/my.cluster.com/vm1], includes = [null]
16:07:47,353 DEBUG [f.p.e.c.f.FsCrawlerImpl] Delete folder /mapr/my.cluster.com/vm1//mapr/my.cluster.com/vm1
16:07:47,353 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[doc], request [SearchRequest{query=path.encoded:eec8ebf874bb5f54d44cace29601ed0, fields=[_source, file.filename], size=10000}]
16:07:47,357 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[folder], request [SearchRequest{query=encoded:eec8ebf874bb5f54d44cace29601ed0, fields=[], size=10000}]
16:07:47,360 DEBUG [f.p.e.c.f.FsCrawlerImpl] Deleting from ES jobs, folder, eec8ebf874bb5f54d44cace29601ed0
16:07:47,361 DEBUG [f.p.e.c.f.c.BulkProcessor] {"delete":{"_index":"jobs","_type":"folder","_id":"eec8ebf874bb5f54d44cace29601ed0"}}
16:07:47,362 DEBUG [f.p.e.c.f.FsCrawlerImpl] Fs crawler is going to sleep for 1m
16:07:51,907 DEBUG [f.p.e.c.f.c.BulkProcessor] Going to execute new bulk composed of 1 actions
16:07:51,913 DEBUG [f.p.e.c.f.c.ElasticsearchClient] bulk response: BulkResponse{items=[BulkItemTopLevelResponse{index=null, delete=BulkItemResponse{failed=false, index='jobs', type='folder', id='eec8ebf874bb5f54d44cace29601ed0', opType=null, failureMessage='null'}}]}
16:07:51,913 DEBUG [f.p.e.c.f.c.BulkProcessor] Executed bulk composed of 1 actions

```

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 17, 2017, 8:10am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/15 "2017-02-17T08:10:32Z")

</div>

@dadoonet 2nd log after adding new document (bago5.pdf)

```auto
16:09:47,394 DEBUG [f.p.e.c.f.FsCrawlerImpl] Fs crawler is now waking up again...
16:09:47,396 DEBUG [f.p.e.c.f.FsCrawlerImpl] Fs crawler thread [jobs] is now running. Run #4...
16:09:47,401 DEBUG [f.p.e.c.f.FsCrawlerImpl] indexing [/mapr/my.cluster.com/vm1] content
16:09:47,402 DEBUG [f.p.e.c.f.f.FileAbstractor] Listing local files from /mapr/my.cluster.com/vm1
16:09:47,404 DEBUG [f.p.e.c.f.f.FileAbstractor] 2 local files found
16:09:47,405 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null], excludes = [[~*]]
16:09:47,405 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], excludes = [[~*]]
16:09:47,405 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null]
16:09:47,405 DEBUG [f.p.e.c.f.FsCrawlerImpl] [changelog.txt] can be indexed: [true]
16:09:47,405 DEBUG [f.p.e.c.f.FsCrawlerImpl] - file: changelog.txt
16:09:47,406 DEBUG [f.p.e.c.f.FsCrawlerImpl] - not modified: creation date 2016-12-12T13:48:24 , file date 2016-12-12T13:48:24, last scan date 2017-02-17T16:08:45.365
16:09:47,406 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [bago5.pdf], includes = [null], excludes = [[~*]]
16:09:47,406 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [bago5.pdf], excludes = [[~*]]
16:09:47,406 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [bago5.pdf], includes = [null]
16:09:47,406 DEBUG [f.p.e.c.f.FsCrawlerImpl] [bago5.pdf] can be indexed: [true]
16:09:47,406 DEBUG [f.p.e.c.f.FsCrawlerImpl] - file: bago5.pdf
16:09:47,406 DEBUG [f.p.e.c.f.FsCrawlerImpl] - not modified: creation date 2017-02-16T11:08:43 , file date 2017-02-16T11:08:43, last scan date 2017-02-17T16:08:45.365
16:09:47,406 DEBUG [f.p.e.c.f.FsCrawlerImpl] Looking for removed files in [/mapr/my.cluster.com/vm1]...
16:09:47,407 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[doc], request [SearchRequest{query=path.encoded:66e1f91ce6a0761b736dbb8117e542e, fields=[_source, file.filename], size=10000}]
16:09:47,414 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null], excludes = [[~*]]
16:09:47,414 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], excludes = [[~*]]
16:09:47,415 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [changelog.txt], includes = [null]
16:09:47,415 DEBUG [f.p.e.c.f.FsCrawlerImpl] Looking for removed directories in [/mapr/my.cluster.com/vm1]...
16:09:47,415 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[folder], request [SearchRequest{query=encoded:66e1f91ce6a0761b736dbb8117e542e, fields=[], size=10000}]
16:09:47,420 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [/mapr/my.cluster.com/vm1], includes = [null], excludes = [[~*]]
16:09:47,420 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [/mapr/my.cluster.com/vm1], excludes = [[~*]]
16:09:47,420 DEBUG [f.p.e.c.f.u.FsCrawlerUtil] filename = [/mapr/my.cluster.com/vm1], includes = [null]
16:09:47,420 DEBUG [f.p.e.c.f.FsCrawlerImpl] Delete folder /mapr/my.cluster.com/vm1//mapr/my.cluster.com/vm1
16:09:47,421 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[doc], request [SearchRequest{query=path.encoded:eec8ebf874bb5f54d44cace29601ed0, fields=[_source, file.filename], size=10000}]
16:09:47,424 DEBUG [f.p.e.c.f.c.ElasticsearchClient] search [jobs]/[folder], request [SearchRequest{query=encoded:eec8ebf874bb5f54d44cace29601ed0, fields=[], size=10000}]
16:09:47,427 DEBUG [f.p.e.c.f.FsCrawlerImpl] Deleting from ES jobs, folder, eec8ebf874bb5f54d44cace29601ed0
16:09:47,427 DEBUG [f.p.e.c.f.c.BulkProcessor] {"delete":{"_index":"jobs","_type":"folder","_id":"eec8ebf874bb5f54d44cace29601ed0"}}
16:09:47,428 DEBUG [f.p.e.c.f.FsCrawlerImpl] Fs crawler is going to sleep for 1m
16:09:51,931 DEBUG [f.p.e.c.f.c.BulkProcessor] Going to execute new bulk composed of 1 actions
16:09:51,937 DEBUG [f.p.e.c.f.c.ElasticsearchClient] bulk response: BulkResponse{items=[BulkItemTopLevelResponse{index=null, delete=BulkItemResponse{failed=false, index='jobs', type='folder', id='eec8ebf874bb5f54d44cace29601ed0', opType=null, failureMessage='null'}}]}
16:09:51,938 DEBUG [f.p.e.c.f.c.BulkProcessor] Executed bulk composed of 1 actions

```

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 17, 2017, 8:24am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/16 "2017-02-17T08:24:06Z")

</div>

@dadoonet i notice that on the second log the changelog.txt repeats after (changelog.txt) can be indexed:[true] while the bago5.pdf didn't.

\*sigh I dont know whats going on I tried repeating every process

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 17, 2017, 8:35am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/17 "2017-02-17T08:35:32Z")

</div>

Please format your code using `</>` icon as explained in [this guide](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will make your post more readable.

Or use markdown style like:

````
```
CODE
```

````

I updated your answers.

The problem here is

```auto
creation date: 2017-02-16T11:08:43
file date : 2017-02-16T11:08:43
last scan date: 2017-02-17T16:08:45.365

```

So your FS does not change the file modification date when you move it to the folder.  
The only way I believe to fix it is to do:

```auto
touch /mapr/my.cluster.com/vm1/bago5.pdf

```

So it will get a more recent date.

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 17, 2017, 8:44am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/18 "2017-02-17T08:44:51Z")

</div>

But this directories are NFS mounted. Which the user only drag and drop the files. So the only way is to Touch?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 17, 2017, 8:56am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/19 "2017-02-17T08:56:11Z")

</div>

Not sure if I can do anything on FSCrawler side.

May be I need to detect what kind of implementation is the underlying FS and find if anything in the Java API can help.

But if I don't have any information about the fact that a file has been added, I can't detect that.

May be the parent directory changed? So I could detect it? But really unsure.

If you have yourself a way to detect changes easily like with a shell script, then you can think of using FSCrawler as a gateway to elasticsearch and activate the REST endpoint.

See [https://github.com/dadoonet/fscrawler#rest-service](https://github.com/dadoonet/fscrawler#rest-service)

Then you can send within your script a file with:

```auto
curl -F "file=@/path/to/yourfile.txt" "http://127.0.0.1:8080/fscrawler/_upload"

```

I hope this helps.

Can you open an issue in FSCrawler with all details (sounds like you are using MapR) and a scenario to reproduce it? I'll try to play with MapR if time allows.

---

<div class="post-metadata">

**Author:** ![Mark\_Dendrix\_Garcia](https://avatars.discourse-cdn.com/v4/letter/m/8baadc/32.png) [@Mark\_Dendrix\_Garcia](https://discuss.elastic.co/u/Mark_Dendrix_Garcia)\
**Post date:** [February 17, 2017, 9:02am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/20 "2017-02-17T09:02:54Z")

</div>

@dadoonet

I see its a bit clearer now, so FSCrawler first run assures that all files are index because it doesnt look on timestamps, wherein the second run will check all the files timestamp and determine what files are new and indexed it?

Is this correct analogy ?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 17, 2017, 9:17am UTC](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128/21 "2017-02-17T09:17:23Z")

</div>

exact. You can always restart from scratch by using `--restart` option which will remove the status file and will reindex everything.

[Next page](https://discuss.elastic.co/t/ingesting-documents-pdf-word-txt-to-elasticsearch/75128.md?page=2)
