# Searching on multiple fields from Index created by Fscrawler

**URL:** <https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275>\
**Category:** Elasticsearch\
**Created:** [October 21, 2018, 8:52am UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275 "2018-10-21T08:52:53Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jasmeet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasmeet/32/69393_2.png) [@Jasmeet](https://discuss.elastic.co/u/Jasmeet)\
**Post date:** [October 21, 2018, 8:52am UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275/1 "2018-10-21T08:52:53Z")

</div>

Hi, I am new to Elasticsearch and have tried out creating index with Fscrawler. After creating custom Analyzers, when i try to search on the fields"content.phonetic" and "content.shingle", i do not get a search hit. Can someone please guide me to where i am going wrong.  
The \_settings.json file used in Fscrawler is given below

//MY CODE

```auto
{
  "settings": {
    "index.mapping.total_fields.limit": 2000,
    "analysis": {
      "analyzer": {
        "default": {
			"type": "custom",
			"tokenizer": "standard",
			"filter": ["lowercase","custom_edge_ngram"]
				},
				
			"dbl_metaphone":{
			"type": "custom",
			"tokenizer": "standard",
			"filter": ["dbl_metaphone"]
				},	
				
			"shingle":{
			"type": "custom",
			"tokenizer": "standard",
			"filter": ["shingle-filter"]
				}	
			
				},
"filter": {

			"custom_edge_ngram": {
				"type": "edge_ngram",
				"min_gram": 2,
				"max_gram": 10
				},
			"dbl_metaphone": {
              "type": "phonetic",
              "encoder": "double_metaphone"
            },		
			"shingle-filter": {
				"max_shingle_size": "5",
				"min_shingle_size": "2",
				"output_unigrams": "false",
				"type": "shingle"
			         }
			}}}},}
},

 "mappings": {
    "_doc": {
      "dynamic_templates": [
        {
          "raw_as_text": {
            "path_match": "meta.raw.*",
            "mapping": {
              "type": "text",
              "fields": {
                  "keyword": {
                  "type": "keyword",
                  "ignore_above": 256
                		} } } } } ],
      "properties": {
        "attachment": {
          "type": "binary",
          "doc_values": false
        },
        "attributes": {
          "properties": {
            "group": {
              "type": "keyword"
            },
            "owner": {
              "type": "keyword"
            }
          }
        },
        "content": {
          "type": "text"
		  "index_analyzer": "default",
		  "search_analyzer" : "standard",
				"fields:{
					"phonetic":{
					"type":"text",
					"analyzer":"dbl_metaphone"
								},
					"shingle":{
					"type":"text",
					"analyzer":"shingle"
					}
					}
					},

```

code continues ----------------------------------

-----ADDING DOCUMENT TO TEST INDEX--------

```auto
POST /test/_doc
{
"content": "Learning Elastic Stack 6"
}

```

TRYING TO SEARCH ON INDEX-----\>No result on "content.phonetic" and "content.shingle" but result obtained on the field "content".

```auto
GET /test/_search
{
 "query": {
    "multi_match": {
       "query": "learning",
       "fields": ["content.phonetic"]
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 29, 2018, 6:54am UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275/2 "2018-10-29T06:54:39Z")

</div>

Could you provide a full recreation script as described in [About the Elasticsearch category](https://discuss.elastic.co/t/about-the-elasticsearch-category/21). It will help to better understand what you are doing. Please, try to keep the example as simple as possible.

A full reproduction script will help readers to understand, reproduce and if needed fix your problem. It will also most likely help to get a faster answer.

Please make it as simple as possible. There is no need here for the tons of mappings you pasted to reproduce your problem.  
Also, I suggest that you try the `_analyze` API to understand how elasticsearch is indexing your text. That's normally helping a lot.

---

<div class="post-metadata">

**Author:** ![Jasmeet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasmeet/32/69393_2.png) [@Jasmeet](https://discuss.elastic.co/u/Jasmeet)\
**Post date:** [October 29, 2018, 5:07pm UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275/3 "2018-10-29T17:07:23Z")

</div>

Thanks, I will try \_analyze API as suggested.  
Meanwhile, could you please guide me on the following issues-

1. How to identify which files have been indexed by fscrawler. When i tried indexing some files in a folder, the number of files indexed in elasticsearch is different from the total number of files present in the folder. --Trace option does not clearly list the non-indexed/skipped files or a list of files crawled.
2. Is there a way to pause crawling in fscralwer or restart from where the last crawling stopped?
3. I am using fscrawler on a Windows machine and the documentation of fscrawler is not very clear on 'touch' command on files. Will all the files in a folder be crawled irrespective of the date of creation ?
4. When i restart fscrawler, it appears to crawl all files in the folder irrespective of whether they were previously indexed. Is there a way to tell the fscralwer not to crawl the previously indexed files and look only for new files in the folder?  
Thanks in advance  
JS

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 31, 2018, 4:13pm UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275/4 "2018-10-31T16:13:48Z")

</div>

> How to identify which files have been indexed by fscrawler

I guess that the only way to do that is by searching in elasticsearch, gathering all the filenames and compare to what `tree` would give back?

> --Trace option does not clearly list the non-indexed/skipped files or a list of files crawled.

I think that `--trace` or `--debug` are printing what are the files meant to be indexed and if we skip or not the indexation.

> Is there a way to pause crawling in fscralwer or restart from where the last crawling stopped?

No.

> <https://github.com/dadoonet/fscrawler/issues/493>
>
> Hello team.
> Is there any way to force fscrawler to continue his work since the …last file, that was crawled, if some exception happens.
> Example: we have a very big folder (about 30TB). I run crawler. Crawler has been working for 1 week, for example, about 3 millions files were crawled and after this some network error happens and crawler throws some timeout or another network exception. In this case I have to delete index from elasticsearch and run crawler again or run crawler with --restart flag (and I don't know exactly, what happens here).
> So I don't know, how to force crawler to continue his work since the last place, when he was stopped. It takes huge time and we are still not able to crawl this folder for at least 2 months, because in case with any network exception we have to start from the start.
> We use option in config file: "continue\_on\_error" : true, but it doesn't help in case with unhandled exception. Crawler just stops his work.
> Is there any way to solve our problem?

> I am using fscrawler on a Windows machine and the documentation of fscrawler is not very clear on 'touch' command on files. Will all the files in a folder be crawled irrespective of the date of creation ?

Yes at the first run. Then on the next run, only files that changed will be indexed. Unless you use `--restart` to restart indexing all files.

> When i restart fscrawler, it appears to crawl all files in the folder irrespective of whether they were previously indexed. Is there a way to tell the fscralwer not to crawl the previously indexed files and look only for new files in the folder?

That's what FSCrawler is supposed to be doing. But for that the first run needs to have been completed so FSCrawler can write on disk the last run date.  
If this status file is not existing, FSCrawler will start again from scratch and will reindex all.

---

<div class="post-metadata">

**Author:** ![Jasmeet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jasmeet/32/69393_2.png) [@Jasmeet](https://discuss.elastic.co/u/Jasmeet)\
**Post date:** [October 31, 2018, 5:23pm UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275/5 "2018-10-31T17:23:33Z")

</div>

Thanks a ton. FS has helped a lot. Hope more features are added to it..

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [October 31, 2018, 5:42pm UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275/6 "2018-10-31T17:42:23Z")

</div>

Sure. PR are welcomed! 😉

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 28, 2018, 5:42pm UTC](https://discuss.elastic.co/t/searching-on-multiple-fields-from-index-created-by-fscrawler/153275/7 "2018-11-28T17:42:24Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
