# FSCrawler - Elastic Mapping Changes

**URL:** <https://discuss.elastic.co/t/fscrawler-elastic-mapping-changes/226867>\
**Category:** Elasticsearch\
**Created:** [April 7, 2020, 10:01am UTC](https://discuss.elastic.co/t/fscrawler-elastic-mapping-changes/226867 "2020-04-07T10:01:07Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Sarath\_Pullabhotla](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sarath_pullabhotla/32/65563_2.png) [@Sarath\_Pullabhotla](https://discuss.elastic.co/u/Sarath_Pullabhotla)\
**Post date:** [April 7, 2020, 10:01am UTC](https://discuss.elastic.co/t/fscrawler-elastic-mapping-changes/226867/1 "2020-04-07T10:01:07Z")

</div>

Hi,

Can anyone please guide me on the below.

Q1)

Is it possible for us to change the tag names in elastic search while ingesting documents using fscrawler.

Example: The content that a file has is getting ingested under "\_source.content" tag in elastic. Can we change this "\_source.content" to "\_source.passage".

Q2)

I believe fscrawler is internally using Apache TIKA for extracting metadata and content from a file. Is there any possibility to split the content that we are saving in elastic search? Instead of saving under a single tag "\_source.content", is there any existing way to split and save the document (page/paragraph wise) under same/different index.

Regards,  
Sarath

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 7, 2020, 11:18am UTC](https://discuss.elastic.co/t/fscrawler-elastic-mapping-changes/226867/2 "2020-04-07T11:18:26Z")

</div>

> [@Sarath\_Pullabhotla](#):
>
> Is it possible for us to change the tag names in Elasticsearch while ingesting documents using fscrawler.

You can define an [ingest pipeline](https://www.elastic.co/guide/en/elasticsearch/reference/current/ingest.html) which uses a [rename processor](https://www.elastic.co/guide/en/elasticsearch/reference/current/rename-processor.html).  
Then [set this pipeline in FSCrawler](https://fscrawler.readthedocs.io/en/latest/admin/fs/elasticsearch.html#ingest-node).

> [@Sarath\_Pullabhotla](#):
>
> Is there any possibility to split the content that we are saving in Elasticsearch?

No it's not possible today. There have been a similar ask here:

> <https://github.com/dadoonet/fscrawler/issues/795>
>
> Hello,
> 
> After some tests with words documents, i see that all the content are …saved in \_source": { "content": field under elasticSearch.
> Is there a way to create an elasticSearch doc by Words Paragraph/Title.
> 
> For example :
> 
> \*\*1 - titre 1\*\*
> content 1
> \*\*2 - titre 2\*\*
> content 2
> ...
> Will generate 2 elastics doc : 
> 
> Doc 1
> "titre" : "titre 1"
> "content" : "content 1"
> and some global metadata such as docpath, creationdate ...
> 
> Doc 2
> "titre" : "titre 2"
> "content" : "content 2"
> and some global metadata such as docpath, creationdate ...
> 
> I know that i can do it by parsing the XML file of docs words and genearte a bulk with that data.
> 
> But is it possible to do that with some fscrawler configuration ?
> 
> Thanks,
> Olivier

IIRC it could be done but for sure this is not going to be implemented anytime soon.  
Best option for now would be to preprocess the PDF document, and generate one file per page before starting FSCrawler or calling its REST API.

---

<div class="post-metadata">

**Author:** ![Sarath\_Pullabhotla](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sarath_pullabhotla/32/65563_2.png) [@Sarath\_Pullabhotla](https://discuss.elastic.co/u/Sarath_Pullabhotla)\
**Post date:** [April 7, 2020, 12:17pm UTC](https://discuss.elastic.co/t/fscrawler-elastic-mapping-changes/226867/3 "2020-04-07T12:17:44Z")

</div>

Thank you David.

Also, when FScrawler crawls through all the documents for indexing them, will the content of the document saved on my Elastic cluster?? or is there any way it works internally?  
Because even when I stop my crawler I will still be able to read the content while searching.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 7, 2020, 2:19pm UTC](https://discuss.elastic.co/t/fscrawler-elastic-mapping-changes/226867/4 "2020-04-07T14:19:51Z")

</div>

Everything is sent to elasticsearch and indexed and stored there.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 5, 2020, 2:19pm UTC](https://discuss.elastic.co/t/fscrawler-elastic-mapping-changes/226867/5 "2020-05-05T14:19:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
