# Fscrawler creating custome mapping

**URL:** <https://discuss.elastic.co/t/fscrawler-creating-custome-mapping/167950>\
**Category:** Elasticsearch\
**Created:** [February 12, 2019, 4:06am UTC](https://discuss.elastic.co/t/fscrawler-creating-custome-mapping/167950 "2019-02-12T04:06:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![vikas\_singhji](https://avatars.discourse-cdn.com/v4/letter/v/53a042/32.png) [@vikas\_singhji](https://discuss.elastic.co/u/vikas_singhji)\
**Post date:** [February 12, 2019, 4:06am UTC](https://discuss.elastic.co/t/fscrawler-creating-custome-mapping/167950/1 "2019-02-12T04:06:52Z")

</div>

I'm using FSCrawler 2.6 and it's working great for indexing the pdf document. However one issue i am facing is that it is putting all the contents of the PDF in the "content" field(i am newbie in this field). So, my question is that "is there any way that i can have my custom mapping for the data of pdf i.e. latitude/longitude, number or if not that... line wise(like content.line1, content.line2, content.line3...) ?".

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 12, 2019, 4:35am UTC](https://discuss.elastic.co/t/fscrawler-creating-custome-mapping/167950/2 "2019-02-12T04:35:09Z")

</div>

Parsing text to extract meaningful content (entities) is a difficult thing.  
The only option I can see for now is by using

> **[spinscale/elasticsearch-ingest-opennlp](https://github.com/spinscale/elasticsearch-ingest-opennlp)**
>
> An Elasticsearch ingest processor to do named entity extraction using Apache OpenNLP - spinscale/elasticsearch-ingest-opennlp

In FSCrawler you can configure the ingest pipeline name to apply after the text has been extracted. See [https://fscrawler.readthedocs.io/en/latest/admin/fs/elasticsearch.html](https://fscrawler.readthedocs.io/en/latest/admin/fs/elasticsearch.html)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 12, 2019, 4:35am UTC](https://discuss.elastic.co/t/fscrawler-creating-custome-mapping/167950/3 "2019-03-12T04:35:11Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
