# How to remove \\n and \\t chars when using FS Crawler?

**URL:** <https://discuss.elastic.co/t/how-to-remove-n-and-t-chars-when-using-fs-crawler/197194>\
**Category:** Elasticsearch\
**Created:** [August 28, 2019, 7:18pm UTC](https://discuss.elastic.co/t/how-to-remove-n-and-t-chars-when-using-fs-crawler/197194 "2019-08-28T19:18:14Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Danilo\_Felix](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danilo_felix/32/53204_2.png) [@Danilo\_Felix](https://discuss.elastic.co/u/Danilo_Felix)\
**Post date:** [August 28, 2019, 7:18pm UTC](https://discuss.elastic.co/t/how-to-remove-n-and-t-chars-when-using-fs-crawler/197194/1 "2019-08-28T19:18:15Z")

</div>

Hello all,

I'm new with FS Crawler, so please be patient with me.

Using version 2.6, I'm crawling many files from a big Windows's folders tree. There are PDF, XLS, DOC, etc.

It's working, but the content attribute goes to Elasticsearch full of /t and /n characters.

```
  {
    "_index": "clgp-docs",
    "_type": "_doc",
    "_id": "c2c62de2f7d5af413ed074c845129751",
    "_score": 1.0,
    "_source": {
      "content": "\nHome\nHistória »\nSaudade\nSindipetro »\nDocumentos »\nSindicalize-se\n\n \n\nNotícias por base »\nEleição 2015 – 2018\nNotícias por assunto »\nBroncas do Petrolino\nAgenda da Diretoria\nTV Petroleira\n\nCategorizado | ACT, Agenda da Diretoria, Direitos, Petrobrás, Reuniões com RHs\n\nSindipetro-LP cobra soluções para problemas relacionados ao Benefício Educacional\n\nPostado em19 fevereiro 2016.\n\nO Benefício Educacional, uma das conquistas da categoria e importante ferramenta...  

```

Is it possible to change these special chars to space chars during the ingestion process? How to do it?

I aprecciatte if you could expose a \_setting.json example file.

Thanks

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 30, 2019, 3:57am UTC](https://discuss.elastic.co/t/how-to-remove-n-and-t-chars-when-using-fs-crawler/197194/2 "2019-08-30T03:57:01Z")

</div>

Welcome!

I believe that the only option from now is to use an ingest pipeline which does the removal at index time on elasticsearch side. See [https://fscrawler.readthedocs.io/en/latest/admin/fs/elasticsearch.html#using-ingest-node-pipeline](https://fscrawler.readthedocs.io/en/latest/admin/fs/elasticsearch.html#using-ingest-node-pipeline)

May be with this processor [https://www.elastic.co/guide/en/elasticsearch/reference/7.3/gsub-processor.html](https://www.elastic.co/guide/en/elasticsearch/reference/7.3/gsub-processor.html)  
or this one [https://www.elastic.co/guide/en/elasticsearch/reference/7.3/script-processor.html](https://www.elastic.co/guide/en/elasticsearch/reference/7.3/script-processor.html)

Could you also open an issue in FSCrawler project so I can think about adding a regex parser in the future in FSCrawler itself?

---

<div class="post-metadata">

**Author:** ![Danilo\_Felix](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danilo_felix/32/53204_2.png) [@Danilo\_Felix](https://discuss.elastic.co/u/Danilo_Felix)\
**Post date:** [August 30, 2019, 1:56pm UTC](https://discuss.elastic.co/t/how-to-remove-n-and-t-chars-when-using-fs-crawler/197194/3 "2019-08-30T13:56:30Z")

</div>

Thanks about the ideia! I'll try it.

This would be a good improvement to the next versions os FS Crawler. I'll open a issue there, as you sugested.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 27, 2019, 1:56pm UTC](https://discuss.elastic.co/t/how-to-remove-n-and-t-chars-when-using-fs-crawler/197194/4 "2019-09-27T13:56:31Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
