# Remove strings from body\_content in Crawler

**URL:** <https://discuss.elastic.co/t/remove-strings-from-body-content-in-crawler/295287>\
**Category:** Elastic Search\
**Created:** [January 24, 2022, 10:10pm UTC](https://discuss.elastic.co/t/remove-strings-from-body-content-in-crawler/295287 "2022-01-24T22:10:32Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ansamHox](https://avatars.discourse-cdn.com/v4/letter/a/54ee81/32.png) [@ansamHox](https://discuss.elastic.co/u/ansamHox)\
**Post date:** [January 24, 2022, 10:10pm UTC](https://discuss.elastic.co/t/remove-strings-from-body-content-in-crawler/295287/1 "2022-01-24T22:10:32Z")

</div>

I have a WordPress page, and crawler always puts Menu, Footer, and all of it in `body_content` field,e.g.

`Company name, Company description Home What We Do Mission Our Process Technology Who We Are About Us Our Team Community Get Involved Students Corner Donations Portal Support the Project News Press News & Events Newsletters Get Social Contact MAIN CONTENT OF ARTICLE/PAGE.. All rights reserved. Site by XXX..REST OF THE FOOTER `

Any ideas how to remove this from Crawling, or how to completely ignore these strings in search / mappings? Thanks

---

<div class="post-metadata">

**Author:** ![Sean\_Story](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sean_story/32/69987_2.png) [@Sean\_Story](https://discuss.elastic.co/u/Sean_Story)\
**Post date:** [February 1, 2022, 10:27pm UTC](https://discuss.elastic.co/t/remove-strings-from-body-content-in-crawler/295287/2 "2022-02-01T22:27:13Z")

</div>

Hi @ansamHox ,

I think that [these instructions for meta tags](https://www.elastic.co/guide/en/app-search/current/web-crawler-reference.html#web-crawler-reference-meta-tags-content-extraction) is what would help you avoid this. However, I'm not familiar with the level of control you have over WordPress's HTML. Right now (v7.16.3) if you don't have control over the source HTML of a website, there's not a lot you can do to control the crawled content. However, this is a feature that is being worked on, so stay tuned in future versions for other content extraction controls.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 4, 2022, 8:33am UTC](https://discuss.elastic.co/t/remove-strings-from-body-content-in-crawler/295287/3 "2022-11-04T08:33:54Z")

</div>


