# Elastic App Search Crawler

**URL:** <https://discuss.elastic.co/t/elastic-app-search-crawler/351699>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [January 24, 2024, 9:07am UTC](https://discuss.elastic.co/t/elastic-app-search-crawler/351699 "2024-01-24T09:07:24Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![\_Pontes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/_pontes/32/131074_2.png) [@\_Pontes](https://discuss.elastic.co/u/_Pontes)\
**Post date:** [January 24, 2024, 9:07am UTC](https://discuss.elastic.co/t/elastic-app-search-crawler/351699/1 "2024-01-24T09:07:24Z")

</div>

Hi there,  
Is there a way to make the Elasticsearch crawler from indexing the content of and HTML tags and their content?  
Specifically, we'd like to remove them from the headings and main\_content (extracted by default). By default, the tags are stripped, but its contents are kept (which seems odd for these types of tag).  
I am aware of the data-elastic-exclude attribute, but in our case, these tags are generated automatically by a framework over which we have limited control. We considered regex, but because the tags are extracted and processed by default this isn't feasible. If this is not possible, I suppose we could extract the headings into a separate fields, and apply regex to remove the and content, though we'd prefer to use the existing default fields.

---

<div class="post-metadata">

**Author:** ![Sander\_Philipse](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sander_philipse/32/101890_2.png) [@Sander\_Philipse](https://discuss.elastic.co/u/Sander_Philipse)\
**Post date:** [January 26, 2024, 8:03pm UTC](https://discuss.elastic.co/t/elastic-app-search-crawler/351699/2 "2024-01-26T20:03:30Z")

</div>

Hi @_Pontes, we usually recommend using ingest pipelines to manipulate the content on ingest: [Customize crawler field values using an ingest pipeline | Enterprise Search documentation [8.12] | Elastic](https://www.elastic.co/guide/en/enterprise-search/current/crawler-custom-values-ingest-pipeline.html)

---

<div class="post-metadata">

**Author:** ![\_Pontes](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/_pontes/32/131074_2.png) [@\_Pontes](https://discuss.elastic.co/u/_Pontes)\
**Post date:** [January 29, 2024, 8:37am UTC](https://discuss.elastic.co/t/elastic-app-search-crawler/351699/3 "2024-01-29T08:37:12Z")

</div>

Yes, that's the approach we ended up following, used a GSUB step in the pipeline to extract the undesired CSS.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 26, 2024, 8:37am UTC](https://discuss.elastic.co/t/elastic-app-search-crawler/351699/4 "2024-02-26T08:37:19Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
