# How can i disable content extraction?

**URL:** <https://discuss.elastic.co/t/how-can-i-disable-content-extraction/355428>\
**Category:** Elastic Search\
**Tags:** elastic-app-search\
**Created:** [March 14, 2024, 4:55pm UTC](https://discuss.elastic.co/t/how-can-i-disable-content-extraction/355428 "2024-03-14T16:55:51Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![maddy30](https://avatars.discourse-cdn.com/v4/letter/m/51bf81/32.png) [@maddy30](https://discuss.elastic.co/u/maddy30)\
**Post date:** [March 14, 2024, 4:55pm UTC](https://discuss.elastic.co/t/how-can-i-disable-content-extraction/355428/1 "2024-03-14T16:55:51Z")

</div>

I am using App search based engines. By default my crawler extracts the content from web pages and pdf's. But when i am running the crawl for one particular app search engine, i only want the meta data of both the web pages and pdf's to be extracted but not the content from it. how can i achieve it? any help would be appreciated. Thanks.

---

<div class="post-metadata">

**Author:** ![Sean\_Story](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/sean_story/32/69987_2.png) [@Sean\_Story](https://discuss.elastic.co/u/Sean_Story)\
**Post date:** [March 14, 2024, 5:56pm UTC](https://discuss.elastic.co/t/how-can-i-disable-content-extraction/355428/2 "2024-03-14T17:56:58Z")

</div>

Hi @maddy30 ,

Looks like this might be related to your other question here: [How can i update the pipeline used for a app search engine?](https://discuss.elastic.co/t/how-can-i-update-the-pipeline-used-for-a-app-search-engine/355427)

The configurations to extract content from files (like PDFs) are made at a deployment level, not on an engine-by-engine basis. What you could do is [add conditionals to your ingest pipeline](https://www.elastic.co/guide/en/elasticsearch/reference/current/ingest.html#conditionally-run-processor) to run certain processors only if the URL matches a certain domain or pattern.

Alternatively, you can take the approach I suggest in the other post to use different pipelines per index, and have some pipelines remove the `body_content` from your documents before indexing it.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 11, 2024, 5:57pm UTC](https://discuss.elastic.co/t/how-can-i-disable-content-extraction/355428/3 "2024-04-11T17:57:44Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
