# Grep-like results on elasticsearch index

**URL:** <https://discuss.elastic.co/t/grep-like-results-on-elasticsearch-index/326524>\
**Category:** Elasticsearch\
**Created:** [February 26, 2023, 4:20pm UTC](https://discuss.elastic.co/t/grep-like-results-on-elasticsearch-index/326524 "2023-02-26T16:20:55Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![John10](https://avatars.discourse-cdn.com/v4/letter/j/71c47a/32.png) [@John10](https://discuss.elastic.co/u/John10)\
**Post date:** [February 26, 2023, 4:20pm UTC](https://discuss.elastic.co/t/grep-like-results-on-elasticsearch-index/326524/1 "2023-02-26T16:20:55Z")

</div>

If you have 5,000 pdf documents and you want to return every instance of the word "dog" across all of those documents (including the context where it occurs -- page number, the line before and after the match, etc.), you could use a grep-like utility (pdfgrep, etc.); however, this is (relatively) slow and doesn't use any sort of index.

Elasticsearch works at the document level and returns documents as opposed to individual matches _within and across_ each and every document. The use-case is just different. It looks like it might be possible to have fscrawler index something like paragraphs or sentences as nested objects of the document and then return the hits within those nested objects across the entire index.

1. Is that really feasible?

2. Would that actually be faster than just using grep? (I don't see how it couldn't be when dealing with thousands of documents but I wonder.)

3. Is there some other obvious solution I'm missing?

Basically, what I'm after is using grep with a full-text index to get grep-like results with the speed advantages of a full-text index.

Thanks,  
John

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 26, 2023, 7:42pm UTC](https://discuss.elastic.co/t/grep-like-results-on-elasticsearch-index/326524/2 "2023-02-26T19:42:45Z")

</div>

FSCrawler can not (yet) index paragraphs or even pages.  
So it will just help to know in which document you can see the term.

That said you can use the [highlighter](https://www.elastic.co/guide/en/elasticsearch/reference/current/highlighting.html) to get more information about the context.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [March 26, 2023, 7:43pm UTC](https://discuss.elastic.co/t/grep-like-results-on-elasticsearch-index/326524/3 "2023-03-26T19:43:09Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
