# Index and query source code

**URL:** <https://discuss.elastic.co/t/index-and-query-source-code/38979>\
**Category:** Elasticsearch\
**Created:** [January 12, 2016, 12:25pm UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979 "2016-01-12T12:25:21Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Joris\_Rau](https://avatars.discourse-cdn.com/v4/letter/j/e9a140/32.png) [@Joris\_Rau](https://discuss.elastic.co/u/Joris_Rau)\
**Post date:** [January 12, 2016, 12:25pm UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/1 "2016-01-12T12:25:21Z")

</div>

Hi,

I want to use Elasticsearch to use a regexp query on source code. However I could not find any information about indexing or querying source code using elasticsearch. I know that GitHub is using Elasticsearch for their own source search. So what's the best way to go?

Regards,

Joris

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [January 14, 2016, 6:48am UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/2 "2016-01-14T06:48:44Z")

</div>

That's a pretty broad question, unfortunately we wouldn't be able to share any of how GH does this as it's their proprietary information.

But the biggest thing is likely to be analysis and what sort of regexp querying you want to do, as that sort of query is going to be pretty resource intensive.

---

<div class="post-metadata">

**Author:** ![Joris\_Rau](https://avatars.discourse-cdn.com/v4/letter/j/e9a140/32.png) [@Joris\_Rau](https://discuss.elastic.co/u/Joris_Rau)\
**Post date:** [January 14, 2016, 10:58am UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/3 "2016-01-14T10:58:08Z")

</div>

Thanks for your reply!

I was thinking about using a 3-gram analyzer for the source code and a different 3-gram analyzer which creates 3-grams out of regex queries. Is that a good idea? Do you have a hint for me which could point me to a better solution?

Regards,

Joris

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [January 14, 2016, 11:09am UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/4 "2016-01-14T11:09:59Z")

</div>

You might want to look at something that works with camelCase e.g. [https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-pattern-analyzer.html#\_camelcase\_tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-pattern-analyzer.html#_camelcase_tokenizer)

---

<div class="post-metadata">

**Author:** ![Joris\_Rau](https://avatars.discourse-cdn.com/v4/letter/j/e9a140/32.png) [@Joris\_Rau](https://discuss.elastic.co/u/Joris_Rau)\
**Post date:** [January 14, 2016, 11:33am UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/5 "2016-01-14T11:33:19Z")

</div>

Thanks. That's an interesting idea. In my case I still want to be able to use all the special characters for the search. So maybe I should combine that with a hierarchical anaylzer? Is that a good idea?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [January 14, 2016, 11:39am UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/6 "2016-01-14T11:39:04Z")

</div>

Using multiple indexing strategies is a common technique. See [https://www.elastic.co/guide/en/elasticsearch/reference/2.1/multi-fields.html#\_multi\_fields\_with\_multiple\_analyzers](https://www.elastic.co/guide/en/elasticsearch/reference/2.1/multi-fields.html#_multi_fields_with_multiple_analyzers)

---

<div class="post-metadata">

**Author:** ![Joris\_Rau](https://avatars.discourse-cdn.com/v4/letter/j/e9a140/32.png) [@Joris\_Rau](https://discuss.elastic.co/u/Joris_Rau)\
**Post date:** [January 14, 2016, 11:57am UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/7 "2016-01-14T11:57:57Z")

</div>

Alright. Awesome. Thanks!  
One last question: I take it there is no proper way to run a full text regex query in elasticsearch, right? I am just asking because you already mention it in the docs that the regexp query runs only on terms.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [January 14, 2016, 12:16pm UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/8 "2016-01-14T12:16:53Z")

</div>

Full text regex query would be slow so yes, we only work with terms produced by the tokenization process.

---

<div class="post-metadata">

**Author:** ![Joris\_Rau](https://avatars.discourse-cdn.com/v4/letter/j/e9a140/32.png) [@Joris\_Rau](https://discuss.elastic.co/u/Joris_Rau)\
**Post date:** [January 14, 2016, 12:29pm UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/9 "2016-01-14T12:29:03Z")

</div>

Thanks. Now I feel enlightened 🙂 .

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:24pm UTC](https://discuss.elastic.co/t/index-and-query-source-code/38979/10 "2017-07-05T23:24:32Z")

</div>


