# How to source search queries from a file

**URL:** <https://discuss.elastic.co/t/how-to-source-search-queries-from-a-file/142440>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [July 31, 2018, 8:54pm UTC](https://discuss.elastic.co/t/how-to-source-search-queries-from-a-file/142440 "2018-07-31T20:54:00Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![username123](https://avatars.discourse-cdn.com/v4/letter/u/7ba0ec/32.png) [@username123](https://discuss.elastic.co/u/username123)\
**Post date:** [July 31, 2018, 8:54pm UTC](https://discuss.elastic.co/t/how-to-source-search-queries-from-a-file/142440/1 "2018-07-31T20:54:01Z")

</div>

Let's say I have a huge file with the queries that I have recorded over time:

```
"GET /primary/_analyze?analyzer=foldedlower&text=yomu
"GET /primary/_analyze?analyzer=foldedlower&text=yonde
"GET /primary/_analyze?analyzer=foldedlower&text=Yo+oigo+con+mis+orejas
"GET /primary/_analyze?analyzer=foldedlower&text=Yo+tengo+dos+ojos
"GET /primary/_analyze?analyzer=foldedlower&text=You%27re+eligible

```

I want to be able to point my `query` operation type to this file, so that the queries created by it would be sourced from the file. I can clean the data out and have everything after `&text=` utilized, but I am not certain how I would plug it into configuration. What would be a correct path to take?

---

<div class="post-metadata">

**Author:** ![danielmitterdorfer](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/danielmitterdorfer/32/110510_2.png) [@danielmitterdorfer](https://discuss.elastic.co/u/danielmitterdorfer)\
**Post date:** [August 1, 2018, 5:39am UTC](https://discuss.elastic.co/t/how-to-source-search-queries-from-a-file/142440/2 "2018-08-01T05:39:01Z")

</div>

Hi,

I think this is best solved by implementing a so-called [parameter source](http://esrally.readthedocs.io/en/stable/adding_tracks.html#custom-parameter-sources). There are two options to implement them: either as a (Python) function or as a class. In your case I'd go for a class. In the constructor you can read the file and prepare a suitable data structure (e.g. store the queries in a list) and in the `params()` method you just pick a query randomly or iterate through them - whatever suits your use case. In our official tracks, we to something very similar in the geonames track which you can use as a [starting point for your own parameter source](https://github.com/elastic/rally-tracks/blob/bc4f01ebe4837e508d3a255b039ad107e3c29f01/geonames/track.py#L5-L43).

Note that the `params()` method of the parameter source is called on a performance-critical path so you should do any preprocessing of the file already in the constructor, otherwise you might introduce an accidental bottleneck in the load generator. You can use Rally's [profiling support](http://esrally.readthedocs.io/en/stable/command_line_reference.html#clr-enable-driver-profiling) to profile it and double-check that there are no problems.

Daniel

---

<div class="post-metadata">

**Author:** ![username123](https://avatars.discourse-cdn.com/v4/letter/u/7ba0ec/32.png) [@username123](https://discuss.elastic.co/u/username123)\
**Post date:** [August 2, 2018, 7:28pm UTC](https://discuss.elastic.co/t/how-to-source-search-queries-from-a-file/142440/3 "2018-08-02T19:28:34Z")

</div>

I see!

Thank you Daniel! That was pretty much the direction I was thinking of going.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 30, 2018, 7:28pm UTC](https://discuss.elastic.co/t/how-to-source-search-queries-from-a-file/142440/4 "2018-08-30T19:28:35Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
