# Recommendation for large synonym file

**URL:** <https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740>\
**Category:** Elasticsearch\
**Created:** [February 25, 2016, 4:31pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740 "2016-02-25T16:31:30Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![timpb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timpb/32/8065_2.png) [@timpb](https://discuss.elastic.co/u/timpb)\
**Post date:** [February 25, 2016, 4:31pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/1 "2016-02-25T16:31:30Z")

</div>

Hi all,

I am looking into the synonym token filter and the Elasticsearch documentation recommends that when you work with large synonym datasets you should set the `synonyms_path` to a file over inserting synonyms directly into the configuration file.

But the documentation tells not why. Is it just because of maintainability? That you do not want scroll through thousands of synonyms to check your configuration/mapping?

And how does Elastcisearch handle this synonym file? Are the contents of the file loaded into memory after an configuration update?

Thanks in advance!

Regards,

Tim

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 25, 2016, 4:47pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/2 "2016-02-25T16:47:57Z")

</div>

Probably because the cluster state which contains index settings will become too big?

---

<div class="post-metadata">

**Author:** ![frankkoornstra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/frankkoornstra/32/6780_2.png) [@frankkoornstra](https://discuss.elastic.co/u/frankkoornstra)\
**Post date:** [February 26, 2016, 8:45am UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/3 "2016-02-26T08:45:08Z")

</div>

I'm keen to know an answer to this as well. Provisioning synonyms through index settings is way easier (done through REST API) than provisioning through files.

@dadoonet: are you sure that all synonyms get sent along with cluster state? It doesn't seem logical to me since synonyms are part of the index settings, not the cluster state. Sending along all filters with cluster state seems weird.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 26, 2016, 3:53pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/4 "2016-02-26T15:53:14Z")

</div>

Index settings are part of the index metadata and index metadata is part of the cluster state.

---

<div class="post-metadata">

**Author:** ![frankkoornstra](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/frankkoornstra/32/6780_2.png) [@frankkoornstra](https://discuss.elastic.co/u/frankkoornstra)\
**Post date:** [February 26, 2016, 4:03pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/5 "2016-02-26T16:03:15Z")

</div>

Alright, thanks for that! Putting them in a file seems better then 🙂

---

<div class="post-metadata">

**Author:** ![timpb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/timpb/32/8065_2.png) [@timpb](https://discuss.elastic.co/u/timpb)\
**Post date:** [February 26, 2016, 4:08pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/6 "2016-02-26T16:08:09Z")

</div>

Thanks for the clarification!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [February 26, 2016, 4:31pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/7 "2016-02-26T16:31:13Z")

</div>

I'm unsure if it's better. TBH I'd really love it to be loaded from a document stored into an elasticsearch index than from the file system. Because, it's harder to maintain on the FS and distribute on all nodes a consistent file.

I opened this feature request. Will see where it goes: [https://github.com/elastic/elasticsearch/issues/16824](https://github.com/elastic/elasticsearch/issues/16824)

---

<div class="post-metadata">

**Author:** ![Ivan](https://avatars.discourse-cdn.com/v4/letter/i/df788c/32.png) [@Ivan](https://discuss.elastic.co/u/Ivan)\
**Post date:** [February 26, 2016, 5:09pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/8 "2016-02-26T17:09:48Z")

</div>

Years ago I wrote a collection of token filters that read its values from a  
database. Been in production all this time. Always wanted to reboot that  
project into a public release, but it was held up due to a change in the  
way analyzers are created/stored in Elasticsearch. The change was pushed  
into the 3.x (now 5.x) branch. Since 5.0 will be released soon (alpha at  
least), I should revisit the project.

Ivan

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:13pm UTC](https://discuss.elastic.co/t/recommendation-for-large-synonym-file/42740/9 "2017-07-05T23:13:18Z")

</div>


