# ElasticSearch as primary DB for document library

**URL:** <https://discuss.elastic.co/t/elasticsearch-as-primary-db-for-document-library/307852>\
**Category:** Elasticsearch\
**Created:** [June 22, 2022, 8:54am UTC](https://discuss.elastic.co/t/elasticsearch-as-primary-db-for-document-library/307852 "2022-06-22T08:54:01Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Gresmir](https://avatars.discourse-cdn.com/v4/letter/g/e47c2d/32.png) [@Gresmir](https://discuss.elastic.co/u/Gresmir)\
**Post date:** [June 22, 2022, 8:54am UTC](https://discuss.elastic.co/t/elasticsearch-as-primary-db-for-document-library/307852/1 "2022-06-22T08:54:01Z")

</div>

My task is a full-text search system for a really large amount of documents (tens of millions). Now I have documents as RTF file and their metadata, so all this will be indexed in Elasticsearch. These documents are unchangeable (they can be only deleted). I don't really expect many new documents per day and I choose the time these documents are inserted. So is it a good idea to use elastic as primary DB in this case?

Maybe I'll store the RTF file separately, but I really don't see the point of storing all this data somewhere else.

---

<div class="post-metadata">

**Author:** ![whatgeorgemade](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/whatgeorgemade/32/103246_2.png) [@whatgeorgemade](https://discuss.elastic.co/u/whatgeorgemade)\
**Post date:** [June 22, 2022, 9:30am UTC](https://discuss.elastic.co/t/elasticsearch-as-primary-db-for-document-library/307852/2 "2022-06-22T09:30:13Z")

</div>

Elasticsearch should be fine for this use-case. Just be sure to keep snapshots. Keeping the original documents is a good idea, so you're able to completely rebuild the indices if the worst happens, or you hit bugs in ingest and need to start from scratch.

Depending on how you define 'really large amount', Elasticsearch could be overkill. If you don't need to distribute the data over multiple nodes and have some Java expertise, using vanilla Lucene is worth thinking about.

---

<div class="post-metadata">

**Author:** ![Gresmir](https://avatars.discourse-cdn.com/v4/letter/g/e47c2d/32.png) [@Gresmir](https://discuss.elastic.co/u/Gresmir)\
**Post date:** [June 22, 2022, 12:05pm UTC](https://discuss.elastic.co/t/elasticsearch-as-primary-db-for-document-library/307852/3 "2022-06-22T12:05:47Z")

</div>

I have a couple of million documents, according to approximate calculations in RTF files, it will be near a petabyte.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 20, 2022, 12:05pm UTC](https://discuss.elastic.co/t/elasticsearch-as-primary-db-for-document-library/307852/4 "2022-07-20T12:05:58Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
