# Search timeouts during indexing

**URL:** <https://discuss.elastic.co/t/search-timeouts-during-indexing/284153>\
**Category:** Elasticsearch\
**Created:** [September 14, 2021, 7:49am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153 "2021-09-14T07:49:58Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Drept](https://avatars.discourse-cdn.com/v4/letter/d/ce7236/32.png) [@Drept](https://discuss.elastic.co/u/Drept)\
**Post date:** [September 14, 2021, 7:49am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/1 "2021-09-14T07:49:58Z")

</div>

I have an index of around 90 million documents in Elasticsearch. I manage it via django-elasticsearch-dsl. The settings are the following:

- number\_of\_shards: 1
- number\_of\_replicas: 0

The mapping is:

```auto
{
  "series" : {
    "mappings" : {
      "properties" : {
        "description" : {
          "type" : "text"
        },
        "all_actors" : {
          "type" : "text"
        },
        "episode_title" : {
          "type" : "text"
        },
        "actors_keyword" : {
          "type" : "keyword",
          "ignore_above" : 1000
        },
        "series_title" : {
          "type" : "keyword",
          "ignore_above" : 1000
        },
        "language" : {
          "type" : "keyword",
          "ignore_above" : 1000
        },
        "number_of_actors" : {
          "type" : "short"
        },
        "translated_title" : {
          "type" : "text"
        },
        "tags" : {
          "type" : "text"
        },
        "tags_keyword" : {
          "type" : "keyword",
          "ignore_above" : 1000
        },
        "url" : {
          "type" : "text"
        },
        "year" : {
          "type" : "short",
          "null_value" : 0
        }
      }
    }
  }
}

```

Now, the problem is: when I start adding new objects to PostgreSQL and signals are being sent to Elasticsearch, the search becomes extremely slow and starts throwing timeouts. The rate of adding new objects is moderate: around 300 objects per minute.

The same occurs when I delete even 100 objects from PostgreSQL. I am confident that this is due to the slow processing of signals. After some 30 seconds (I believe when the processings of the signals is over), everything works fine again, but if I have to process some 2 million objects over 2 weeks, Elasticsearch will be unresponsive for 2 weeks too (at least).

What I have tried to remedy the issue:

- setting refresh\_interval to "60s";
- increasing the JVM heap size twice (to 4 GB out of the total 16 GB available);
- profiling the data from the slow logs: the results show that the queries are really quick, so the problem is not in the queries;
- checking the shard health: no issues identified;
- doing the same operations with other – smaller – indices. This shows that there are no speed problems with the smaller indices (or they are hardly noticeable).

Could you help me understand what might be causing the timeouts in this case?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 14, 2021, 7:52am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/2 "2021-09-14T07:52:12Z")

</div>

What is the output from the `_cluster/stats?pretty&human` API?

---

<div class="post-metadata">

**Author:** ![Drept](https://avatars.discourse-cdn.com/v4/letter/d/ce7236/32.png) [@Drept](https://discuss.elastic.co/u/Drept)\
**Post date:** [September 14, 2021, 8:10am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/4 "2021-09-14T08:10:32Z")

</div>

Here it is in the gist: [https://gist.github.com/Dterb/6812ffee42ebd8e714eeefd5a7d70033](https://gist.github.com/Dterb/6812ffee42ebd8e714eeefd5a7d70033)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [September 14, 2021, 8:13am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/5 "2021-09-14T08:13:28Z")

</div>

Thanks.

It's surprising that this is happening, what sort of storage does the node have?  
Is there anything in your Elasticsearch logs?  
What about your Elasticsearch slow logs and hot threads?  
Are you using the Monitoring functionality?

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 14, 2021, 8:18am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/6 "2021-09-14T08:18:47Z")

</div>

What kind of hardware is the node deployed on? What type of storage are you using? Local SSD?

---

<div class="post-metadata">

**Author:** ![Drept](https://avatars.discourse-cdn.com/v4/letter/d/ce7236/32.png) [@Drept](https://discuss.elastic.co/u/Drept)\
**Post date:** [September 14, 2021, 8:23am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/7 "2021-09-14T08:23:55Z")

</div>

The last records in /etc/log/elasticsearch/elasticsearch.log are one-hour-old: just showing that I restarted Elasticsearch. The next records are 12-hour-old, so no, there is nothing strange in there.

Slow logs: multiple queries consuming 40+ seconds. When profiling them, the results show that the queries actually take from 1 to 4 seconds (as expected when no signals are being sent to the database).

Hot threads. Here are two prints in the gists: [one](https://gist.github.com/Dterb/34d94dae30859282d62281d613999620), [two](https://gist.github.com/Dterb/f3b816ba3120c05a1106fe5cbf583034).

Monitoring functionality. No, I have not been using it.

---

<div class="post-metadata">

**Author:** ![Drept](https://avatars.discourse-cdn.com/v4/letter/d/ce7236/32.png) [@Drept](https://discuss.elastic.co/u/Drept)\
**Post date:** [September 14, 2021, 8:27am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/8 "2021-09-14T08:27:26Z")

</div>

I am using DigitalOcean: Shared CPU, 8 vCPUs, 16 GB RAM, 160 GB SSD, 6 TB Transfer.

The indices are stored on the droplet's SSD. The tablespaces for the underlying database objects are stored on a separate block storage volume (their size is far above the 160 GB of the SSD).

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 14, 2021, 9:14am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/9 "2021-09-14T09:14:00Z")

</div>

So you have slow storage fronted by an SSD cache? In that case it is possible that indexing and merging may affect the cache, causing slower searches until it has again loaded the most queried data. I would recommend monitoring iowait during indexing and initial querying to see if this spikes. If it does you may need to switch to faster storage if you are to eliminate these performance issues.

---

<div class="post-metadata">

**Author:** ![Drept](https://avatars.discourse-cdn.com/v4/letter/d/ce7236/32.png) [@Drept](https://discuss.elastic.co/u/Drept)\
**Post date:** [September 14, 2021, 9:56am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/10 "2021-09-14T09:56:39Z")

</div>

I've tried to measure the I/O wait by means of the _top_ command and the _wa_ parameter it is showing.

In general, _wa_ is dynamically changing from 0.3 to 5.5. When I am deleting hundreds of objects and am searching in Elasticsearch, it is rising to 7.5 to 10.0. The greatest spike I saw was 12.7 % Cpu(s).

Does this seem to be the problem? Or these figures are rather normal?

UPD. Actually, even several minutes after, when the indexing should have been over, _wa_ still varies from 0.3 to 9.8.

UPD 2. The _read - vda_ parameter is surging during these periods:

 ![read-vda](https://us1.discourse-cdn.com/elastic/original/3X/2/8/28fa40e8edc941374c939f519e765694e9ad6bde.png)

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [September 14, 2021, 10:10am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/11 "2021-09-14T10:10:03Z")

</div>

Then it sounds like I/O performance (or rather lack of it) is the cause. Make sure you are using bulk requests when indexing/updating as this may help some. To solve the problem you probably need better and faster storage though. Indexing and subsequent merging is often quite I/O intensive.

---

<div class="post-metadata">

**Author:** ![Drept](https://avatars.discourse-cdn.com/v4/letter/d/ce7236/32.png) [@Drept](https://discuss.elastic.co/u/Drept)\
**Post date:** [September 15, 2021, 8:49am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/12 "2021-09-15T08:49:06Z")

</div>

Thanks, your answered really helped.

The problem was in the frequency of refreshes (of course, looking more deeply, in I/O) and in the fact that django\_elasticsearch\_dsl does not use the Bulk API by default.

Even '60s' was too much for a large index, and the refreshes took all the 60 seconds or even more, which made the search unresponsive.

I had to:

- disable all signals in django\_elasticsearch\_dsl (as thwy were triggering individual actions for every objects;
- set _refresh\_interval_ to '-1' (i.e. to disable automatic refreshes).

Now, I am using bulk updates by specifying the objects that need to be indexed/deleted manually and am refreshing the indices once per week. That resolved the issue.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 13, 2021, 8:49am UTC](https://discuss.elastic.co/t/search-timeouts-during-indexing/284153/13 "2021-10-13T08:49:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
