# Bulk indexing raise read timeout error

**URL:** <https://discuss.elastic.co/t/bulk-indexing-raise-read-timeout-error/798>\
**Category:** Elasticsearch\
**Created:** [May 17, 2015, 5:08am UTC](https://discuss.elastic.co/t/bulk-indexing-raise-read-timeout-error/798 "2015-05-17T05:08:23Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fool\_LeoTao](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fool_leotao/32/44882_2.png) [@Fool\_LeoTao](https://discuss.elastic.co/u/Fool_LeoTao)\
**Post date:** [May 17, 2015, 5:08am UTC](https://discuss.elastic.co/t/bulk-indexing-raise-read-timeout-error/798/1 "2015-05-17T05:08:23Z")

</div>

When using bulk api to index with python client,it's ok at begin.But sooner an readtime error raised like the following:

```auto
bulk_index start processing...
1 chunk bulk index spend: 20.0
2 chunk bulk index spend: 17.0
3 chunk bulk index spend: 17.0
4 chunk bulk index spend: 18.0
5 chunk bulk index spend: 18.0
6 chunk bulk index spend: 21.0
7 chunk bulk index spend: 19.0
8 chunk bulk index spend: 20.0
Traceback (most recent call last):
  File "es_index.py", line 54, in <module>
    bulk_index()
  File "es_index.py", line 19, in _
    rv = func(*args, **kwargs)
  File "es_index.py", line 48, in bulk_index
    chunk_size=100000, timeout=30)
  File "../es/wrappers.py", line 81, in bulk
    for chunk_len, errors in streaming_bulk_index(client, actions, **kwargs):
  File "../es/wrappers.py", line 58, in streaming_bulk_index
    raise e
elasticsearch.exceptions.ConnectionTimeout: ConnectionTimeout caused by - ReadTimeoutError(HTTPConnectionPool(host=u'219.224.135.97', port=9200): Read timed out. (read timeout=10))

```

I don't understand:

1. Read timeout seems like a problem concerning query,but when bulk indexing,why a read timeout error raised?
2. I use es-1.5.2 and just make elasticsearch.yml the following config which means the left config just use default.By the way, `ES_HEAP_SIZE` is set to 5g.

```auto
index.number_of_shards: 5
index.number_of_replicas: 0
index.store.type: mmapfs
indices.memory.index_buffer_size: 30%
index.translog.flush_threshold_ops: 50000
refresh_interval: 60s

```

My python code is simple like that:

```auto
es = Elasticsearch()

def bulk_index():
    actions = doc_generator()
    res = bulk(es, actions, index='test', doc_type='test',
               expand_action_callback=expand_action,
               chunk_size=100000, timeout=30)
    print 'res: ', res

```

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [May 17, 2015, 8:05am UTC](https://discuss.elastic.co/t/bulk-indexing-raise-read-timeout-error/798/2 "2015-05-17T08:05:49Z")

</div>

You get read timeouts from the server because the client is misbehaving. Cluster power, chunk size, timeout length and API use are not harmonized.

1. You do not let finish the indexing in 30 seconds, one reason is, the chunk is too large

2. You do not evaluate the bulk responses before continuing

Use a smaller `chunk_size` like 1000 und most important for convenient API usage, use [https://elasticsearch-py.readthedocs.org/en/master/helpers.html#elasticsearch.helpers.bulk](https://elasticsearch-py.readthedocs.org/en/master/helpers.html#elasticsearch.helpers.bulk) for evaluating the number of successfully indexed documents before you continue.

---

<div class="post-metadata">

**Author:** ![spuder](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spuder/32/44895_2.png) [@spuder](https://discuss.elastic.co/u/spuder)\
**Post date:** [May 31, 2015, 4:39am UTC](https://discuss.elastic.co/t/bulk-indexing-raise-read-timeout-error/798/3 "2015-05-31T04:39:16Z")

</div>

Possibly related [https://github.com/logstash-plugins/logstash-output-elasticsearch/issues/141#issuecomment-107113994](https://github.com/logstash-plugins/logstash-output-elasticsearch/issues/141#issuecomment-107113994)

---

<div class="post-metadata">

**Author:** ![fool\_01](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fool_01/32/7298_2.png) [@fool\_01](https://discuss.elastic.co/u/fool_01)\
**Post date:** [January 21, 2016, 4:47am UTC](https://discuss.elastic.co/t/bulk-indexing-raise-read-timeout-error/798/4 "2016-01-21T04:47:49Z")

</div>

I faced the same issue and finally the issue got resolved by the use of request\_timeout parameter instead of timeout.   
  
So the call must be like this helpers.bulk(es,actions,chunk\_size=some\_value,request\_timeout=some\_value)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 11:22pm UTC](https://discuss.elastic.co/t/bulk-indexing-raise-read-timeout-error/798/5 "2017-07-05T23:22:46Z")

</div>


