# Pyes bulk loading no server available error

**URL:** <https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162>\
**Category:** Elasticsearch\
**Created:** [March 26, 2011, 9:58pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162 "2011-03-26T21:58:00Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![Wolf\_2](https://avatars.discourse-cdn.com/v4/letter/w/ac8455/32.png) [@Wolf\_2](https://discuss.elastic.co/u/Wolf_2)\
**Post date:** [March 26, 2011, 9:58pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/1 "2011-03-26T21:58:00Z")

</div>

Hi,

First of all I cannot find any documentation other than python ES help  
pages or reading source code to setup bulk loading with pyes. However  
I am now running a demonstration with a limited capacity before errors  
occur. My report is below.

I am using the pyes module bulk loading application with elasticsearch  
0.15.2 . I am attempting to load a 6G documents of twitter like  
content onto a 14 instance Xen virtual cluster. Each instance has 2  
processors, 4GB of data and 50GB of storage. I am using the ES  
facility constructor with all with the array of Xen instances, 10  
retries and bulk\_size runs of 400, 1000, 2000 or 10000.

My index has 28 shards with no replicas. I have a mapping setup with a  
text field being indexed.

On a local server I have increased my file descriptor limit to 65535  
and used 4 shards but I have not been able to eliminate the error.

I am not able to load more 2G of documents without a server timeout  
error.

During operation I am able to load more than 2K of documents per  
second but I need to sustain the input without errors.

What should I do?

Below is a listing of my code and the error message:

import json  
import datetime  
import time  
import thread  
import os  
from thrift import Thrift  
from thrift.transport import TTransport  
from thrift.transport import TSocket  
from thrift.protocol.TBinaryProtocol import TBinaryProtocolAccelerated  
from pyes import \*  
from gzip import GzipFile

class DocWriter:  
def **init** (self, hosts):  
self.key\_id = 0  
self.port = '9500'  
self.server\_properties = { "base": "192.168.0.", "port": "9500" }  
self.index = { "name": "twitter" }  
self.type = { "name": "tweet" }  
self.hosts = map(self.set\_host, hosts)  
self.property\_files = { "base": "/root/F7setup", "index":  
"index\_settings", "mapping": "tweet\_mapping" }  
self.data\_path = '/root/F7setup/Cluster/data'  
self.connect = self.connection()  
self.set\_files()  
try:  
self.set\_index()  
except Exception:  
print "Index: " + self.index['name'] + " already exists"  
else:  
print "Index: " + self.index['name'] + " is set"  
self.set\_mapping()  
def set\_host(self, host):  
return self.server\_properties['base'] + host + ":" +  
self.server\_properties['port']  
def connection(self):  
return ES(self.hosts, bulk\_size=10000, max\_retries=10)

def delete\_index(self):  
self.connect.delete\_index(self.index['name'])  
def set\_index(self):  
with open(self.property\_files['base'] + "/" +  
self.property\_files['index'], 'r') as f:  
index=json.loads(f.read())  
self.index=index=index['index']  
self.connect.create\_index(index['name'], index['properties'])  
def set\_mapping(self):  
with open(self.property\_files['base']  
+"/"+self.property\_files['mapping'], 'r') as f:  
self.type=json.loads(f.read())  
self.type=self.type['type']  
self.connect.put\_mapping(self.type['name'], self.type,  
self.index['name'])  
def set\_files(self):  
self.data\_files = self.get\_files()  
def get\_files(self):  
return map(lambda x: self.data\_path + "/" + x,  
os.listdir(self.data\_path))  
def write\_files\_to\_es(self, file\_count\_offset=0):  
start = datetime.datetime.now()  
self.set\_files()  
del self.data\_files[0:file\_count\_offset]  
for file in self.data\_files:  
self.write\_block\_to\_es(file)  
print "document count: " + str(self.key\_id)  
print file + " written to es"  
finish = datetime.datetime.now()  
print "time span: ", finish - start  
print " duration: ", finish - start  
def write\_file\_to\_es(self, file):  
with GzipFile(file, 'r') as f:  
line = f.readline()  
while 0 \< len(line):  
self.key\_id = self.key\_id + 1  
line = json.loads(line, object\_hook=self.fix\_obj)  
response = self.connect.index(line, self.index['name'],  
self.type['name'], self.key\_id)  
try: response['ok']  
except NameError:  
print "file name: ", file, " count: ", self.key\_id  
print response  
f.close  
return self.key\_id  
else:  
line = f.readline()  
f.closed  
return self.key\_id  
def write\_block\_to\_es(self, file):  
with GzipFile(file, 'r') as f:  
line = f.readline()  
while 0 \< len(line):  
self.key\_id = self.key\_id + 1  
line = json.loads(line, object\_hook=self.fix\_obj)  
response = self.connect.index(line, self.index['name'],  
self.type['name'], self.key\_id, bulk=True)  
line = f.readline()

f.closed  
return self.key\_id  
def write\_file\_line\_to\_es(self, file, line\_no):  
line\_cnt = 1  
with GzipFile(file, 'r') as f:  
line = f.readline()  
while 0 \< len(line):  
self.key\_id = self.key\_id + 1  
if line\_cnt == line\_no:  
print line  
line = json.loads(line, object\_hook=self.fix\_obj)  
response = self.connect.index(line, self.index['name'],  
self.type['name'], self.key\_id)  
break  
line\_cnt = line\_cnt + 1  
line = f.readline()  
f.closed  
return self.key\_id  
def fix\_obj(self, obj):  
if 'in\_reply\_to\_user\_id' in obj:  
obj['in\_reply\_to\_user\_id'] = str(obj['in\_reply\_to\_user\_id'])  
if 'retweeted\_status' in obj:  
if 'retweet\_count' in obj['retweeted\_status']:

# print(obj['retweeted\_status']['retweet\_count'])

```
obj['retweeted_status']['retweet_count'] =

```

str(obj['retweeted\_status']['retweet\_count'])  
obj['retweeted\_status']=obj['retweeted\_status']  
elif 'retweet\_count' in obj:  
obj['retweet\_count'] = str(obj['retweet\_count'])  
return obj  
def iterate(self, loop\_count, delay=0):  
time.sleep(delay)

while 0 \< loop\_count:  
loop\_count = loop\_count - 1  
self.write\_file\_to\_es()  
return self.index

error message:

Traceback (most recent call last):  
File "", line 1, in   
File "doc\_writer.py", line 58, in write\_files\_to\_es  
self.write\_block\_to\_es(file)  
File "doc\_writer.py", line 87, in write\_block\_to\_es  
response = self.connect.index(line, self.index['name'],  
self.type['name'], self.key\_id, bulk=True)  
File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
es.py", line 519, in index  
self.flush\_bulk()  
File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
es.py", line 543, in flush\_bulk  
self.force\_bulk()  
File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
es.py", line 551, in force\_bulk  
self.\_send\_request("POST", "/\_bulk", self.bulk\_data.getvalue())  
File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
es.py", line 197, in \_send\_request  
response = self.connection.execute(request)  
File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
connection.py", line 174, in \_client\_call  
raise NoServerAvailable  
pyes.exceptions.NoServerAvailable

---

<div class="post-metadata">

**Author:** ![Clinton\_Gormley](https://avatars.discourse-cdn.com/v4/letter/c/50afbb/32.png) [@Clinton\_Gormley](https://discuss.elastic.co/u/Clinton_Gormley)\
**Post date:** [March 26, 2011, 10:47pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/2 "2011-03-26T22:47:27Z")

</div>

Hi Wolf

I wonder if you're running into swap problems, or your garbage  
collections are slow because some of the JVM is being swapped out?

Try turning swap off, it may help.

Also, see this blog post about what's in master:

[http://www.elasticsearch.org/blog/2011/03/23/update-settings.html](http://www.elasticsearch.org/blog/2011/03/23/update-settings.html)

hth

clint

---

<div class="post-metadata">

**Author:** ![Wolf\_2](https://avatars.discourse-cdn.com/v4/letter/w/ac8455/32.png) [@Wolf\_2](https://discuss.elastic.co/u/Wolf_2)\
**Post date:** [March 26, 2011, 10:47pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/3 "2011-03-26T22:47:43Z")

</div>

Hi I am getting better results by extending the timeout limit to 10  
sec but there are probably other issues to address.

Regards,

Wolf

On Mar 26, 2:58 pm, Wolf [wolfk...@gmail.com](mailto:wolfk...@gmail.com) wrote:

> Hi,
> 
> First of all I cannot find any documentation other than python ES help  
> pages or reading source code to setup bulk loading with pyes. However  
> I am now running a demonstration with a limited capacity before errors  
> occur. My report is below.
> 
> I am using the pyes module bulk loading application with elasticsearch  
> 0.15.2 . I am attempting to load a 6G documents of twitter like  
> content onto a 14 instance Xen virtual cluster. Each instance has 2  
> processors, 4GB of data and 50GB of storage. I am using the ES  
> facility constructor with all with the array of Xen instances, 10  
> retries and bulk\_size runs of 400, 1000, 2000 or 10000.
> 
> My index has 28 shards with no replicas. I have a mapping setup with a  
> text field being indexed.
> 
> On a local server I have increased my file descriptor limit to 65535  
> and used 4 shards but I have not been able to eliminate the error.
> 
> I am not able to load more 2G of documents without a server timeout  
> error.
> 
> During operation I am able to load more than 2K of documents per  
> second but I need to sustain the input without errors.
> 
> What should I do?
> 
> Below is a listing of my code and the error message:
> 
> import json  
> import datetime  
> import time  
> import thread  
> import os  
> from thrift import Thrift  
> from thrift.transport import TTransport  
> from thrift.transport import TSocket  
> from thrift.protocol.TBinaryProtocol import TBinaryProtocolAccelerated  
> from pyes import \*  
> from gzip import GzipFile
> 
> class DocWriter:  
> def **init** (self, hosts):  
> self.key\_id = 0  
> self.port = '9500'  
> self.server\_properties = { "base": "192.168.0.", "port": "9500" }  
> self.index = { "name": "twitter" }  
> self.type = { "name": "tweet" }  
> self.hosts = map(self.set\_host, hosts)  
> self.property\_files = { "base": "/root/F7setup", "index":  
> "index\_settings", "mapping": "tweet\_mapping" }  
> self.data\_path = '/root/F7setup/Cluster/data'  
> self.connect = self.connection()  
> self.set\_files()  
> try:  
> self.set\_index()  
> except Exception:  
> print "Index: " + self.index['name'] + " already exists"  
> else:  
> print "Index: " + self.index['name'] + " is set"  
> self.set\_mapping()  
> def set\_host(self, host):  
> return self.server\_properties['base'] + host + ":" +  
> self.server\_properties['port']  
> def connection(self):  
> return ES(self.hosts, bulk\_size=10000, max\_retries=10)
> 
> def delete\_index(self):  
> self.connect.delete\_index(self.index['name'])  
> def set\_index(self):  
> with open(self.property\_files['base'] + "/" +  
> self.property\_files['index'], 'r') as f:  
> index=json.loads(f.read())  
> self.index=index=index['index']  
> self.connect.create\_index(index['name'], index['properties'])  
> def set\_mapping(self):  
> with open(self.property\_files['base']  
> +"/"+self.property\_files['mapping'], 'r') as f:  
> self.type=json.loads(f.read())  
> self.type=self.type['type']  
> self.connect.put\_mapping(self.type['name'], self.type,  
> self.index['name'])  
> def set\_files(self):  
> self.data\_files = self.get\_files()  
> def get\_files(self):  
> return map(lambda x: self.data\_path + "/" + x,  
> os.listdir(self.data\_path))  
> def write\_files\_to\_es(self, file\_count\_offset=0):  
> start = datetime.datetime.now()  
> self.set\_files()  
> del self.data\_files[0:file\_count\_offset]  
> for file in self.data\_files:  
> self.write\_block\_to\_es(file)  
> print "document count: " + str(self.key\_id)  
> print file + " written to es"  
> finish = datetime.datetime.now()  
> print "time span: ", finish - start  
> print " duration: ", finish - start  
> def write\_file\_to\_es(self, file):  
> with GzipFile(file, 'r') as f:  
> line = f.readline()  
> while 0 \< len(line):  
> self.key\_id = self.key\_id + 1  
> line = json.loads(line, object\_hook=self.fix\_obj)  
> response = self.connect.index(line, self.index['name'],  
> self.type['name'], self.key\_id)  
> try: response['ok']  
> except NameError:  
> print "file name: ", file, " count: ", self.key\_id  
> print response  
> f.close  
> return self.key\_id  
> else:  
> line = f.readline()  
> f.closed  
> return self.key\_id  
> def write\_block\_to\_es(self, file):  
> with GzipFile(file, 'r') as f:  
> line = f.readline()  
> while 0 \< len(line):  
> self.key\_id = self.key\_id + 1  
> line = json.loads(line, object\_hook=self.fix\_obj)  
> response = self.connect.index(line, self.index['name'],  
> self.type['name'], self.key\_id, bulk=True)  
> line = f.readline()
> 
> f.closed  
> return self.key\_id  
> def write\_file\_line\_to\_es(self, file, line\_no):  
> line\_cnt = 1  
> with GzipFile(file, 'r') as f:  
> line = f.readline()  
> while 0 \< len(line):  
> self.key\_id = self.key\_id + 1  
> if line\_cnt == line\_no:  
> print line  
> line = json.loads(line, object\_hook=self.fix\_obj)  
> response = self.connect.index(line, self.index['name'],  
> self.type['name'], self.key\_id)  
> break  
> line\_cnt = line\_cnt + 1  
> line = f.readline()  
> f.closed  
> return self.key\_id  
> def fix\_obj(self, obj):  
> if 'in\_reply\_to\_user\_id' in obj:  
> obj['in\_reply\_to\_user\_id'] = str(obj['in\_reply\_to\_user\_id'])  
> if 'retweeted\_status' in obj:  
> if 'retweet\_count' in obj['retweeted\_status']:
> 
> # print(obj['retweeted\_status']['retweet\_count'])
> 
> ```
> obj['retweeted_status']['retweet_count'] =
> 
> ```
> 
> str(obj['retweeted\_status']['retweet\_count'])  
> obj['retweeted\_status']=obj['retweeted\_status']  
> elif 'retweet\_count' in obj:  
> obj['retweet\_count'] = str(obj['retweet\_count'])  
> return obj  
> def iterate(self, loop\_count, delay=0):  
> time.sleep(delay)
> 
> while 0 \< loop\_count:  
> loop\_count = loop\_count - 1  
> self.write\_file\_to\_es()  
> return self.index
> 
> error message:
> 
> Traceback (most recent call last):  
> File "", line 1, in   
> File "doc\_writer.py", line 58, in write\_files\_to\_es  
> self.write\_block\_to\_es(file)  
> File "doc\_writer.py", line 87, in write\_block\_to\_es  
> response = self.connect.index(line, self.index['name'],  
> self.type['name'], self.key\_id, bulk=True)  
> File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
> es.py", line 519, in index  
> self.flush\_bulk()  
> File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
> es.py", line 543, in flush\_bulk  
> self.force\_bulk()  
> File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
> es.py", line 551, in force\_bulk  
> self.\_send\_request("POST", "/\_bulk", self.bulk\_data.getvalue())  
> File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
> es.py", line 197, in \_send\_request  
> response = self.connection.execute(request)  
> File "/usr/lib/python2.7/site-packages/pyes-0.14.1-py2.7.egg/pyes/  
> connection.py", line 174, in \_client\_call  
> raise NoServerAvailable  
> pyes.exceptions.NoServerAvailable

---

<div class="post-metadata">

**Author:** ![Wolf\_2](https://avatars.discourse-cdn.com/v4/letter/w/ac8455/32.png) [@Wolf\_2](https://discuss.elastic.co/u/Wolf_2)\
**Post date:** [March 27, 2011, 11:56pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/4 "2011-03-27T23:56:57Z")

</div>

Thank you for the pointers Clint.

I set swapoff -a on all servers, cut down jvm memory allocation to 3/4  
of 4G physical capacity, set index.refresh\_interval: -1 and  
index.merge\_factor=30

I also added an another server to total 15 and extended the shards to  
30

I am still getting some rolling performance congestion every 40-60 or  
so block transfers of 10K documents, occasionally crashes still occur  
after 5.5M documents are loaded with a NoServerAvailable exception.

Hopefully I can get a more sustained load input and replication as my  
next objectives.

after setting the # of shard replicas to 1, refresh interval to 1s and  
merge factor to 10 the cluster status went to yellow!

what do I do now?

Regards,

Wolf

On Mar 26, 3:47 pm, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) wrote:

> Hi Wolf
> 
> I wonder if you're running into swap problems, or your garbage  
> collections are slow because some of the JVM is being swapped out?
> 
> Try turning swap off, it may help.
> 
> Also, see this blog post about what's in master:
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/blog/2011/03/23/update-settings.html)
> 
> hth
> 
> clint

---

<div class="post-metadata">

**Author:** ![Wolf\_2](https://avatars.discourse-cdn.com/v4/letter/w/ac8455/32.png) [@Wolf\_2](https://discuss.elastic.co/u/Wolf_2)\
**Post date:** [March 27, 2011, 11:57pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/5 "2011-03-27T23:57:46Z")

</div>

Thank you for the pointers Clint.

I set swapoff -a on all servers, cut down jvm memory allocation to 3/4  
of 4G physical capacity, set index.refresh\_interval: -1 and  
index.merge\_factor=30

I also added an another server to total 15 and extended the shards to  
30

I am still getting some rolling performance congestion every 40-60 or  
so block transfers of 10K documents, occasionally crashes still occur  
after 5.5M documents are loaded with a NoServerAvailable exception.

Hopefully I can get a more sustained load input and replication as my  
next objectives.

after setting the # of shard replicas to 1, refresh interval to 1s and  
merge factor to 10 the cluster status went to yellow!

what do I do now?

Regards,

Wolf

On Mar 26, 3:47 pm, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) wrote:

> Hi Wolf
> 
> I wonder if you're running into swap problems, or your garbage  
> collections are slow because some of the JVM is being swapped out?
> 
> Try turning swap off, it may help.
> 
> Also, see this blog post about what's in master:
> 
> [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/blog/2011/03/23/update-settings.html)
> 
> hth
> 
> clint

---

<div class="post-metadata">

**Author:** ![Wolf\_2](https://avatars.discourse-cdn.com/v4/letter/w/ac8455/32.png) [@Wolf\_2](https://discuss.elastic.co/u/Wolf_2)\
**Post date:** [March 28, 2011, 12:33am UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/6 "2011-03-28T00:33:54Z")

</div>

I have progressively been improving my performance with the python API  
although better block performance still has inconsistencies and server  
not found exception blowups.

Should I use the Java API for bulk loading? Possibly this would give  
me better control.

Peter Karich has a intro article on the API to get me started:

[http://java.dzone.com/articles/get-started-elasticsearch](http://java.dzone.com/articles/get-started-elasticsearch)

Regards,

Wolf

On Mar 27, 4:57 pm, Wolf [wolfk...@gmail.com](mailto:wolfk...@gmail.com) wrote:

> Thank you for the pointers Clint.
> 
> I set swapoff -a on all servers, cut down jvm memory allocation to 3/4  
> of 4G physical capacity, set index.refresh\_interval: -1 and  
> index.merge\_factor=30
> 
> I also added an another server to total 15 and extended the shards to  
> 30
> 
> I am still getting some rolling performance congestion every 40-60 or  
> so block transfers of 10K documents, occasionally crashes still occur  
> after 5.5M documents are loaded with a NoServerAvailable exception.
> 
> Hopefully I can get a more sustained load input and replication as my  
> next objectives.
> 
> after setting the # of shard replicas to 1, refresh interval to 1s and  
> merge factor to 10 the cluster status went to yellow!
> 
> what do I do now?
> 
> Regards,
> 
> Wolf
> 
> On Mar 26, 3:47 pm, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) wrote:
> 
> > Hi Wolf
> 
> > I wonder if you're running into swap problems, or your garbage  
> > collections are slow because some of the JVM is being swapped out?
> 
> > Try turning swap off, it may help.
> 
> > Also, see this blog post about what's in master:
> 
> > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/blog/2011/03/23/update-settings.html)
> 
> > hth
> 
> > clint

---

<div class="post-metadata">

**Author:** ![kimchy](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/kimchy/32/44952_2.png) [@kimchy](https://discuss.elastic.co/u/kimchy)\
**Post date:** [March 28, 2011, 10:18am UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/7 "2011-03-28T10:18:15Z")

</div>

Hey,

I am not familiar with the pyes API to great details, when does the server not found exception is being raised? Maybe you can _gist_ the code that uses pyes,

-shay.banon  
On Monday, March 28, 2011 at 2:33 AM, Wolf wrote:

> I have progressively been improving my performance with the python API  
> although better block performance still has inconsistencies and server  
> not found exception blowups.
> 
> Should I use the Java API for bulk loading? Possibly this would give  
> me better control.
> 
> Peter Karich has a intro article on the API to get me started:
> 
> [http://java.dzone.com/articles/get-started-elasticsearch](http://java.dzone.com/articles/get-started-elasticsearch)
> 
> Regards,
> 
> Wolf
> 
> On Mar 27, 4:57 pm, Wolf [wolfk...@gmail.com](mailto:wolfk...@gmail.com) wrote:
> 
> > Thank you for the pointers Clint.
> > 
> > I set swapoff -a on all servers, cut down jvm memory allocation to 3/4  
> > of 4G physical capacity, set index.refresh\_interval: -1 and  
> > index.merge\_factor=30
> > 
> > I also added an another server to total 15 and extended the shards to  
> > 30
> > 
> > I am still getting some rolling performance congestion every 40-60 or  
> > so block transfers of 10K documents, occasionally crashes still occur  
> > after 5.5M documents are loaded with a NoServerAvailable exception.
> > 
> > Hopefully I can get a more sustained load input and replication as my  
> > next objectives.
> > 
> > after setting the # of shard replicas to 1, refresh interval to 1s and  
> > merge factor to 10 the cluster status went to yellow!
> > 
> > what do I do now?
> > 
> > Regards,
> > 
> > Wolf
> > 
> > On Mar 26, 3:47 pm, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) wrote:
> > 
> > > Hi Wolf
> > 
> > > I wonder if you're running into swap problems, or your garbage  
> > > collections are slow because some of the JVM is being swapped out?
> > 
> > > Try turning swap off, it may help.
> > 
> > > Also, see this blog post about what's in master:
> > 
> > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/blog/2011/03/23/update-settings.html)
> > 
> > > hth
> > 
> > > clint

---

<div class="post-metadata">

**Author:** ![Wolf\_2](https://avatars.discourse-cdn.com/v4/letter/w/ac8455/32.png) [@Wolf\_2](https://discuss.elastic.co/u/Wolf_2)\
**Post date:** [March 28, 2011, 6:36pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/8 "2011-03-28T18:36:13Z")

</div>

Hi Shay,

Here is the gist link:

```
git://gist.github.com/890954.git

```

I run the program as follows:

hosts = [\<list of last digit of ipv4 address, 192.168.0. is assumed by  
program\>]

from doc\_writer import DocWriter

dw = DocWriter(hosts) # if no argument given the local host is the  
default

dw.write\_files\_to\_es()

* * *

I was able to run this program for 6 hours last night and load 50M  
tweets with only the text field indexed.

I have not attempted to set up a single replica yet since last time I  
got a yellow state.

The JVM swap on seems to be the primary cause of the Server Not Found  
Exception, Clint gave me the concept of this.

I am loading 10K tweets per bulk insert typically in 2.4 to 4.6  
seconds but I still do notice a rolling congestion in the network  
wherein a delay of up to 1 minute between bulk inserts occurs.

I may move to Java API to control the loading and application of the  
system more handily.

I need to load 20M tweets per day and support up to real time queries.  
I can extend my VM instances if necessary.

Wolf

* * *

here is the index setting I was using

{ "index":  
{ "name": "twitter",  
"properties":  
{ "index":  
{ "numberOfShards": 30, "numberOfReplicas": 0,  
"refresh\_interval" : "10s",  
"merge.policy.merge\_factor" : 30,  
"analysis":  
{ "analyzer":  
{ "collation":  
{ "tokenizer": "keyword",  
"filter": ["myCollator"]  
},  
"my\_analyzer":  
{ "type": "igo"  
}  
},  
"filter":  
{ "myCollator":  
{ "type": "icu\_collation",  
"language": "ja"  
}  
}  
}  
}  
}  
}  
}

On Mar 28, 3:18 am, Shay Banon [shay.ba...@elasticsearch.com](mailto:shay.ba...@elasticsearch.com) wrote:

> Hey,
> 
> I am not familiar with the pyes API to great details, when does the server not found exception is being raised? Maybe you can _gist_ the code that uses pyes,
> 
> -shay.banon
> 
> On Monday, March 28, 2011 at 2:33 AM, Wolf wrote:
> 
> > I have progressively been improving my performance with the python API  
> > although better block performance still has inconsistencies and server  
> > not found exception blowups.
> 
> > Should I use the Java API for bulk loading? Possibly this would give  
> > me better control.
> 
> > Peter Karich has a intro article on the API to get me started:
> 
> > [http://java.dzone.com/articles/get-started-elasticsearch](http://java.dzone.com/articles/get-started-elasticsearch)
> 
> > Regards,
> 
> > Wolf
> 
> > On Mar 27, 4:57 pm, Wolf [wolfk...@gmail.com](mailto:wolfk...@gmail.com) wrote:
> > 
> > > Thank you for the pointers Clint.
> 
> > > I set swapoff -a on all servers, cut down jvm memory allocation to 3/4  
> > > of 4G physical capacity, set index.refresh\_interval: -1 and  
> > > index.merge\_factor=30
> 
> > > I also added an another server to total 15 and extended the shards to  
> > > 30
> 
> > > I am still getting some rolling performance congestion every 40-60 or  
> > > so block transfers of 10K documents, occasionally crashes still occur  
> > > after 5.5M documents are loaded with a NoServerAvailable exception.
> 
> > > Hopefully I can get a more sustained load input and replication as my  
> > > next objectives.
> 
> > > after setting the # of shard replicas to 1, refresh interval to 1s and  
> > > merge factor to 10 the cluster status went to yellow!
> 
> > > what do I do now?
> 
> > > Regards,
> 
> > > Wolf
> 
> > > On Mar 26, 3:47 pm, Clinton Gormley [clin...@iannounce.co.uk](mailto:clin...@iannounce.co.uk) wrote:
> 
> > > > Hi Wolf
> 
> > > > I wonder if you're running into swap problems, or your garbage  
> > > > collections are slow because some of the JVM is being swapped out?
> 
> > > > Try turning swap off, it may help.
> 
> > > > Also, see this blog post about what's in master:
> 
> > > > [Elasticsearch Platform — Find real-time answers at scale | Elastic](http://www.elasticsearch.org/blog/2011/03/23/update-settings.html)
> 
> > > > hth
> 
> > > > clint

---

<div class="post-metadata">

**Author:** ![Alberto\_Paro\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alberto_paro_2/32/1137_2.png) [@Alberto\_Paro\_2](https://discuss.elastic.co/u/Alberto_Paro_2)\
**Post date:** [March 30, 2011, 8:20pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/9 "2011-03-30T20:20:21Z")

</div>

Il giorno 28/mar/2011, alle ore 20.36, Wolf ha scritto:

To reduce problems with memory allocation, usually I set the ES\_MIN\_MEM=ES\_MAX\_MEM.

To improve the import speed, try to use an eventlet pool in which you define a connection for every server to parallelize the inserting.

(I'm writing the documentation for pyes, so be patient. But the tests dir cover a lot of cases)

Also I'll implement some helpers to set the index in "bulk mode" as [http://www.elasticsearch.org/blog/2011/03/23/update-settings.html](http://www.elasticsearch.org/blog/2011/03/23/update-settings.html) after checking that ES is a least a 0.16 version

Let me know, if you discovery some problems with pyes (possibly via github or IRC), so I can improve it.

Hi,  
Alberto Paro

---

<div class="post-metadata">

**Author:** ![Wolf\_2](https://avatars.discourse-cdn.com/v4/letter/w/ac8455/32.png) [@Wolf\_2](https://discuss.elastic.co/u/Wolf_2)\
**Post date:** [March 31, 2011, 8:20am UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/10 "2011-03-31T08:20:40Z")

</div>

Hi Alberto,

Thank you for getting back.

Presently I am generating and pyes ES instance with a list of node IP  
addresses. Does this generate the eventlet pool? If not could you be  
more specific. My code is on gist:

Here is the gist link:

```
    git://gist.github.com/890954.git

```

* * *

I run the program as follows from the python command prompt:

hosts = [\<list of last digit of ipv4 address, 192.168.0. is assumed by  
program\>]

from doc\_writer import DocWriter

dw = DocWriter(hosts) # if no argument is given the local host is the  
default

dw.write\_files\_to\_es()

* * *

I have set the ES\_MIN\_MEM = ES\_MAX\_MEM

Regards,

Wolf

On Mar 30, 1:20 pm, Alberto Paro [alberto.p...@gmail.com](mailto:alberto.p...@gmail.com) wrote:

> Il giorno 28/mar/2011, alle ore 20.36, Wolf ha scritto:
> 
> To reduce problems with memory allocation, usually I set the ES\_MIN\_MEM=ES\_MAX\_MEM.
> 
> To improve the import speed, try to use an eventlet pool in which you define a connection for every server to parallelize the inserting.
> 
> (I'm writing the documentation for pyes, so be patient. But the tests dir cover a lot of cases)
> 
> Also I'll implement some helpers to set the index in "bulk mode" ashttp://www.elasticsearch.org/blog/2011/03/23/update-settings.htmlafter checking that ES is a least a 0.16 version
> 
> Let me know, if you discovery some problems with pyes (possibly via github or IRC), so I can improve it.
> 
> Hi,  
> Alberto Paro

---

<div class="post-metadata">

**Author:** ![Alberto\_Paro\_2](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/alberto_paro_2/32/1137_2.png) [@Alberto\_Paro\_2](https://discuss.elastic.co/u/Alberto_Paro_2)\
**Post date:** [March 31, 2011, 8:32pm UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/11 "2011-03-31T20:32:16Z")

</div>

Il giorno 31/mar/2011, alle ore 10.20, Wolf ha scritto:

> Hi Alberto,
> 
> Thank you for getting back.
> 
> Presently I am generating and pyes ES instance with a list of node IP  
> addresses. Does this generate the eventlet pool? If not could you be  
> more specific. My code is on gist:

I'll give you some references [greenpool – Green Thread Pools — Eventlet 0.33.0 documentation](http://eventlet.net/doc/modules/greenpool.html) :

The flow is this one put your code:

def bulk\_data\_insert(server, filedata):  
#create es connection  
# load data  
# do bulk insert  
#results

def mygenerator(servers, filestoprocess):  
count = 0  
for filename in filestoprocess:  
#load data from filename  
yield servers[count%len(servers)], data  
count += 1

def main():  
#init green threadpool size==num server  
#give them to process to pool  
pool = GreenPool(len(servers))  
for result in pool.imap(bulk\_data\_insert, mygenerator(servers, filestoprocess)):  
print result

NOTE:  
Using eventlet you have not blocking calls. You can consider eventlet as "node.js python version".  
If you sent too much data and you don't have a lot of memory/CPU on ES server, you'll put it on high load.

Hi,  
Alberto

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 4:09am UTC](https://discuss.elastic.co/t/pyes-bulk-loading-no-server-available-error/4162/12 "2017-07-06T04:09:02Z")

</div>


