# Loading and indexing dbpedia datasets in ES

**URL:** <https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739>\
**Category:** Elasticsearch\
**Created:** [March 11, 2019, 11:30am UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739 "2019-03-11T11:30:26Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![Furabio](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/furabio/32/43115_2.png) [@Furabio](https://discuss.elastic.co/u/Furabio)\
**Post date:** [March 11, 2019, 11:30am UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739/1 "2019-03-11T11:30:26Z")

</div>

Hello,  
I'm new to Elasticsearch and I am trying to load and index dbpedia datasets (RDF triples) into ES. The datasets are available in ttl format from [https://wiki.dbpedia.org/downloads-2016-04](https://wiki.dbpedia.org/downloads-2016-04).

My question is how do I load and index this into ES? Should I first convert the datato json format?

Thanks for your help.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [March 11, 2019, 11:38am UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739/2 "2019-03-11T11:38:04Z")

</div>

See [https://www.youtube.com/watch?v=ZzWT-2xdaek](https://www.youtube.com/watch?v=ZzWT-2xdaek)  
The comments section includes a link to some code

---

<div class="post-metadata">

**Author:** ![Furabio](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/furabio/32/43115_2.png) [@Furabio](https://discuss.elastic.co/u/Furabio)\
**Post date:** [March 11, 2019, 12:20pm UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739/3 "2019-03-11T12:20:46Z")

</div>

Hi, thanks for the reply.  
I've already seen that tutorial but I miss the passage of the loading of the dataset in ES. When I try to run the python script, I get a connection error, although elasticsearch is running in the cloud. When elasticsearch is running locally instead, the python script is executed successfully, but on cmd I have a java.io.IOException and I see this error "[o.e.h.n.Netty4HttpServerTransport] [my\_node] caught exception while handling client http traffic, closing connection"

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [March 11, 2019, 1:34pm UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739/4 "2019-03-11T13:34:16Z")

</div>

> [@Furabio](#):
>
> I get a connection error, although elasticsearch is running in the cloud

You need to setup the connection details correctly.  
This is an example of a python client connecting to an [elastic cloud](https://www.elastic.co/cloud/) cluster:

```
import certifi
from elasticsearch.client import Elasticsearch

remoteEs = Elasticsearch(
		["xxxxxMY_CLOUD_ENDPOINT xxxxxx.found.io"],
		port=9243,
		http_auth="MY_USERNAME:MY_PASSWORD",
		use_ssl=True,
		verify_certs=True,
		ca_certs=certifi.where()
	)

response = remoteEs.search(index="MY_INDEX", body = myQuery)

```

---

<div class="post-metadata">

**Author:** ![Furabio](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/furabio/32/43115_2.png) [@Furabio](https://discuss.elastic.co/u/Furabio)\
**Post date:** [March 11, 2019, 7:03pm UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739/5 "2019-03-11T19:03:23Z")

</div>

Thanks again.  
I solved it: the python script is executed successfully, and I have verified that the index is created into ES successfully.  
However, I have one last problem: the index is empty.  
I noticed that this depends on the fact that the script never enters the final loop, which should iterate over all the triples in the dataset (`for line in file:`, where `file` is obtained by `with os.popen('bzip2 -cd ' + filename) as file:`).

Also, I noticed that the import of bz2 is unused. Can the above error depend on this?

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [March 11, 2019, 7:17pm UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739/6 "2019-03-11T19:17:12Z")

</div>

To be honest the code was probably originally written by searching stackoverflow for “how to read a bz2 file using python”. This probably isn’t the forum to discuss your problems with reading the raw data but I’d start by checking the file name is right.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 8, 2019, 7:17pm UTC](https://discuss.elastic.co/t/loading-and-indexing-dbpedia-datasets-in-es/171739/7 "2019-04-08T19:17:16Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
