# Indexing RDF datasets

**URL:** <https://discuss.elastic.co/t/indexing-rdf-datasets/84281>\
**Category:** Elasticsearch\
**Created:** [May 2, 2017, 2:07pm UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281 "2017-05-02T14:07:50Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![AZammit](https://avatars.discourse-cdn.com/v4/letter/a/8edcca/32.png) [@AZammit](https://discuss.elastic.co/u/AZammit)\
**Post date:** [May 2, 2017, 2:07pm UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/1 "2017-05-02T14:07:50Z")

</div>

Hi all, I am new to ElasticSearch and currently I am doing research and implementing a concept for a small project. I would like to index the DBpedia RDF datasets using ElasticSearch. The RDF datasets will be stored in Apache Fuseki and I would like to stream these datasets into ElasticSearch for indexing. I found the following possibilities:

1. [https://github.com/elastic/elasticsearch-river-wikipedia](https://github.com/elastic/elasticsearch-river-wikipedia)  
Rivers Deprecated

2. [https://github.com/eea/eea.elasticsearch.river.rdf](https://github.com/eea/eea.elasticsearch.river.rdf)  
Rivers Deprecated.

3. [https://github.com/elastic/stream2es](https://github.com/elastic/stream2es)  
Suggests to use Logstash, although there already seems to be functionality to stream Wikipedia datasets into Elasticsearch.

4. Logstash  
Regarding Logstash, I am a bit lost since from my understanding Logstash gives you the facility to stream logs into Elasticsearch.

On which option I should concentrate my efforts? Are there any alternatives? It seems that there is no ready made solution to index RDF datasets.

---

<div class="post-metadata">

**Author:** ![nik9000](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nik9000/32/44947_2.png) [@nik9000](https://discuss.elastic.co/u/nik9000)\
**Post date:** [May 2, 2017, 3:41pm UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/2 "2017-05-02T15:41:19Z")

</div>

Logstash probably.

I dunno about Elaticsearch for RDF in general because it can't arbitrarily join and RDF is all joins. You can use Elasticsearch for the full text querying though.

---

<div class="post-metadata">

**Author:** ![Nilabhsagar](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/nilabhsagar/32/10647_2.png) [@Nilabhsagar](https://discuss.elastic.co/u/Nilabhsagar)\
**Post date:** [May 2, 2017, 5:07pm UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/3 "2017-05-02T17:07:07Z")

</div>

Elasticsearch might be a wrong choice here. I will suggest look into Marklogic. It should solve your requirement.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 2, 2017, 5:23pm UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/4 "2017-05-02T17:23:02Z")

</div>

I indexed the DBpedia link structure in elasticsearch and explored it using the Graph UI which can be used to give priority to significant links in the data (significant != popular). There's a video demo here [1] and if it looks like it is of interest I can share how this demo was put together with you.

[1] See 32 minutes in to [https://www.elastic.co/elasticon/conf/2016/sf/graph-capabilities-in-the-elastic-stack](https://www.elastic.co/elasticon/conf/2016/sf/graph-capabilities-in-the-elastic-stack)

---

<div class="post-metadata">

**Author:** ![AZammit](https://avatars.discourse-cdn.com/v4/letter/a/8edcca/32.png) [@AZammit](https://discuss.elastic.co/u/AZammit)\
**Post date:** [May 2, 2017, 5:55pm UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/5 "2017-05-02T17:55:30Z")

</div>

That is great Mark! This is exactly what I need.  
Yes please, Mark I want to know how the demo was put together.

The Graph UI is incredible; I tested it using the Shakespeare dataset and the experience was just awesome. For sure it will be awesome using the Graph UI on the DBpedia datasets.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 3, 2017, 9:08am UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/6 "2017-05-03T09:08:03Z")

</div>

Check this gist [1] for a python script to load dbpedia data [2] into 5.3+ elasticsearch.

Each elasticsearch doc is a single wikipedia article with an array of the other articles it links to.  
Using the Graph api/UI in x-pack [3] you can explore strongly-associated subjects (those subjects that are found to be commonly paired together in articles' `linked_subjects` field).

Cheers,  
Mark

[1] [https://gist.github.com/markharwood/21c723039425b4b3e4277b2bffa5c54c](https://gist.github.com/markharwood/21c723039425b4b3e4277b2bffa5c54c)  
[2] [http://downloads.dbpedia.org/3.6/en/page\_links\_en.nt.bz2](http://downloads.dbpedia.org/3.6/en/page_links_en.nt.bz2)  
[3] [https://www.elastic.co/downloads/x-pack](https://www.elastic.co/downloads/x-pack)

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [May 8, 2017, 11:47am UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/7 "2017-05-08T11:47:57Z")

</div>

Demo using 5.4 [https://youtu.be/ZzWT-2xdaek](https://youtu.be/ZzWT-2xdaek)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 5, 2017, 11:51am UTC](https://discuss.elastic.co/t/indexing-rdf-datasets/84281/8 "2017-06-05T11:51:25Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
