# Building Graph Relationships Between Documents

**URL:** <https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452>\
**Category:** Kibana\
**Tags:** elastic-stack-graph\
**Created:** [September 13, 2018, 12:54pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452 "2018-09-13T12:54:48Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrew\_Stroz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_stroz/32/104061_2.png) [@Andrew\_Stroz](https://discuss.elastic.co/u/Andrew_Stroz)\
**Post date:** [September 13, 2018, 12:54pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/1 "2018-09-13T12:54:48Z")

</div>

I have a index that contains a field with the text contents of .docx, .pptx and .pdf documents.

I have another index that holds documents that have a field with a string of characters representing a piece of equipment.

I would like to run a graph query that shows all of the related documents retrieved from full text search to all of the pieces of equipment.

Is this possible?

Thanks.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 13, 2018, 1:04pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/2 "2018-09-13T13:04:59Z")

</div>

> [@Andrew\_Stroz](#):
>
> I would like to run a graph query that shows all of the related documents retrieved from full text search to all of the pieces of equipment.

Is there a field the two indices share?

> [@Andrew\_Stroz](#):
>
> I have another index that holds documents that have a field with a string of characters representing a piece of equipment.

For the sake of argument let's call that field "equipment\_id" and assume it is of the type `keyword`

> [@Andrew\_Stroz](#):
>
> I have a index that contains a field with the text contents of .docx, .pptx and .pdf documents.

I'm guessing the field is of the type `text` and hidden in the text are some references to items of equipment. If the pattern of an equipment\_id is sufficiently unique (e.g. always an 11 digit number) then it might be possible to use a regex to extract these values from the text and place into `keyword` type field called `equipment_id` which is an array. Let's also assume each document has a keyword field called `doc_id`.

Given this setup it would be possible to create a graph of `doc_id` and `equipment_id` values and how they are connected purely using the document index (ignoring the equipment index).

This is mostly speculation about your data so I think you may need to fill in some more details about the problem here.

---

<div class="post-metadata">

**Author:** ![Andrew\_Stroz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_stroz/32/104061_2.png) [@Andrew\_Stroz](https://discuss.elastic.co/u/Andrew_Stroz)\
**Post date:** [September 13, 2018, 1:44pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/3 "2018-09-13T13:44:41Z")

</div>

> [@Mark\_Harwood](#):
>
> I'm guessing the field is of the type `text` and hidden in the text are some references to items of equipment.

Yes this is true.

> [@Mark\_Harwood](#):
>
> If the pattern of an equipment\_id is sufficiently unique (e.g. always an 11 digit number) then it might be possible to use a regex to extract these values from the text and place into `keyword` type field called `equipment_id` which is an array. Let's also assume each document has a keyword field called `doc_id` .

This does not hold true. The equipment\_id field is not sufficiently unique to perform regex to extract the values. That is why I was hoping I could relate `equipment_id` from one set of documents to the full text search results for that `equipment_id` in the indexed word/ppt/pdf.

Thanks.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 13, 2018, 2:01pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/4 "2018-09-13T14:01:58Z")

</div>

If your app can't isolate the numbers from the text at ingest time then elasticsearch will equally have a hard time doing any analysis on this data at query time.

---

<div class="post-metadata">

**Author:** ![Andrew\_Stroz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_stroz/32/104061_2.png) [@Andrew\_Stroz](https://discuss.elastic.co/u/Andrew_Stroz)\
**Post date:** [September 13, 2018, 2:06pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/5 "2018-09-13T14:06:55Z")

</div>

Is there not some way to visualize with graph the `equipment_id` as a central node and all of its edges are connected to full text search query results? Preferably with the strength of the node being related to the score returned from the full text search.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 13, 2018, 2:18pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/6 "2018-09-13T14:18:21Z")

</div>

> [@Andrew\_Stroz](#):
>
> Is there not some way to visualize with graph the `equipment_id` as a central node and all of its edges are connected to full text search query results?

Not a clean way, no. We rely on nodes being identified by a combination of fieldname and term which will make life complex if you can't extract equipment IDs out of the text into a field called equipmentID.

---

<div class="post-metadata">

**Author:** ![Andrew\_Stroz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_stroz/32/104061_2.png) [@Andrew\_Stroz](https://discuss.elastic.co/u/Andrew_Stroz)\
**Post date:** [September 13, 2018, 3:00pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/7 "2018-09-13T15:00:23Z")

</div>

Is there any way for a node to be identified by a document? Or a node to be identified as by a combination of fieldname and term but as the result of a search query.

I would really like to have the ability to relate a 'node being identified by a combination of fieldname and term' to documents that that match a query for that term.

This would allow me to harness the power of Elastic as a full text search service and the visualization that graph offers.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 13, 2018, 3:25pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/8 "2018-09-13T15:25:32Z")

</div>

> [@Andrew\_Stroz](#):
>
> This would allow me to harness the power of Elastic as a full text search service and the visualization that graph offers.

Is this a useful graph visualization? If I understand your requirement it is a star-shaped graph with a single central "query" node and lines connecting out to matching satellite "doc" nodes.

That sounds more usefully drawn as a horizontal bar chart with a bar per doc and bar lengths being doc score?

---

<div class="post-metadata">

**Author:** ![Andrew\_Stroz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrew_stroz/32/104061_2.png) [@Andrew\_Stroz](https://discuss.elastic.co/u/Andrew_Stroz)\
**Post date:** [September 13, 2018, 4:09pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/9 "2018-09-13T16:09:58Z")

</div>

The documents that are full text searched based on `equipment_id` also have other metadata that relate each-other ie. same document type, document status, etc.

The visualization I want to see is star like at the center but documents with `equipment_id` found in the full text search query are related to each other using other metadata that are in the document.

I will try and play around with graph in my Kibana instance to gain a better understanding of the relationships I can build.

Thanks for your help.

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 13, 2018, 4:33pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/10 "2018-09-13T16:33:48Z")

</div>

> [@Andrew\_Stroz](#):
>
> relate each-other ie. same document type, document status,

Useful graphs are those that use fields with high-cardinality (many unique terms).  
Examples include bank accounts, email addresses or hashtags [1]. These create sparser, interesting shapes. Meaningful relationships exist between rarer terms.

If you choose fields with a small number of values (eg "gender" or your doc type/status fields) then you end up with "hairball" graphs with too many lines, connecting all the nodes. These tend to be much less interesting connections.

[1] [http://hivemindmap.com/](http://hivemindmap.com/)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 11, 2018, 4:33pm UTC](https://discuss.elastic.co/t/building-graph-relationships-between-documents/148452/11 "2018-10-11T16:33:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
