# Ingesting documents (.pbix) to elasticsearch

**URL:** <https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549>\
**Category:** Elasticsearch\
**Created:** [November 6, 2020, 3:35pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549 "2020-11-06T15:35:38Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![Yuval\_David](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuval_david/32/78377_2.png) [@Yuval\_David](https://discuss.elastic.co/u/Yuval_David)\
**Post date:** [November 6, 2020, 3:35pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/1 "2020-11-06T15:35:38Z")

</div>

so I would like to create an application which would get pbix files insert those files into elastic and then we could create a rest application which could query elastic for full text search over those pbix files.

however, I have problems in putting pbix files into elastic. does anyone have any idea how to do it?

to be more precise does someone knows how to extract that data from .pbix files into some sort of document that can be stored in elastic?  
i would like to be able to do full text search on the pbix files content

i have only seen guides on how to do the oppsite(i.e use elastic data in power bi)

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 8, 2020, 11:01pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/2 "2020-11-08T23:01:25Z")

</div>

What are pbix files?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 9, 2020, 9:06am UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/3 "2020-11-09T09:06:18Z")

</div>

[PBIX File - What is a .pbix file and how do I open it?](https://fileinfo.com/extension/pbix) says:

> A PBIX file is a document created by Power BI Desktop, a Microsoft application used to create reports and visualizations. It contains queries, data models, visualizations, settings, and reports added by the user.

I'm not sure what @Yuval_David would actually want to index from this content. I mean that it does not contain data as per say.

---

<div class="post-metadata">

**Author:** ![Yuval\_David](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuval_david/32/78377_2.png) [@Yuval\_David](https://discuss.elastic.co/u/Yuval_David)\
**Post date:** [November 9, 2020, 9:28pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/4 "2020-11-09T21:28:08Z")

</div>

but it does contain text so i would like to index all the text that appers in the pbix file

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 9, 2020, 9:48pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/5 "2020-11-09T21:48:58Z")

</div>

Ok, then you will need to figure out a way to extract that from the file. It's not something that is native to the Elastic Stack.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 10, 2020, 3:12am UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/6 "2020-11-10T03:12:45Z")

</div>

Adding that I have no idea if FSCrawler supports it. May be give it a try?

---

<div class="post-metadata">

**Author:** ![Yuval\_David](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuval_david/32/78377_2.png) [@Yuval\_David](https://discuss.elastic.co/u/Yuval_David)\
**Post date:** [November 10, 2020, 11:49am UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/7 "2020-11-10T11:49:26Z")

</div>

> [@dadoonet](#):
>
> Adding that I have no idea if FSCrawler supports it. May be give it a try?

i am not sure that it is supported by FTCrawler i even tried to look, thought that maybe someone who dealt with this kind of files know how to extract that data and ingest it into elastic

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 10, 2020, 12:50pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/8 "2020-11-10T12:50:09Z")

</div>

FSCrawler supports whatever you have in that list:

> **[Apache Tika – Supported Document Formats](https://tika.apache.org/1.24.1/formats.html#Supported_Document_Formats)**

If I understand correctly, `pbix` files are XML files, you should be able to read it with whatever parser (can be logstash if you wish) and transform the content to a json document which is then sent to Elasticsearch.

Not sure it will make sense though. But may be you can share somewhere a typical `pbix` document so we can look at it? And tell us exactly what you want to index from this document.

---

<div class="post-metadata">

**Author:** ![Yuval\_David](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuval_david/32/78377_2.png) [@Yuval\_David](https://discuss.elastic.co/u/Yuval_David)\
**Post date:** [November 10, 2020, 1:23pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/9 "2020-11-10T13:23:24Z")

</div>

i can but it wont help you much its an compressed file.  
you can download one here [Sales & Returns sample report](https://go.microsoft.com/fwlink/?linkid=2113239).  
(taken from here : [https://docs.microsoft.com/en-us/power-bi/create-reports/sample-datasets](https://docs.microsoft.com/en-us/power-bi/create-reports/sample-datasets))  
but since its compressed i can quite figure out where to take the data from, you can chance the extention to .zip and then extract the file but i still cant figure much out

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 10, 2020, 1:38pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/10 "2020-11-10T13:38:34Z")

</div>

So what kind of content do you want to index from this file?

---

<div class="post-metadata">

**Author:** ![Yuval\_David](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuval_david/32/78377_2.png) [@Yuval\_David](https://discuss.elastic.co/u/Yuval_David)\
**Post date:** [November 10, 2020, 2:34pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/11 "2020-11-10T14:34:53Z")

</div>

as i said when you open this kind of file with the power bi, you get some sort of a report, with text inside it -\> thats what i want to index

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 10, 2020, 4:38pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/12 "2020-11-10T16:38:38Z")

</div>

Can you share a screen capture of the file you shared, opened in PowerBI which shows the text you would like to see indexed in elasticsearch?

---

<div class="post-metadata">

**Author:** ![Yuval\_David](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yuval_david/32/78377_2.png) [@Yuval\_David](https://discuss.elastic.co/u/Yuval_David)\
**Post date:** [November 10, 2020, 5:25pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/13 "2020-11-10T17:25:11Z")

</div>

as you can see there are serval tabs etc... alot of text/numerical values to index (i dont care if its all stored as one big text field

 ![image](https://us1.discourse-cdn.com/elastic/original/3X/7/1/7112c4b57015db344ed54602a8c4a6fe8409ef16.png)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 12, 2020, 3:18pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/14 "2020-11-12T15:18:46Z")

</div>

So you want to index something like:

```auto
{
   "content": "Category Breakdown Power BI $52K Word $36K OneNote $31K PowerPoint $30K ... Store Breakdown Fama $40K Contoso $39K .... "
}

```

That's it?

I gave a try with Apache Tika (via FSCrawler) and here is what has been extracted:

[https://gist.githubusercontent.com/dadoonet/57e0e2d2eebc379249f3ea75c6e991da/raw/4050aad09d68fbbfeb0b3ab71d1074cbcb9377d9/extracted.txt](https://gist.githubusercontent.com/dadoonet/57e0e2d2eebc379249f3ea75c6e991da/raw/4050aad09d68fbbfeb0b3ab71d1074cbcb9377d9/extracted.txt)

Not sure it helps to index that type of content.

So I don't think there is a solution out of the box here. You probably need to build something by yourself.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 10, 2020, 3:18pm UTC](https://discuss.elastic.co/t/ingesting-documents-pbix-to-elasticsearch/254549/15 "2020-12-10T15:18:52Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
