# Indexing PDFs directly

**URL:** <https://discuss.elastic.co/t/indexing-pdfs-directly/199124>\
**Category:** Elasticsearch\
**Created:** [September 11, 2019, 6:59pm UTC](https://discuss.elastic.co/t/indexing-pdfs-directly/199124 "2019-09-11T18:59:55Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![jtgajda](https://avatars.discourse-cdn.com/v4/letter/j/f4b2a3/32.png) [@jtgajda](https://discuss.elastic.co/u/jtgajda)\
**Post date:** [September 11, 2019, 6:59pm UTC](https://discuss.elastic.co/t/indexing-pdfs-directly/199124/1 "2019-09-11T18:59:56Z")

</div>

I was able to upload 1650 PDF documents into Elasticsearch using the ingestor plugin.  
However, it looks like I am locked into the schema that was generated and am also not happy with search performance.  
As such, I am trying to load PDF using my own mapping that excludes the pdf content field that is being searched.  
The documents do successfully load in but I am unable to search within the PDF content though I am able to search my custom metadata.  
Following is my mapping and the python ES code used to load the data.

Mapping:  
mapping = '''  
{  
"\_source": {  
"excludes": [  
"content"  
]  
},  
"properties": {  
"title": {  
"type": "text"  
},  
"ip": {  
"type": "text"  
},  
"content": {  
"type": "text",  
"store": "true"  
},  
"query": {  
"properties": {  
"match\_phrase": {  
"properties": {  
"content": {  
"type": "text",  
"fields": {  
"keyword": {  
"type": "keyword",  
"ignore\_above": 256  
}  
}  
}  
}  
}  
}  
}  
}  
}  
'''

Python snippet:

# Read PDF as binary, convert to BASE64, back to string and then index.

with open(file, "rb") as pdf\_file:  
enc\_pdf = base64.b64encode(pdf\_file.read()).decode('ascii')  
respdf = es.index(index='jdocs2', doc\_type='\_doc', id=mfn, body={'content': enc\_pdf, "ip": ip, "title": title })

The file indexes successfully but I get no results when trying query the PDF content for known terms.  
res = es.search(index="jdocs2", body={"size": 2, "query": {"match": {"title": "guide"}}})

Am I missing a step prior to indexing or are there other suggestions?  
Any help would be appreciated!

Regards,  
Jeff Gajda

---

<div class="post-metadata">

**Author:** ![Mark\_Harwood](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mark_harwood/32/10538_2.png) [@Mark\_Harwood](https://discuss.elastic.co/u/Mark_Harwood)\
**Post date:** [September 13, 2019, 2:22pm UTC](https://discuss.elastic.co/t/indexing-pdfs-directly/199124/2 "2019-09-13T14:22:02Z")

</div>

You'd need to share an example of your parsed `content` field.  
I don't know much about PDF formats but I expect there's much more to parsing than base64 decoding.  
Have you tried using the TIKA library directly to take more control over which fields to index?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 13, 2019, 4:52pm UTC](https://discuss.elastic.co/t/indexing-pdfs-directly/199124/3 "2019-09-13T16:52:25Z")

</div>

Also have a look at FSCrawler project. It extracts a lot of metadata as well. Might be useful.

---

<div class="post-metadata">

**Author:** ![jtgajda](https://avatars.discourse-cdn.com/v4/letter/j/f4b2a3/32.png) [@jtgajda](https://discuss.elastic.co/u/jtgajda)\
**Post date:** [September 16, 2019, 2:44pm UTC](https://discuss.elastic.co/t/indexing-pdfs-directly/199124/4 "2019-09-16T14:44:40Z")

</div>

Mark, David

Following is a sample of the output of the 64encoding and conversion back to string via the following python snipped:  
with open(file, "rb") as pdf\_file:  
enc\_pdf = base64.b64encode(pdf\_file.read()).decode('ascii')  
print(enc\_pdf)

JVBERi0xLjYNJeLjz9MNCjQ2NSAwIG9iag08PC9GaWx0ZXIvRmxhdGVEZWNvZGUvRmlyc3QgMzcvTGVu  
Z3RoIDM1Mi9OIDUvVHlwZS9PYmpTdG0+PnN0cmVhbQ0KAMLoHK6ppeTW8z4SOnBk4K3L3ZEs9dKsisPJ  
jr1SRSAjvBFrd63ny8sVh4UWWovapio2+/M/QxKwOuVwhyu4agobDDv8IA2zJDg3SD9QL0KC/1xYVUQq  
Bm8TPkq+7XkuZyvV1nIi5wosuzjW6hu+sLgmOhXAGIeWVM3v5QeaBLhzBedK0UnJpEtBGO69ZsmbZK1L  
T3sckFFZ26Avf5Nbv2mcqKtpcDxdYwHWDPZMFb2RnxsN6MamUD3s+6WY......

I have not tried using the TIKA library as I am not that knowledgeab le yet. I will try the FSCrawler as suggested and may come back with more questions.  
Thank you both for your insight.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 14, 2019, 2:44pm UTC](https://discuss.elastic.co/t/indexing-pdfs-directly/199124/5 "2019-10-14T14:44:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
