# Some pdf can't be indexed

**URL:** <https://discuss.elastic.co/t/some-pdf-cant-be-indexed/149666>\
**Category:** Elasticsearch\
**Created:** [September 24, 2018, 12:27pm UTC](https://discuss.elastic.co/t/some-pdf-cant-be-indexed/149666 "2018-09-24T12:27:54Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Thierry\_Pasqualini](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/thierry_pasqualini/32/35805_2.png) [@Thierry\_Pasqualini](https://discuss.elastic.co/u/Thierry_Pasqualini)\
**Post date:** [September 24, 2018, 12:27pm UTC](https://discuss.elastic.co/t/some-pdf-cant-be-indexed/149666/1 "2018-09-24T12:27:54Z")

</div>

Hello,  
With some pdf, the result of the indexation via fscrawler is an empty content field (see below).

Any idea?  
Thanks.

```
{
* "_index": "lesdocsmv",
* "_type": "_doc",
* "_id": "b3d35554-eb3a-4bea-947a-b998ebf4f387",
* "_version": 1,
* "_score": 1,
* "_source": {
  * "content": " ",
  * "meta": {
    * "format": "application/pdf; version=1.4",
    * "creator_tool": "Canon iR-ADV C5235 ",
    * "created": "2018-09-06T07:11:55.000+0000",
    * "raw": {
      * "pdf:PDFVersion": "1.4",
      * "xmp:CreatorTool": "Canon iR-ADV C5235 ",
      * "access_permission:modify_annotations": "true",
      * "access_permission:can_print_degraded": "true",
      * "dcterms:created": "2018-09-06T07:11:55Z",
      * "dc:format": "application/pdf; version=1.4",
      * "xmpMM:DocumentID": "uuid:42d3905b-0000-8887-177f-b25700000000",
      * "pdf:docinfo:creator_tool": "Canon iR-ADV C5235 ",
```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 24, 2018, 12:41pm UTC](https://discuss.elastic.co/t/some-pdf-cant-be-indexed/149666/2 "2018-09-24T12:41:15Z")

</div>

Could you share your PDF document? You can DM it to me.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 24, 2018, 1:21pm UTC](https://discuss.elastic.co/t/some-pdf-cant-be-indexed/149666/4 "2018-09-24T13:21:28Z")

</div>

@Thierry_Pasqualini So the document is an image and does not contain text.  
The only way to extract text from it is to configure OCR.

Have a look at [https://fscrawler.readthedocs.io/en/fscrawler-2.5/user/tips.html?highlight=ocr#ocr-integration](https://fscrawler.readthedocs.io/en/fscrawler-2.5/user/tips.html?highlight=ocr#ocr-integration)

HTH

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 22, 2018, 1:22pm UTC](https://discuss.elastic.co/t/some-pdf-cant-be-indexed/149666/5 "2018-10-22T13:22:06Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
