# Ingest-attachment not parsing docx

**URL:** <https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671>\
**Category:** Elasticsearch\
**Created:** [May 29, 2018, 12:03pm UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671 "2018-05-29T12:03:45Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![ottadini](https://avatars.discourse-cdn.com/v4/letter/o/b19c9b/32.png) [@ottadini](https://discuss.elastic.co/u/ottadini)\
**Post date:** [May 29, 2018, 12:03pm UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/1 "2018-05-29T12:03:45Z")

</div>

I have ES 6.0.1 on my localhost, and the same on a AWS server. On my localhost I installed the ingest-attachment plugin, and tried to index some docx files. They are not parsed, and ES returns content-length of 0. On the server, using the same code, they are parsed as expected.

on localhost the document looks like:

```
{
  "_index": "testing",
  "_type": "documents",
  "_id": "45061422-cf3a-4d23-8b20-0ad27a272735",
  "_version": 1,
  "found": true,
  "_source": {
    "path": """D:\draft 1.docx""",
    "filename": "draft 1.docx",
    "attachment": {
      "content_type": "application/x-tika-ooxml",
      "content_length": 0
    }
  }
}

```

On the server it looks like:

```
{
  "_index": "testing",
  "_type": "documents",
  "_id": "afb12bcf-a8c8-43b4-838f-10336b28e91a",
  "_version": 1,
  "found": true,
  "_source": {
    "path": """D:\draft 1.docx""",
    "filename": "draft 1.docx",
    "attachment": {
      "date": "2018-05-09T05:51:00Z",
      "content_type": "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
      "author": "Tom W",
      "language": "en",
      "title": "Letter",
      "content": """<snipped>""",
      "content_length": 28769
    }
  }
}

```

Why the difference in the content\_type? What have I done wrong on my localhost?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 29, 2018, 2:07pm UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/2 "2018-05-29T14:07:17Z")

</div>

Interesting. What kind of servers do you have?  
May be something related to a Locale setting?

---

<div class="post-metadata">

**Author:** ![ottadini](https://avatars.discourse-cdn.com/v4/letter/o/b19c9b/32.png) [@ottadini](https://discuss.elastic.co/u/ottadini)\
**Post date:** [May 30, 2018, 5:05am UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/3 "2018-05-30T05:05:22Z")

</div>

I've tried on yet another server, and no problem there either.

I think it must be an issue with apache tika. My laptop has jdk v10, while the servers have v8. If I parse the docx document directly in tika, it does work, but spits out some warnings and errors. Perhaps they are enough to ruin the party for ES. Here's an example of the output from tika:

on starting the jar:

> WARNING: An illegal reflective access operation has occurred  
> WARNING: Illegal reflective access by org.apache.poi.openxml4j.util.ZipSecureFile$1 (file:/D:/opt/tika-1.17/tika-app-1.17.jar) to field java.io.FilterInputStream.in  
> WARNING: Please consider reporting this to the maintainers of org.apache.poi.openxml4j.util.ZipSecureFile$1  
> WARNING: Use --illegal-access=warn to enable warnings of further illegal reflective access operations  
> WARNING: All illegal access operations will be denied in a future release

on parsing the docx:

> X-Parsed-By: org.apache.tika.parser.DefaultParser  
> X-Parsed-By: org.apache.tika.parser.microsoft.ooxml.OOXMLParser  
> X-TIKA:EXCEPTION:embedded\_stream\_exception: java.lang.ClassCastException: org.apache.poi.openxml4j.util.ZipSecureFile$ThresholdInputStream cannot be cast to java.base/java.util.zip.ZipFile$ZipFileInputStream

and

> X-TIKA:EXCEPTION:embedded\_stream\_exception: java.lang.ClassCastException: org.apache.poi.openxml4j.util.ZipSecureFile$ThresholdInputStream cannot be cast to java.base/java.util.zip.ZipFile$ZipFileInputStream

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 30, 2018, 5:16am UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/4 "2018-05-30T05:16:51Z")

</div>

What happens if you are using jdk8 on your laptop?

---

<div class="post-metadata">

**Author:** ![ottadini](https://avatars.discourse-cdn.com/v4/letter/o/b19c9b/32.png) [@ottadini](https://discuss.elastic.co/u/ottadini)\
**Post date:** [May 30, 2018, 6:08am UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/5 "2018-05-30T06:08:15Z")

</div>

I don't have it installed. I'll try it soon and report back here.

---

<div class="post-metadata">

**Author:** ![ottadini](https://avatars.discourse-cdn.com/v4/letter/o/b19c9b/32.png) [@ottadini](https://discuss.elastic.co/u/ottadini)\
**Post date:** [May 30, 2018, 6:47am UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/6 "2018-05-30T06:47:45Z")

</div>

According to the docs, ElasticSearch honours the JAVA\_HOME environment variable. I installed jdk8 (alongside an existing jdk10), then changed JAVA\_HOME to point to the jdk8. I cannot get docx files to parse still.

I don't know what else to try, apart from uninstalling jdk10, which I don't want to do.

---

<div class="post-metadata">

**Author:** ![ottadini](https://avatars.discourse-cdn.com/v4/letter/o/b19c9b/32.png) [@ottadini](https://discuss.elastic.co/u/ottadini)\
**Post date:** [May 30, 2018, 6:57am UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/7 "2018-05-30T06:57:47Z")

</div>

Idiotic mistake - I forgot to restart the terminal. It had kept the old JAVA\_HOME env variable.

Now it's happily indexing docx files with jdk8.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 30, 2018, 10:24am UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/8 "2018-05-30T10:24:21Z")

</div>

Great. Then I believe this is something to ask to the Tika team on [https://issues.apache.org/jira/browse/TIKA](https://issues.apache.org/jira/browse/TIKA)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 27, 2018, 10:24am UTC](https://discuss.elastic.co/t/ingest-attachment-not-parsing-docx/133671/9 "2018-06-27T10:24:21Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
