# Stackoverflow on Elasticsearch file indexation with ingest-attachment

**URL:** https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455
**Category:** Elasticsearch
**Created:** [October 27, 2020, 3:16pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455 "2020-10-27T15:16:24Z")
**Posts on this page:** 9
**Page:** 1

<div class="post-metadata">

### Author: ![Mathulbrich](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mathulbrich/32/63127_2.png) [@Mathulbrich](https://discuss.elastic.co/u/Mathulbrich)
#### Post date: [October 27, 2020, 3:16pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/1 "2020-10-27T15:16:24Z")

</div>

Hi,

Our company recently had some connectivity issues with Elasticsearch. Suddenly, on some days of the week the cluster had RED status and in the logs the only information found was that the node had been disconnected.

On further investigation, we ended up discovering that one of our customers was trying to index an electronic docx file that contained a embedded PDF inside it, and when it try to make the request for the ingest-attachment plugin, Elastic service was killed by StackoverflowError. See the collected logs from the node:

```
[2020-10-25T01:09:42,214][INFO][o.e.c.m.MetaDataCreateIndexService] [0f0qpyW] [2c9b808575556dcd0175556e03e30000] creating index, cause [api], templates [], shards [1]/[0], mappings [_doc]
[2020-10-25T01:09:42,422][INFO][o.e.c.r.a.AllocationService] [0f0qpyW] Cluster health status changed from [YELLOW] to [GREEN] (reason: [shards started [[2c9b808575556dcd0175556e03e30000][0]] ...]).
[2020-10-25T01:09:42,540][INFO][o.e.c.m.MetaDataMappingService] [0f0qpyW] [2c9b808575556dcd0175556e03e30000/_z-DGoGXRyGOZkDNJp7pKw] update_mapping [_doc]
[2020-10-25T01:09:45,473][ERROR][o.e.b.ElasticsearchUncaughtExceptionHandler] [0f0qpyW] fatal error in thread [elasticsearch[0f0qpyW][write][T#2]], exiting
java.lang.StackOverflowError: null
	at java.util.HashMap.hash(HashMap.java:339) ~[?:?]
	at java.util.LinkedHashMap.get(LinkedHashMap.java:440) ~[?:?]
	at org.apache.pdfbox.cos.COSDictionary.getDictionaryObject(COSDictionary.java:188) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree.getKids(PDPageTree.java:135) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree.access$200(PDPageTree.java:38) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:166) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:169) ~[?:?]

```

This embedded pdf is not even being opened by Adobe Acrobat Reader, accusing some error that it may be corrupted. But would it make sense for Elastic to die for that? Shouldn't he ignore cases like this so as not to compromise the entire cluster?

**Elasticsearch version:** 6.8.0  
**Host:** AWS ElasticSearch Service

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [October 27, 2020, 6:04pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/2 "2020-10-27T18:04:42Z")

</div>

I'm not sure which version of Tika was used for this old version, but I'd first try with a brand new cluster using 7.9.3 first.

If you can reproduce the issue, could you share your docx file to the Tika project and open an issue?

BTW did you look at [Cloud by Elastic](https://www.elastic.co/cloud), also available if needed from [AWS Marketplace](https://aws.amazon.com/marketplace/pp/B01N6YCISK) ?

Cloud by elastic is one way to have access to **all features** , all managed by us. Think about what is there yet like Security, Monitoring, Reporting, SQL, Canvas, Maps UI, Alerting and built-in solutions named [Observability](https://www.elastic.co/observability), [Security](https://www.elastic.co/security), [Enterprise Search](https://www.elastic.co/enterprise-search) and what is coming next 🙂 ...

---

<div class="post-metadata">

### Author: ![Mathulbrich](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mathulbrich/32/63127_2.png) [@Mathulbrich](https://discuss.elastic.co/u/Mathulbrich)
#### Post date: [October 28, 2020, 6:29pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/4 "2020-10-28T18:29:52Z")

</div>

I have tested most recently version (7.9.3), but same problem occurs.

Seeing the log with more attention, the problem occurs on PDPageTree wich comes from Apache PDFBox (wich I'm investigating now).

```
[2020-10-28T14:52:27,849][ERROR][o.e.b.ElasticsearchUncaughtExceptionHandler] [es-node] fatal error in thread [elasticsearch[es-node][write][T#7]], exiting
java.lang.StackOverflowError: null
	at java.util.regex.Pattern$BmpCharPredicate.lambda$union$2(Pattern.java:5692) ~[?:?]
	at java.util.regex.Pattern$BmpCharProperty.match(Pattern.java:4019) ~[?:?]
	at java.util.regex.Pattern$GroupHead.match(Pattern.java:4855) ~[?:?]
	at java.util.regex.Pattern$BranchConn.match(Pattern.java:4763) ~[?:?]
	at java.util.regex.Pattern$GroupTail.match(Pattern.java:4886) ~[?:?]
	at java.util.regex.Pattern$BmpCharProperty.match(Pattern.java:4020) ~[?:?]
	at java.util.regex.Pattern$GroupHead.match(Pattern.java:4855) ~[?:?]
	at java.util.regex.Pattern$Branch.match(Pattern.java:4800) ~[?:?]
	at java.util.regex.Pattern$Branch.match(Pattern.java:4798) ~[?:?]
	at java.util.regex.Pattern$Branch.match(Pattern.java:4798) ~[?:?]
	at java.util.regex.Pattern$BranchConn.match(Pattern.java:4763) ~[?:?]
	at java.util.regex.Pattern$GroupTail.match(Pattern.java:4886) ~[?:?]
	at java.util.regex.Pattern$BmpCharPropertyGreedy.match(Pattern.java:4394) ~[?:?]
	at java.util.regex.Pattern$GroupHead.match(Pattern.java:4855) ~[?:?]
	at java.util.regex.Pattern$Branch.match(Pattern.java:4800) ~[?:?]
	at java.util.regex.Pattern$BranchConn.match(Pattern.java:4763) ~[?:?]
	at java.util.regex.Pattern$GroupTail.match(Pattern.java:4886) ~[?:?]
	at java.util.regex.Pattern$BmpCharProperty.match(Pattern.java:4020) ~[?:?]
	at java.util.regex.Pattern$BmpCharPropertyGreedy.match(Pattern.java:4394) ~[?:?]
	at java.util.regex.Pattern$GroupHead.match(Pattern.java:4855) ~[?:?]
	at java.util.regex.Pattern$Branch.match(Pattern.java:4800) ~[?:?]
	at java.util.regex.Pattern$BmpCharProperty.match(Pattern.java:4020) ~[?:?]
	at java.util.regex.Pattern$Start.match(Pattern.java:3673) ~[?:?]
	at java.util.regex.Matcher.search(Matcher.java:1729) ~[?:?]
	at java.util.regex.Matcher.find(Matcher.java:773) ~[?:?]
	at java.util.Formatter.parse(Formatter.java:2702) ~[?:?]
	at java.util.Formatter.format(Formatter.java:2655) ~[?:?]
	at java.util.Formatter.format(Formatter.java:2609) ~[?:?]
	at java.lang.String.format(String.java:3292) ~[?:?]
	at java.util.logging.SimpleFormatter.format(SimpleFormatter.java:176) ~[?:?]
	at java.util.logging.StreamHandler.publish(StreamHandler.java:199) ~[?:?]
	at java.util.logging.ConsoleHandler.publish(ConsoleHandler.java:95) ~[?:?]
	at java.util.logging.Logger.log(Logger.java:979) ~[?:?]
	at java.util.logging.Logger.doLog(Logger.java:1006) ~[?:?]
	at java.util.logging.Logger.logp(Logger.java:1172) ~[?:?]
	at org.apache.commons.logging.impl.Jdk14Logger.log(Jdk14Logger.java:87) ~[?:?]
	at org.apache.commons.logging.impl.Jdk14Logger.warn(Jdk14Logger.java:260) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree.getKids(PDPageTree.java:159) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree.access$200(PDPageTree.java:41) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:183) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]
	at org.apache.pdfbox.pdmodel.PDPageTree$PageIterator.enqueueKids(PDPageTree.java:186) ~[?:?]

```

But I think that problems that come from plugins used by ES should not affect the integrity of the node. Maybe it would be interesting to have some treat to just return an error in the request instead of killing the entire node. There is a flag that can be set in the pipelines (ignore\_failure), but this did not solve this situation. Don't you think so?

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [October 28, 2020, 8:36pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/5 "2020-10-28T20:36:23Z")

</div>

Of course it should not. But I think that some exceptions are not really catch-able easily.

Please share your document so we can reproduce and report in the upstream project.

---

<div class="post-metadata">

### Author: ![Mathulbrich](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mathulbrich/32/63127_2.png) [@Mathulbrich](https://discuss.elastic.co/u/Mathulbrich)
#### Post date: [October 28, 2020, 9:00pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/6 "2020-10-28T21:00:43Z")

</div>

> [@dadoonet](#):
>
> Of course it should not. But I think that some exceptions are not really catch-able easily.
> 
> Please share your document so we can reproduce and report in the upstream project.

I removed sensitive data to not expose our client's confidential informations, but any document with that embedded PDF will result in error.

> **[Google Drive: Sign-in](https://accounts.google.com/v3/signin/identifier?continue=https%3A%2F%2Fdrive.google.com%2Ffile%2Fd%2F1ZJ2cxQ_KvduirRLPsdgWn81eHDPZ_bax%2Fview%3Fusp%3Dsharing&followup=https%3A%2F%2Fdrive.google.com%2Ffile%2Fd%2F1ZJ2cxQ_KvduirRLPsdgWn81eHDPZ_bax%2Fview%3Fusp%3Dsharing&ifkv=Ab5oB3r4tOP_HXV4hJqYDaCs5iaFAfjH8KQ8GVtTwexnwn5xW4h95fbB3jzEigula2U_Rf0uGzmn&osid=1&passive=1209600&service=wise&flowName=GlifWebSignIn&flowEntry=ServiceLogin&dsh=S1546368332%3A1725077474256881&ddm=0)**
>
> Access Google Drive with a Google account (for personal use) or Google Workspace account (for business use).

Thanks for comprehension, if there's some update at the theme would be nice that be shared here.

---

<div class="post-metadata">

### Author: ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)
#### Post date: [November 4, 2020, 4:41pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/7 "2020-11-04T16:41:39Z")

</div>

I can reproduce it with Tika 1.24.1.

I opened an issue:

> **[\[TIKA-3224\] Stackoverflow with Embedded PDF in DOCX document - ASF JIRA](https://issues.apache.org/jira/browse/TIKA-3224)**

---

<div class="post-metadata">

### Author: ![tallison](https://avatars.discourse-cdn.com/v4/letter/t/258eb7/32.png) [@tallison](https://discuss.elastic.co/u/tallison)
#### Post date: [November 4, 2020, 8:01pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/8 "2020-11-04T20:01:48Z")

</div>

I extracted the PDF and reproduced the problem with straight PDFBox's ExtractText. The PDF file looks corrupt...see my comments here: [https://issues.apache.org/jira/browse/TIKA-3224?focusedCommentId=17226366&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-17226366](https://issues.apache.org/jira/browse/TIKA-3224?focusedCommentId=17226366&page=com.atlassian.jira.plugin.system.issuetabpanels%3Acomment-tabpanel#comment-17226366)

I opened [https://issues.apache.org/jira/browse/PDFBOX-5009](https://issues.apache.org/jira/browse/PDFBOX-5009)

---

<div class="post-metadata">

### Author: ![tallison](https://avatars.discourse-cdn.com/v4/letter/t/258eb7/32.png) [@tallison](https://discuss.elastic.co/u/tallison)
#### Post date: [November 4, 2020, 8:11pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/10 "2020-11-04T20:11:59Z")

</div>

We strongly encourage keeping Tika processing out of the same JVM/VM/M/rack/data center, as your indexer or even the ingest process.

This can be done with tika-batch, the ForkParser or tika-server. These three options remove the potential for catastrophic problems affecting the indexing process.

We do what we can when we find problems on Apache Tika, but we know and loudly proclaim that robust parsing of untrusted documents must be run in an isolated JVM.

We're happy to help you @dadoonet make FSCrawler and/or ingest-attachment more robust if you have an interest...

---

<div class="post-metadata">

### Author: ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)
#### Post date: [December 2, 2020, 8:11pm UTC](https://discuss.elastic.co/t/stackoverflow-on-elasticsearch-file-indexation-with-ingest-attachment/253455/11 "2020-12-02T20:11:59Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
