# Attachment content was been truncated

**URL:** <https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907>\
**Category:** Elasticsearch\
**Created:** [May 12, 2016, 1:52pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907 "2016-05-12T13:52:35Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![carlosq](https://avatars.discourse-cdn.com/v4/letter/c/8e8cbc/32.png) [@carlosq](https://discuss.elastic.co/u/carlosq)\
**Post date:** [May 12, 2016, 1:52pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907/1 "2016-05-12T13:52:35Z")

</div>

Hi,

when I index an attachment which is MS word document and it contains double quotes.  
after indexing, the content after the double quotes was truncated.

for example, this is the attachment:

Test attachment, "test document", after this, it has more words.

after indexing, the content does not have "after this, it has more words."

I'm using Mapper Attachment Plugin V2.3.2

Thank you in advance,

Carlos

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 12, 2016, 3:03pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907/2 "2016-05-12T15:03:22Z")

</div>

There is a limit by default of extracted content.  
IIRC it's 10000 characters.

---

<div class="post-metadata">

**Author:** ![MarkOStewart](https://avatars.discourse-cdn.com/v4/letter/m/898d66/32.png) [@MarkOStewart](https://discuss.elastic.co/u/MarkOStewart)\
**Post date:** [May 12, 2016, 4:20pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907/3 "2016-05-12T16:20:53Z")

</div>

Try to cut and paste the text out of MS Word and into a text editor like Notepad. Save the file and try your import again.

I never use MS Word except to do word processing.  
Word may not be your issue however I have found nothing but trouble using text from MS Word Documents as Word uses RTF and injects a bunch of hidden stuff that causes issues when parsing data.

Just a thought.  
MS

---

<div class="post-metadata">

**Author:** ![carlosq](https://avatars.discourse-cdn.com/v4/letter/c/8e8cbc/32.png) [@carlosq](https://discuss.elastic.co/u/carlosq)\
**Post date:** [May 12, 2016, 5:15pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907/4 "2016-05-12T17:15:15Z")

</div>

my testing document has only less 100 characters.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 12, 2016, 5:52pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907/5 "2016-05-12T17:52:16Z")

</div>

Can you provide a full script which reproduces your issue?

---

<div class="post-metadata">

**Author:** ![carlosq](https://avatars.discourse-cdn.com/v4/letter/c/8e8cbc/32.png) [@carlosq](https://discuss.elastic.co/u/carlosq)\
**Post date:** [May 12, 2016, 6:03pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907/6 "2016-05-12T18:03:18Z")

</div>

I use the example from the link [https://gist.github.com/karmi/5594127](https://gist.github.com/karmi/5594127).  
the only change I made is the test.doc. I added double quotes in the test.doc

for example (before change):  
Test  
RTF document.

Lorem  
ipsum dolor.

(after change):  
Test  
RTF document.  
"test document"  
Lorem  
ipsum dolor.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 5, 2017, 10:52pm UTC](https://discuss.elastic.co/t/attachment-content-was-been-truncated/49907/7 "2017-07-05T22:52:00Z")

</div>


