# Getting the extracted content from the attachment mapper plugin

**URL:** <https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528>\
**Category:** Elasticsearch\
**Created:** [November 2, 2012, 12:14pm UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528 "2012-11-02T12:14:49Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [November 2, 2012, 12:14pm UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528/1 "2012-11-02T12:14:49Z")

</div>

Hi there,

I just played around with the attachment mapper plugin and wondered if  
I can access the parsedContent (as in AttachmentMapper.java:309),  
which contains the tika-parsed content of the document, in any way.  
When doing a simple GET on the document I only see the base64 encoded  
value which I pushed.

I'd like to do some special text extraction in my documents (like  
searching for dates in them) after indexing. Alternatively I could  
call the tika code a second time in my own application, seems a bit  
dirty though.

Any hints appreciated.. possibly I just overlooked something when  
skimming through the source and it is totally easy 🙂

--Alexander

--

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 2, 2012, 12:50pm UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528/2 "2012-11-02T12:50:49Z")

</div>

I tried also to play with it some time ago but did not succeed with mimetype autodetection. ☹

I posted here something about it without answer: [https://groups.google.com/forum/m/?fromgroups#!search/Tika$20Pilato$20content$20type/elasticsearch/Ne5\_uOKlAAk](https://groups.google.com/forum/m/?fromgroups#!search/Tika%2420Pilato%2420content%2420type/elasticsearch/Ne5_uOKlAAk)

So if someone answers to Alexander, it would be nice to have a look also at my old post. :-/

--  
David 😉  
Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs

Le 2 nov. 2012 à 13:14, Alexander Reelsen [alr@spinscale.de](mailto:alr@spinscale.de) a écrit :

Hi there,

I just played around with the attachment mapper plugin and wondered if  
I can access the parsedContent (as in AttachmentMapper.java:309),  
which contains the tika-parsed content of the document, in any way.  
When doing a simple GET on the document I only see the base64 encoded  
value which I pushed.

I'd like to do some special text extraction in my documents (like  
searching for dates in them) after indexing. Alternatively I could  
call the tika code a second time in my own application, seems a bit  
dirty though.

Any hints appreciated.. possibly I just overlooked something when  
skimming through the source and it is totally easy 🙂

--Alexander

--

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [November 2, 2012, 7:34pm UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528/3 "2012-11-02T19:34:04Z")

</div>

I noticed this statement too in the attachment mapper plugin, because I was  
interested in handling binary content with ES.

How do you want the binary content to be exposed? At shard level? Or via  
node transport?

The REST API uses JSON (XContentBuilder) which is the reason for base64.  
With websockets, I can imagine direct exposition of binary streams to a  
Java client, but it's still a lot to do when transporting huge data from  
the shard level to the requesting node without blowing up the heap (e.g.  
chunked streams).

Jörg

On Friday, November 2, 2012 1:51:00 PM UTC+1, David Pilato wrote:

> I tried also to play with it some time ago but did not succeed with  
> mimetype autodetection. ☹
> 
> I posted here something about it without answer:  
> [Redirecting to Google Groups](https://groups.google.com/forum/m/?fromgroups#!search/Tika$20Pilato$20content$20type/elasticsearch/Ne5_uOKlAAk)
> 
> So if someone answers to Alexander, it would be nice to have a look also  
> at my old post. :-/
> 
> --  
> David 😉  
> Twitter : @dadoonet / @elasticsearchfr / @scrutmydocs
> 
> Le 2 nov. 2012 à 13:14, Alexander Reelsen \<[a...@spinscale.de](mailto:a...@spinscale.de) \<javascript:\>\>  
> a écrit :
> 
> Hi there,
> 
> I just played around with the attachment mapper plugin and wondered if  
> I can access the parsedContent (as in AttachmentMapper.java:309),  
> which contains the tika-parsed content of the document, in any way.  
> When doing a simple GET on the document I only see the base64 encoded  
> value which I pushed.
> 
> I'd like to do some special text extraction in my documents (like  
> searching for dates in them) after indexing. Alternatively I could  
> call the tika code a second time in my own application, seems a bit  
> dirty though.
> 
> Any hints appreciated.. possibly I just overlooked something when  
> skimming through the source and it is totally easy 🙂
> 
> --Alexander
> 
> --

--

---

<div class="post-metadata">

**Author:** ![gondo\_2](https://avatars.discourse-cdn.com/v4/letter/g/f08c70/32.png) [@gondo\_2](https://discuss.elastic.co/u/gondo_2)\
**Post date:** [April 23, 2013, 1:42pm UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528/4 "2013-04-23T13:42:22Z")

</div>

did you solve this somehow?  
is it at all possible to access parsed plain text via REST API?

@Jörg Prante: why binary? why not plaintext?

On Friday, 2 November 2012 23:14:52 UTC+11, Alexander Reelsen wrote:

> Hi there,
> 
> I just played around with the attachment mapper plugin and wondered if  
> I can access the parsedContent (as in AttachmentMapper.java:309),  
> which contains the tika-parsed content of the document, in any way.  
> When doing a simple GET on the document I only see the base64 encoded  
> value which I pushed.
> 
> I'd like to do some special text extraction in my documents (like  
> searching for dates in them) after indexing. Alternatively I could  
> call the tika code a second time in my own application, seems a bit  
> dirty though.
> 
> Any hints appreciated.. possibly I just overlooked something when  
> skimming through the source and it is totally easy 🙂
> 
> --Alexander

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [April 23, 2013, 2:41pm UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528/5 "2013-04-23T14:41:08Z")

</div>

Yes. Just store fields you need and ask for it when searching.  
Have a look at this GIST: [Testing FSRiver with Mapper attachment and check metadata extracted · GitHub](https://gist.github.com/dadoonet/5310075)

It gives some clues (using FSRiver but you can run the same test without FSRiver)

## HTH

David Pilato | Technical Advocate | [Elasticsearch.com](http://Elasticsearch.com)  
@dadoonet | @elasticsearchfr | @scrutmydocs

Le 23 avr. 2013 à 15:42, gondo [gondar@webdesigners.sk](mailto:gondar@webdesigners.sk) a écrit :

> did you solve this somehow?  
> is it at all possible to access parsed plain text via REST API?
> 
> @Jörg Prante: why binary? why not plaintext?
> 
> On Friday, 2 November 2012 23:14:52 UTC+11, Alexander Reelsen wrote:  
> Hi there,
> 
> I just played around with the attachment mapper plugin and wondered if  
> I can access the parsedContent (as in AttachmentMapper.java:309),  
> which contains the tika-parsed content of the document, in any way.  
> When doing a simple GET on the document I only see the base64 encoded  
> value which I pushed.
> 
> I'd like to do some special text extraction in my documents (like  
> searching for dates in them) after indexing. Alternatively I could  
> call the tika code a second time in my own application, seems a bit  
> dirty though.
> 
> Any hints appreciated.. possibly I just overlooked something when  
> skimming through the source and it is totally easy 🙂
> 
> --Alexander
> 
> --  
> You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
> To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
> For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![gondo\_2](https://avatars.discourse-cdn.com/v4/letter/g/f08c70/32.png) [@gondo\_2](https://discuss.elastic.co/u/gondo_2)\
**Post date:** [April 24, 2013, 12:48pm UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528/6 "2013-04-24T12:48:38Z")

</div>

hi  
thanks David, that helped a little.  
although it looks like the file is stored as base64 version of original  
file, and text extraction by tika is done during indexation AND also during  
extraction, am i right?  
i was hoping to store just the plaintext (without original document in any  
format), but i guess i ll have to extract text separately by using tika and  
then check and store just the result  
anyways thanks for making it more clear for me

On Friday, 2 November 2012 23:14:52 UTC+11, Alexander Reelsen wrote:

> Hi there,
> 
> I just played around with the attachment mapper plugin and wondered if  
> I can access the parsedContent (as in AttachmentMapper.java:309),  
> which contains the tika-parsed content of the document, in any way.  
> When doing a simple GET on the document I only see the base64 encoded  
> value which I pushed.
> 
> I'd like to do some special text extraction in my documents (like  
> searching for dates in them) after indexing. Alternatively I could  
> call the tika code a second time in my own application, seems a bit  
> dirty though.
> 
> Any hints appreciated.. possibly I just overlooked something when  
> skimming through the source and it is totally easy 🙂
> 
> --Alexander

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:39am UTC](https://discuss.elastic.co/t/getting-the-extracted-content-from-the-attachment-mapper-plugin/9528/7 "2017-07-06T02:39:48Z")

</div>


