# Indexing Binary vs text

**URL:** <https://discuss.elastic.co/t/indexing-binary-vs-text/16655>\
**Category:** Elasticsearch\
**Created:** [March 27, 2014, 7:09pm UTC](https://discuss.elastic.co/t/indexing-binary-vs-text/16655 "2014-03-27T19:09:29Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![IronMan2014](https://avatars.discourse-cdn.com/v4/letter/i/94ad74/32.png) [@IronMan2014](https://discuss.elastic.co/u/IronMan2014)\
**Post date:** [March 27, 2014, 7:09pm UTC](https://discuss.elastic.co/t/indexing-binary-vs-text/16655/1 "2014-03-27T19:09:29Z")

</div>

I have couple of simple questions that I would like to clear up:

#1: For transportClient & cluster of two hosts: Do I have to add both hosts  
to the client, or is it enough to add just one of them and the yml(s) will  
take care of the clustering?

.addTransportAddress(new InetSocketTransportAddress(host[0], port))

.addTransportAddress(new InetSocketTransportAddress(host[1], port));

#2: Assume I have the following document structure:

jdoc{  
"title":"my title"  
"uid":"ux1234"  
"tags":"ES"  
"date":"1/1/2011"  
"content":"Content of doc goes here"  
}

//This is for my Binary attachment for Binaries (PDF)

putMappingResponse = new PutMappingRequestBuilder(  
client.admin().indices() ).setIndices(INDEX\_NAME).setType(INDEX\_TYPE).  
setSource(

```
                                      XContentFactory.jsonBuilder().

```

startObject()

```
                                        .startObject(INDEX_TYPE)

                                        .startObject("properties")

                                          //pdf

                                            .startObject("file")

                                                            .field( 

```

"type", "attachment" )

```
                                               .startObject("fields")

                                                   .startObject("title")

                                                       .field("store", 

```

"yes")

```
                                                   .endObject()

                                                   .startObject("file")

                                                       .field("store", 

```

"yes")

```
                                                       .field( 

```

"term\_vector", "with\_positions\_offsets" )

```
                                                   .endObject()

                                               .endObject()

                                            .endObject()

                                          .endObject()

                                        .endObject()

                                      .endObject()

                                  ).execute().actionGet();

```

void indexDocument(JSONObject jdoc){

bulkProcessor.add(Requests.indexRequest(INDEX\_NAME).type(INDEX\_TYPE).id(  
jDoc.getString("uid")).source(jDoc.toString()));  
}

void indexBinaryDocument(JSONObject jdoc){

XContentBuilder source = jsonBuilder().startObject()

```
                                     .field("file", jDoc.getString(

```

CONTENT)) //from tika Binary 64

```
                                     .field("uid",jDoc.getString(UID))

                                     .field("date",jDoc.getString(DATE))
                                     ....

                                    .endObject();

```

bulkProcessor.add(Requests.indexRequest(INDEX\_NAME).type(INDEX\_TYPE).source  
(source));  
}

My Question:

Based on the document, I either call indexDocument for normal text docs or  
indexBinaryDocument. However, this is confusing, I want to be able to call  
one index function like "indexDocument" above without having to specify  
source again for binary, In other words, if the document is binary, why do  
I have to tell it about the "file" field again, couldn't I just replace the  
"content" field with the 64 base encoded text, everything else in the  
document is the same, only the content field is different? Somehow I feel  
both of should one of the same?

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
To view this discussion on the web visit [https://groups.google.com/d/msgid/elasticsearch/2b59fd33-9d10-4b65-8b7a-f40d03bdbc83%40googlegroups.com](https://groups.google.com/d/msgid/elasticsearch/2b59fd33-9d10-4b65-8b7a-f40d03bdbc83%40googlegroups.com).  
For more options, visit [https://groups.google.com/d/optout](https://groups.google.com/d/optout).

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 1:40am UTC](https://discuss.elastic.co/t/indexing-binary-vs-text/16655/2 "2017-07-06T01:40:02Z")

</div>


