# JavaEsSpark.saveToES not using pre-defined mapping fields while posting the data to ES cluster

**URL:** <https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761>\
**Category:** Elasticsearch\
**Tags:** es-hadoop\
**Created:** [February 28, 2017, 10:39am UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761 "2017-02-28T10:39:42Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![r.ganeshbabu](https://avatars.discourse-cdn.com/v4/letter/r/4da419/32.png) [@r.ganeshbabu](https://discuss.elastic.co/u/r.ganeshbabu)\
**Post date:** [February 28, 2017, 10:39am UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/1 "2017-02-28T10:39:42Z")

</div>

Hello All,

We have been trying to push the data to Elasticsearch cluster by using the following method in spark **"JavaEsSpark.saveToEs"**

I have defined the **"rd\_innovation\_config1"** index mappings and settings in ES cluster and below is the code snipet of JavaRDD to save the document in Elasticsearch

```
	JavaRDD<InnovationConfig> iConfigRDD = initParquetDf.javaRDD().map(y -> InnovationConfig.builder()
								            .isoCntryCode(y.getString(0))
		                        			.chrID(y.getString(1))
		                        			.chrValID(y.getString(2))
		                        			.processType(y.getString(4))
		                        			.activeID(y.getString(5))
		                        			.releaseInd(String.valueOf(y.getLong(6)))
		                        			.build());
			
    JavaEsSpark.saveToEs(iConfigRDD, "rd_innovation_config1/innovation_config");

```

Since **InnovationConfig.builder()** has all the pojo's and in which each has set **@jsonproperty** as like below,

```
package com.ogrds.datamodel.es.rdinnovationconfig;

import java.io.Serializable;
import com.fasterxml.jackson.annotation.JsonClassDescription;
import com.fasterxml.jackson.annotation.JsonProperty;
import com.modeliosoft.modelio.javadesigner.annotations.objid;
import lombok.Builder;
import lombok.Getter;

@objid ("0dcc4051-9213-435e-88da-50c7d6433b8f")
@JsonClassDescription("rd_innovation_config")
@Builder
@Getter
public class InnovationConfig implements Serializable {
    @objid ("99e5975b-401e-4aa6-8736-26e21348e4b5")
    @JsonProperty("ACTIVE_ID")
    public String activeID;

    @objid ("b71d26fd-4203-4ffc-a93a-ceef95aaf91e")
    @JsonProperty("CHR_ID")
    public String chrID;

    @objid ("2fba5712-ff54-4375-a3a2-e94a8a9dbf48")
    @JsonProperty("CHR_VAL_ID")
    public String chrValID;

    @objid ("ef742f1e-ff29-42a1-afca-172049e87b9e")
    @JsonProperty("ISO_COUNTRY_CODE")
    public String isoCntryCode;

    @objid ("f402da18-5123-43c2-b93f-86da44cd787a")
    @JsonProperty("PROCESS_TYPE")
    public String processType;

    @objid ("e2b7beb8-119c-4663-a98d-2055a62c70e7")
    @JsonProperty("RELEASE_IND")
    public String releaseInd;
}

```

But after saveToES operation done the values are not saved to the right field in the pre-defined mapping and seems @jsonproperty are not able to serialize properly???

 ![](https://us1.discourse-cdn.com/elastic/original/2X/a/a523e95b99057fd7e6b23bdfb4b8fd8fc0c0a021.png)

Please correct me if I am doing anything wrong and let me know your thoughts..

Thanks,  
Ganeshbabu R

---

<div class="post-metadata">

**Author:** ![rulanitee](https://avatars.discourse-cdn.com/v4/letter/r/b5e925/32.png) [@rulanitee](https://discuss.elastic.co/u/rulanitee)\
**Post date:** [February 28, 2017, 12:06pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/2 "2017-02-28T12:06:33Z")

</div>

Hi Ganeshbabu

Did you set es.index.auto.create to false in your spark config as follows:

sparkConf.set("es.index.auto.create", "false") -\> that way it does not create the index from InnovationConfig.builder().

Regards

---

<div class="post-metadata">

**Author:** ![r.ganeshbabu](https://avatars.discourse-cdn.com/v4/letter/r/4da419/32.png) [@r.ganeshbabu](https://discuss.elastic.co/u/r.ganeshbabu)\
**Post date:** [February 28, 2017, 12:20pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/3 "2017-02-28T12:20:21Z")

</div>

Hi @rulanitee

No, I have given like this in sparkconf.set("es.index.auto.create", "true")

Regards,  
Ganeshbabu R

---

<div class="post-metadata">

**Author:** ![rulanitee](https://avatars.discourse-cdn.com/v4/letter/r/b5e925/32.png) [@rulanitee](https://discuss.elastic.co/u/rulanitee)\
**Post date:** [February 28, 2017, 12:33pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/4 "2017-02-28T12:33:54Z")

</div>

Hi,

> [@r.ganeshbabu](#):
>
> I have defined the "rd\_innovation\_config1" index mappings and settings in ES cluster and below is the code snipet of JavaRDD to save the document in Elasticsearch

If you already created the index and have done the mappings, then you can set that config to false.

Regards

---

<div class="post-metadata">

**Author:** ![r.ganeshbabu](https://avatars.discourse-cdn.com/v4/letter/r/4da419/32.png) [@r.ganeshbabu](https://discuss.elastic.co/u/r.ganeshbabu)\
**Post date:** [February 28, 2017, 12:55pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/5 "2017-02-28T12:55:33Z")

</div>

> [@rulanitee](#):
>
> If you already created the index and have done the mappings, then you can set that config to false.

Yes @rulanitee I already created the index with mappings before pushing the data to ES but after pushing the data to ES the values are not mapped to the existing fields.

Below is the mapping of rd\_innovation\_config1 index,

You can see clearly there will be two sets of fields are mapped to the same index,

1. Manually created fields at the top (i.e UPPERCASE)
2. Automatically created fields at the bottom (i.e LOWERCASE)

 ![](https://us1.discourse-cdn.com/elastic/original/2X/2/2173ad3135abe63dfb7d1a918fa12c5f59a147e5.png)

Please let me know your thoughts on this...

Regards,  
Ganeshbabu R

---

<div class="post-metadata">

**Author:** ![rulanitee](https://avatars.discourse-cdn.com/v4/letter/r/b5e925/32.png) [@rulanitee](https://discuss.elastic.co/u/rulanitee)\
**Post date:** [February 28, 2017, 2:14pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/6 "2017-02-28T14:14:01Z")

</div>

Hi,

Oh i see now. The json annotations are being ignored. Quick question, do you want to save to ES as json or RDD?

As i see it, you want:

> [@r.ganeshbabu](#):
>
> JavaRDD\<InnovationConfig\> iConfigRDD

to be serialized using the json annotation you specified.

If that is the case, why not convert InnovationConfig to a json string and push to ES the way you expect the data to be. (JavaEsSpark.saveJsonToEs(your\_json\_object)) as [documented](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html#spark-write-json).

If I am not mistaken:

> [@r.ganeshbabu](#):
>
> JavaRDD\<InnovationConfig\> iConfigRDD

it is being serialized as is, as per the [documentation](https://www.elastic.co/guide/en/elasticsearch/hadoop/current/spark.html#spark-write-java).

Hope it helps ...

Regards

---

<div class="post-metadata">

**Author:** ![r.ganeshbabu](https://avatars.discourse-cdn.com/v4/letter/r/4da419/32.png) [@r.ganeshbabu](https://discuss.elastic.co/u/r.ganeshbabu)\
**Post date:** [February 28, 2017, 5:35pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/7 "2017-02-28T17:35:07Z")

</div>

> [@rulanitee](#):
>
> Quick question, do you want to save to ES as json or RDD?

Yes, I want to save as json in ES..

@rulanitee Could you please share some sample code how to get json string from the RDD?

As I am new to this technology and I stuck in getting the json strings from the RDD.

Please let me know your feedback

Thanks,  
Ganeshbabu R

---

<div class="post-metadata">

**Author:** ![rulanitee](https://avatars.discourse-cdn.com/v4/letter/r/b5e925/32.png) [@rulanitee](https://discuss.elastic.co/u/rulanitee)\
**Post date:** [March 1, 2017, 8:02am UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/8 "2017-03-01T08:02:02Z")

</div>

Hi,

Sorry for getting back to you late. Here is a quick and dirty way of sending data as json:

import com.fasterxml.jackson.annotation.JsonClassDescription;  
import com.fasterxml.jackson.annotation.JsonProperty;

@JsonClassDescription("rd\_innovation\_config")  
public class InnovationConfig {

```
@JsonProperty("ACTIVE_ID")
public String activeID;

@JsonProperty("CHR_ID")
public String chrID;

@JsonProperty("CHR_VAL_ID")
public String chrValID;

@JsonProperty("ISO_COUNTRY_CODE")
public String isoCntryCode;

@JsonProperty("PROCESS_TYPE")
public String processType;

@JsonProperty("RELEASE_IND")
public String releaseInd;

```

}

public final class JavaSpark {

```
public static void main (String[] args) throws Exception{
    SparkConf conf = new SparkConf().setAppName("EsJavaClient").setMaster("local");
    conf.set("es.index.auto.create", "false");        
    // and the rest of your settings
    
    JavaSparkContext sc = new JavaSparkContext(conf);

    InnovationConfig innovationConfig = new InnovationConfig();
    innovationConfig.activeID = "0";
    innovationConfig.chrID = "33986";
    innovationConfig.chrValID = "72291112";
    innovationConfig.isoCntryCode = "7";
    innovationConfig.processType = "INITIATIVE";
    innovationConfig.releaseInd = "14733550531568";

    InnovationConfig innovationConfig1 = new InnovationConfig();
    innovationConfig1.activeID = "1";
    innovationConfig1.chrID = "4444444";
    innovationConfig1.chrValID = "5555555";
    innovationConfig1.isoCntryCode = "750";
    innovationConfig1.processType = "ONE_TESTING";
    innovationConfig1.releaseInd = "14733550531568";

    ObjectMapper mapper = new ObjectMapper(); // USING JACKSON, WE CONVERT OUR OBJECT TO JSON
    String innovationConfigJsonString = mapper.writeValueAsString(innovationConfig);
    String innovationConfig1JsonString = mapper.writeValueAsString(innovationConfig1);

    JavaRDD<String> stringRDD = sc.parallelize(ImmutableList.of(innovationConfigJsonString, innovationConfig1JsonString)); // WE ADD THE JSON STRINGS TO AN IMMUTABLE LIST
    JavaEsSpark.saveJsonToEs(stringRDD,"rd_innovation_config1/innovation_config"); // PUSH TO ES
}

```

}

And the result is as follows:

 ![](https://us1.discourse-cdn.com/elastic/original/2X/b/b968930e22d561cf41d6e908e4517af4fce2f1d8.png)

Hope this will set you in the right direction, let me know how it goes ...

Regards

---

<div class="post-metadata">

**Author:** ![james.baiera](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/james.baiera/32/10209_2.png) [@james.baiera](https://discuss.elastic.co/u/james.baiera)\
**Post date:** [March 12, 2017, 5:41pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/9 "2017-03-12T17:41:24Z")

</div>

@r.ganeshbabu ES-Hadoop does not support your JSON annotations on the fields. Instead, it only respects the names of the getter fields for the java bean. The names from those methods are converted to the implied property names and formatted in `lowerCamelCase` style when converting the bean into JSON. If you want to use these annotations, I would suggest calling the `.map()` method on the RDD and within it converting the bean into JSON using which ever object mappers you prefer, and then saving the corresponding JSON to Elasticsearch using `es.json.input = true` property.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 9, 2017, 5:41pm UTC](https://discuss.elastic.co/t/javaesspark-savetoes-not-using-pre-defined-mapping-fields-while-posting-the-data-to-es-cluster/76761/10 "2017-04-09T17:41:51Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
