# Data mismatches happening while sending data to Elastic Search index using pyspark

**URL:** <https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188>\
**Category:** Elasticsearch\
**Tags:** datastreams\
**Created:** [April 16, 2025, 11:07am UTC](https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188 "2025-04-16T11:07:53Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![yolo1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yolo1/32/142597_2.png) [@yolo1](https://discuss.elastic.co/u/yolo1)\
**Post date:** [April 16, 2025, 11:07am UTC](https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188/1 "2025-04-16T11:07:53Z")

</div>

Any idea why data sent through df.write. in pyspark the data doesn't match correctly . in the backend the data is correct.

---

<div class="post-metadata">

**Author:** ![Keith\_Massey](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/keith_massey/32/83666_2.png) [@Keith\_Massey](https://discuss.elastic.co/u/Keith_Massey)\
**Post date:** [April 17, 2025, 2:46pm UTC](https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188/2 "2025-04-17T14:46:07Z")

</div>

Can you provide your index mappings and a pyspark script to reproduce this?

---

<div class="post-metadata">

**Author:** ![RainTown](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/raintown/32/140206_2.png) [@RainTown](https://discuss.elastic.co/u/RainTown)\
**Post date:** [April 17, 2025, 5:52pm UTC](https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188/3 "2025-04-17T17:52:50Z")

</div>

> [@yolo1](#):
>
> the data doesn't match correctly

can you be more specific on the differences you see?

---

<div class="post-metadata">

**Author:** ![yolo1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yolo1/32/142597_2.png) [@yolo1](https://discuss.elastic.co/u/yolo1)\
**Post date:** [April 20, 2025, 2:06pm UTC](https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188/4 "2025-04-20T14:06:41Z")

</div>

Hi, so  
this is a sample hive code

```auto
create table db.sample
(id string,
count bigint,
time timestamp)
stored by 'org.elasticsearch.hadoop.hive.EsStorageHandler'
tblproperties(
'es.nodes.wan.only'='true',
'es.nodes'=esnode,
'es.resource'=index,
'es.mapping.names'='time:@timestamp');
Insert into table db.sample select * from data1;

create table db.sample_2
(id string,
status string,
count bigint,
time timestamp)
stored by 'org.elasticsearch.hadoop.hive.EsStorageHandler'
tblproperties(
'es.nodes.wan.only'='true',
'es.nodes'=esnode,
'es.resource'=index,
'es.mapping.names'='time:@timestamp');
Insert into table db.sample_2 select * from data2;

```

and this is my sample spark code

```auto
df_1 = data1.select("id","count","time")
df_2 = data2.select("id","status","count","time")
df_1.write.format("org.elasticsearch.spark.sql")\
.option('es.nodes.wan.only','true')\
.option('es.nodes',es_node)\
.option('es.resource',index)\
.option('es.mapping.names','time:@timestamp')\
.mode('append')\
.save(index)

df_2.write.format("org.elasticsearch.spark.sql")\
.option('es.nodes.wan.only','true')\
.option('es.nodes',es_node)\
.option('es.resource',index)\
.option('es.mapping.names','time:@timestamp')\
.mode('append')\
.save(index)

```

I am using spark 2.4.4 rn .  
So the issue that i see is whenever i run my spark code each successive time either some data gets duplicated or is missing .

No problem with hive. I am using elasticsearch hadoop v8 jar for this.

Currently since i had a deadline i am now doing processing in spark saving to a temp table and then using hive to transfer the data. IDk why the spark script didn't work. Also i have like 8 dataframes which i am inserting but the data quantity is small . you can assume 450 to 1500 rows and max i think 3000 rows

---

<div class="post-metadata">

**Author:** ![elasticforme](https://avatars.discourse-cdn.com/v4/letter/e/f05b48/32.png) [@elasticforme](https://discuss.elastic.co/u/elasticforme)\
**Post date:** [April 21, 2025, 7:25pm UTC](https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188/5 "2025-04-21T19:25:04Z")

</div>

elasticsearch wil insert data as it comes with \_id autogenerated. are you pushing same record again? if so you have duplicate in elastic.

you need to put lot more info to understand what is going on.

---

<div class="post-metadata">

**Author:** ![yolo1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yolo1/32/142597_2.png) [@yolo1](https://discuss.elastic.co/u/yolo1)\
**Post date:** [May 19, 2025, 10:24am UTC](https://discuss.elastic.co/t/data-mismatches-happening-while-sending-data-to-elastic-search-index-using-pyspark/377188/6 "2025-05-19T10:24:40Z")

</div>

no but the schema for each df was different . I found a workaround , where i am now inserting data to a temporary table in spark and then creating a external table in hive with elasticsearch as storage and writing data to it using select statement.

Still not sure why i was getting mismatches when sending data directly from spark
