# Question on ingesting multiple csv files with repeated data

**URL:** <https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679>\
**Category:** Elasticsearch\
**Created:** [May 28, 2020, 7:27am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679 "2020-05-28T07:27:28Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![richylyq](https://avatars.discourse-cdn.com/v4/letter/r/41988e/32.png) [@richylyq](https://discuss.elastic.co/u/richylyq)\
**Post date:** [May 28, 2020, 7:27am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/1 "2020-05-28T07:27:29Z")

</div>

Hi, does elastic auto remove duplicated data when i ingest multiple csv files with repeated data in them. or is there a way to enable the removal?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 28, 2020, 8:36am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/2 "2020-05-28T08:36:24Z")

</div>

No.

Only if you are using the same `_id` in which case the duplicated row will overwrite the previous one.

---

<div class="post-metadata">

**Author:** ![richylyq](https://avatars.discourse-cdn.com/v4/letter/r/41988e/32.png) [@richylyq](https://discuss.elastic.co/u/richylyq)\
**Post date:** [May 28, 2020, 9:06am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/3 "2020-05-28T09:06:36Z")

</div>

ohhh i see. thanks for the prompt reply @dadoonet 🙇‍♂️  
how do I go about setting the id for the ingest? the data i have is sth like

```auto
1.csv
-------
2020-04-19,22,2899,2888,339,429,11,6588
2020-04-20,23,3398,3782,351,449,11,8014
2020-04-21,27,3566,4682,371,468,11,9125
2020-04-22,25,4209,4999,413,483,12,10141
--------

```

```auto
2.csv
--------
2020-04-20,23,1364,5824,351,441,11,8014
2020-04-21,27,1381,6875,371,460,11,9125
2020-04-22,25,1571,7645,413,475,12,10141
2020-04-23,26,1342,8874,434,490,12,11178

```

The way i am ingesting the data right now looks sth like this, but as per your suggestion, I am not sure on where to set the id, and how to go about doing it. and if the amount of data (in rows) is inconsistent, will the id be able to track that the particular row is repeated?

```auto
for x in sorted(files):
    print(x)
    with open(x) as f:
        reader = csv.DictReader(f)
        helpers.bulk(es, reader, index='testindex')

```

Just found an article on enrich policy with ingest processor,

> [How to enrich logs and metrics using an Elasticsearch ingest node | Elastic Blog](https://www.elastic.co/blog/how-to-enrich-logs-and-metrics-using-an-elasticsearch-ingest-node)

Is it something like this?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 2, 2020, 1:32pm UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/4 "2020-06-02T13:32:19Z")

</div>

If you don't define an `_id`, then elasticsearch will generate one for you.  
If the unique id in your case is the date of the event, then you can think of using the date as the `_id`.

So CSV1 will be like:

```auto
PUT index/_doc/2020-04-19
{ "date": "2020-04-19", ... }
PUT index/_doc/2020-04-20
{ "date": "2020-04-20", ... }

```

Then the second CSV will be something like:

```auto
PUT index/_doc/2020-04-20
{ "date": "2020-04-20", ... }
PUT index/_doc/2020-04-21
{ "date": "2020-04-21", ... }

```

Which means that at the end, you will have the following records:

```auto
{ "date": "2020-04-19", ... }
{ "date": "2020-04-20", ... }
{ "date": "2020-04-21", ... }

```

---

<div class="post-metadata">

**Author:** ![richylyq](https://avatars.discourse-cdn.com/v4/letter/r/41988e/32.png) [@richylyq](https://discuss.elastic.co/u/richylyq)\
**Post date:** [June 3, 2020, 9:37am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/5 "2020-06-03T09:37:32Z")

</div>

Thanks for the information! I think i got the solution for the \_id for my case already.

```auto
p.put_pipeline(id='attachment', body={
    'description': 'setting press release date to _id',
    'processors': [
        {
            "set": {
                "field": "_id",
                "value": "{{Press release date}}"
            }
        }
    ]
})

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 3, 2020, 9:42am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/6 "2020-06-03T09:42:19Z")

</div>

Yes. That will work.

Out of curiosity, why not providing the correct `_id` in your Python script instead of having to reprocess the json in an ingest pipeline?  
It will be faster if you do that on Python side.

---

<div class="post-metadata">

**Author:** ![richylyq](https://avatars.discourse-cdn.com/v4/letter/r/41988e/32.png) [@richylyq](https://discuss.elastic.co/u/richylyq)\
**Post date:** [June 8, 2020, 3:46am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/7 "2020-06-08T03:46:17Z")

</div>

haha, i guess there will always be a better way to work things out, but i think for now i will leave it as it is, i will update it if there is a need to. thanks anyways!

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2020, 3:49am UTC](https://discuss.elastic.co/t/question-on-ingesting-multiple-csv-files-with-repeated-data/234679/8 "2020-07-06T03:49:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
