# Filter out duplicate products

**URL:** <https://discuss.elastic.co/t/filter-out-duplicate-products/165718>\
**Category:** Elasticsearch\
**Created:** [January 25, 2019, 8:21am UTC](https://discuss.elastic.co/t/filter-out-duplicate-products/165718 "2019-01-25T08:21:10Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![vermin](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/vermin/32/40208_2.png) [@vermin](https://discuss.elastic.co/u/vermin)\
**Post date:** [January 25, 2019, 8:21am UTC](https://discuss.elastic.co/t/filter-out-duplicate-products/165718/1 "2019-01-25T08:21:11Z")

</div>

Once a day I parse a product feed and index the products to Elasticsearch.

I want to keep it up-to-date and since delete operations are really expensive in Elasticsearch I choose to use rollover index. Every day the products are written to a different index `products-yyyy-mm-dd`.

Now when I search for products I want to look in todays and yesterdays index but avoid returning duplicates for the same **product\_id**.

Field Collapsing seems like the way to go so by simply doing:

```auto
"collapse" : {
	"field" : "product_id" 
}
"sort": ["imported_at"]

```

I get unique hits by **product\_id**. Unfortunately the total number of hits and aggregation is not affected by this and therefore my filters and counters are not accurate.

**How can I rethink my setup?** Currently duplicate products will have the same `_type` and `_id` but different `_index`.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 22, 2019, 8:21am UTC](https://discuss.elastic.co/t/filter-out-duplicate-products/165718/2 "2019-02-22T08:21:13Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
