# Match data from two sets and mark entries

**URL:** <https://discuss.elastic.co/t/match-data-from-two-sets-and-mark-entries/282731>\
**Category:** Elasticsearch\
**Tags:** docker, painless\
**Created:** [August 28, 2021, 5:22pm UTC](https://discuss.elastic.co/t/match-data-from-two-sets-and-mark-entries/282731 "2021-08-28T17:22:38Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![mydeadvictoria](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mydeadvictoria/32/93843_2.png) [@mydeadvictoria](https://discuss.elastic.co/u/mydeadvictoria)\
**Post date:** [August 28, 2021, 5:22pm UTC](https://discuss.elastic.co/t/match-data-from-two-sets-and-mark-entries/282731/1 "2021-08-28T17:22:38Z")

</div>

I have two relatively large sets of people data (first\_name, last\_name, birth\_date etc).  
And I need to 'match' them, because some people are present in both sets and I need to find and mark them (by adding their ID from the other set). I've come up with this solution:

1. Load set A into Elasticsearch index (6M entries)
2. During load of set B, for each entry, make an update-by-query request which looks for people with the same first\_name(Text), last\_name(Text) & birth\_date(Date) and adds `B_id` field through a simple script (painless)

So for example, `A` index roughly might look like this:

| \_id | first\_name | last\_name | birth\_date | b\_id |
| --- | --- | --- | --- | --- |
| yJzdWIiOiJ | demo1 | qwerty | 2001.11.04 | 4d5323bd91c2 |
| MKPj0ILgq1 | demo2 | demo2 | 1995.11.11 | null |
| oueUg3sBO | demo512 | demo512 | 2000.05.16 | null |

Here, an entry with id `yJzdWIiOiJ` got matched with an entry with id `4d5323bd91c2` from index `B`.

An example request for such flow:

```json
{
    "script":{
        "id":"id-field-adding-script",
        "params":{
            "value":9150456064
        }
    },
    "size":1000,
    "query":{
        "bool":{
            "must":[
                {
                    "match":{
                        "first_name":{
                            "query":"Ryan",
                            "operator":"AND",
                            "prefix_length":0,
                            "max_expansions":50,
                            "fuzzy_transpositions":true,
                            "lenient":false,
                            "zero_terms_query":"NONE",
                            "auto_generate_synonyms_phrase_query":true,
                            "boost":1.0
                        }
                    }
                },
                {
                    "match":{
                        "last_name":{
                            "query":"Jewel",
                            "operator":"AND",
                            "prefix_length":0,
                            "max_expansions":50,
                            "fuzzy_transpositions":true,
                            "lenient":false,
                            "zero_terms_query":"NONE",
                            "auto_generate_synonyms_phrase_query":true,
                            "boost":1.0
                        }
                    }
                },
                {
                    "match":{
                        "birth_date":{
                            "query":"1992-09-11",
                            "operator":"OR",
                            "prefix_length":0,
                            "max_expansions":50,
                            "fuzzy_transpositions":true,
                            "lenient":false,
                            "zero_terms_query":"NONE",
                            "auto_generate_synonyms_phrase_query":true,
                            "boost":1.0
                        }
                    }
                }
            ],
            "adjust_pure_negative":true,
            "boost":1.0
        }
    }
}

```

This request is sent for every entry (via Java Spring backend). it is not as slow as I expected it to be, but I still want to improve it. In my tests, set `A` has 6M entries and set `B` has ~250K. It takes up to 2 hours to both match set `B` with set `A` and import set `B` afterwards (actually it takes 10k entries, matches them and then imports into ES, repeats).

**The question is** : How do I improve this, make it faster? Or maybe you can suggest another approach.

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [August 30, 2021, 11:16am UTC](https://discuss.elastic.co/t/match-data-from-two-sets-and-mark-entries/282731/2 "2021-08-30T11:16:21Z")

</div>

There is no built-in function to solve this very concrete use-case.

You could possibly do this with two single searches.. albeit longer ones 🙂

How about running two scroll/PIT searches, one against index a, one against index b. Sort by last name, first name, birth\_date and then iterate through the index with less documents and see compare to the index with more indices, when you find a matching document.

The idea would be something like leap-frogging/zig zag joining the data and thus find the two matches. Whenever you find a match, run an update query (and potentially bulk those to save some more resources).

Does that make sense?

---

<div class="post-metadata">

**Author:** ![mydeadvictoria](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mydeadvictoria/32/93843_2.png) [@mydeadvictoria](https://discuss.elastic.co/u/mydeadvictoria)\
**Post date:** [August 30, 2021, 4:03pm UTC](https://discuss.elastic.co/t/match-data-from-two-sets-and-mark-entries/282731/3 "2021-08-30T16:03:25Z")

</div>

I'm relatively new to ES, not sure I get the idea. Could you please show some example?

---

<div class="post-metadata">

**Author:** ![vincenbr](https://avatars.discourse-cdn.com/v4/letter/v/8edcca/32.png) [@vincenbr](https://discuss.elastic.co/u/vincenbr)\
**Post date:** [August 30, 2021, 5:07pm UTC](https://discuss.elastic.co/t/match-data-from-two-sets-and-mark-entries/282731/4 "2021-08-30T17:07:06Z")

</div>

When I hear "match data" in elasticsearch, the [percolator](https://www.elastic.co/guide/en/elasticsearch/reference/current/query-dsl-percolate-query.html) automatically comes to my mind...  
I can't dig into it further now, but maybe index set B in a percolation index (with term queries), and percolate documents from set A into it, in order to retrieve matching ids ?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 27, 2021, 5:07pm UTC](https://discuss.elastic.co/t/match-data-from-two-sets-and-mark-entries/282731/5 "2021-09-27T17:07:56Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
