# How to prevent duplicate and has null value documents with fingerprint

**URL:** <https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150>\
**Category:** Logstash\
**Created:** [July 8, 2021, 8:02am UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150 "2021-07-08T08:02:04Z")\
**Posts on this page:** 18\
**Page:** 1

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 8, 2021, 8:02am UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/1 "2021-07-08T08:02:04Z")

</div>

```auto
my csv file =>
name,surname,age,email,phone
Harry,Potter,18,NULL,NULL
Harry,Potter,NULL,harrypotter@gmail.com,+955555555
Harry,Potter,NULL,harrypotter@gmail.com,NULL
Harry,Potter,NULL,NULL,+955555555

```

When I want to detect and delete duplicate documents with  
fingerprint method, it creates a new document for each row.

```auto
filter {
    fingerprint {
        key => "1234ABCD"
        method => "MD5"
        source => ["name","surname","age","email","phone"]
        target => "[@metadata][generated_id]"
    }
}
output {
    stdout { codec => dots }
    elasticsearch {
        index => "null_problem_fingerprint"
        document_id => "%{[@metadata][generated_id]}"
        action => 'update'
    }
}

```

If I specify only the name and surname fields for the source as in the code blog below  
this time it does not read the other rows after reading the first row.

```auto
filter {
    fingerprint {
        key => "1234ABCD"
        method => "MD5"
        source => ["name","surname"]
        target => "[@metadata][generated_id]"
    }
}
output {
    stdout { codec => dots }
    elasticsearch {
        index => "null_problem_fingerprint"
        document_id => "%{[@metadata][generated_id]}"
        action => 'update'
    }
}

```

dear friends please help me! İ want to see just one document like this;

`Harry,Potter,18,harrypotter@gmail.com,+955555555`

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 8, 2021, 1:02pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/2 "2021-07-08T13:02:49Z")

</div>

Hi @Busra_Duygu Welcome to the community.

I think you need to [concatenate the sources](https://www.elastic.co/guide/en/logstash/current/plugins-filters-fingerprint.html#plugins-filters-fingerprint-concatenate_sources) the sources

`concatenate_sources => true`

I also think perhaps you want to use `doc_as_upsert` see [here](https://www.elastic.co/guide/en/logstash/current/plugins-outputs-elasticsearch.html#plugins-outputs-elasticsearch-doc_as_upsert)

`doc_as_upsert => true`

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 8, 2021, 1:39pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/3 "2021-07-08T13:39:12Z")

</div>

> [@stephenb](#):
>
> doc\_as\_upsert =\> true

First of all thank you very much for replying 🙂 but again added all lines to elastic as document.

```auto
filter {
      fingerprint {
        key => "1234ABCD"
        method => "UUID"
        source => ["name","surname","age","email","phone"]
        target => "[@metadata][generated_id]"
        concatenate_sources => "true"
      }
}

output {
    stdout { codec => dots }
    elasticsearch {
        index => "null_problem_fingerprint"
        document_id => "%{[@metadata][generated_id]}"
        doc_as_upsert => "true"
    }
}

```

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 8, 2021, 1:52pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/4 "2021-07-08T13:52:20Z")

</div>

First Try

```
    method => "SHA1"

```

I don't think you want UUID  
_"If set to `UUID` , a [UUID](https://en.wikipedia.org/wiki/Universally_unique_identifier) will be generated. The result will be random and thus not a consistent hash."_

Also when I look at your input data each row IS unique? across all the source values so I would expect each row to be a new unique row in the results

when you specified only

```
    source => ["name","surname"]

```

What was is actual result ... I would expect to be only the LAST ROW since all rows match the the fingerprint criteria so each row updates with the next.

So Better question (and perhaps I should have asked that first) with the 4 input rows you show what do you expect / want the output to be?

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 8, 2021, 1:55pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/5 "2021-07-08T13:55:36Z")

</div>

> [@stephenb](#):
>
> What was is actual result ... I would expect to be only the LAST ROW since all rows match the the fingerprint criteria so each row updates with the next.

Don't forget pipeline.ordered and pipeline.workers.

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 8, 2021, 2:07pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/6 "2021-07-08T14:07:31Z")

</div>

Ahh thanks @Badger

So in logstash.yml following correct?

```auto
pipeline.ordered : true (or auto if workers set to 1)

pipeline.workers : 1

```

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 8, 2021, 2:17pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/7 "2021-07-08T14:17:07Z")

</div>

`Also when I look at your input data each row IS unique? across all the source values so I would expect each row to be a new unique row in the results`  
yes each line in the csv file has unique records but they are all information of the same person.  
For example, in the bill of lading data, a firm makes more than one export or import. The first record has the address of the company, but no phone number. In the second record, the company has a phone number, but not an address. I want to get only company information from bill of lading data. There are two documents of the same company. I want to see only one document. If this is the document I want to see, I want it to contain both phone number and address information.

The content of the document I want to see is as follows:  
`Harry,Potter,18,harrypotter@gmail.com,+955555555`

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 8, 2021, 2:27pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/8 "2021-07-08T14:27:46Z")

</div>

To the logstash.yml file

```auto
pipeline.ordered: auto
pipeline.workers : 1

```

I added these but it still didn't work. The codes are like this;

```auto
filter {
     aggregate {
        task_id => "%{name}"
        code => "map['sql_duration'] = 0"
        map_action => "create"
      }

       fingerprint {
         key => "1234ABCD"
         method => "SHA1"
         source => ["name","surname"]
         target => "[@metadata][generated_id]"
         concatenate_sources => "true"
       }
}
output {
     stdout { codec => dots }
     elasticsearch {
         index => "null_problem_fingerprint"
         document_id => "%{[@metadata][generated_id]}"
         doc_as_upsert => "true"
     }
}

```

Result like this:

```auto
{
  "took" : 0,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "null_problem_fingerprint",
        "_type" : "_doc",
        "_id" : "e6ed065d89c16d449ff08e93418e589d3e217256",
        "_score" : 1.0,
   "_source" : {
          "name" : "Harry",
          "surname" : "Potter",
          "@timestamp" : "2021-07-08T14:23:05.629Z",
          "age" : "18",
          "phone" : "+955555555",
          "@version" : "1",
          "email" : null
        }
      }
    ]
  }
}

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 8, 2021, 2:34pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/9 "2021-07-08T14:34:07Z")

</div>

There is no reason to use a key for a fingerprint filter. It will happily use a hash rather than a digest if you do not set a key.

doc\_as\_upsert will update fields on the document with fields from the event. If those fields are null then they will still get overwritten. You need to remove fields from the event if you do not want them to be set on the document. So for

```
Harry,Potter,NULL,harrypotter@gmail.com,NULL

```

you need to delete the age and phone number fields before sending them to elasticsearch.

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 9, 2021, 5:55am UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/10 "2021-07-09T05:55:30Z")

</div>

I removed the key parameter as you said, but I couldn't figure out how to remove the phone number and age fields from the event. Could you please elaborate a little more?

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 9, 2021, 12:29pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/11 "2021-07-09T12:29:31Z")

</div>

@Badger  
I did what you said I ran the following codes to delete the fields with null values while adding the lines from the csv file to the first index. Then again I tried to block this duplicate data using the fingerprint method` doc_as_upsert => "true"` using this as well but without success.

null\_problem.conf:

```auto
input{
    file { 
      path => ".../null_problem.csv"
      start_position => "beginning"
      sincedb_path => "NUL" 
    }
}
filter{
    csv{        
        autodetect_column_names => "true"
        separator => ","
        skip_header => "true"
        columns => ["name","surname","age","email","phone"]
    }
    if [age] and [email] and [phone] == "" {
      drop { }
    }
    mutate { 
        remove_field =>["path", "host", "message", "@version", "@timestamp", "trade_date"]
    }
}
output{
    elasticsearch { 
        hosts => "http://localhost:9200"
        index => "null_problem"
        document_type => "_doc"
    }
    stdout {}
}

```

null\_problem\_finger :

```auto
input {
  elasticsearch {
    hosts => "localhost"
    index => "null_problem"
    query => '{ "sort": ["_doc"] }'
  }
}
filter {
    fingerprint {
      method => "SHA1"
      source => ["name","surname"]
      target => "[@metadata][generated_id]"
      concatenate_sources => "true"
    }
}
output {
    stdout { codec => dots }
    elasticsearch {
        index => "null_problem_fingerprint"
        document_id => "%{[@metadata][generated_id]}"
        doc_as_upsert => "true"
    }
}

```

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [July 9, 2021, 5:10pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/12 "2021-07-09T17:10:45Z")

</div>

> [@Busra\_Duygu](#):
>
> ```auto
> if [age] and [email] and [phone] == "" {
> drop { }
> }
> 
> ```

That says 'if the age field exists, and the email field exists, and the phone field is equal to "" then delete the event'. It's probably not what you want. If the csv literally contains the string "NULL" then what you want is

```
if [email] == "NULL" { mutate { remove_field => ["email"] } }

```

etc.

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 12, 2021, 5:44am UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/13 "2021-07-12T05:44:41Z")

</div>

@Badger  
I tried the code blog below to delete the empty fields and it did, but I still couldn't get the result I wanted.

```auto
ruby {
        code => "
            def walk_hash(parent, path, hash)
                path << parent if parent
                hash.each do |key, value|
                walk_hash(key, path, value) if value.is_a?(Hash)
                @paths << (path + [key]).map {|p| '[' + p + ']' }.join('')
                end
                path.pop
            end
            @paths = []
            walk_hash(nil, [], event.to_hash)
            @paths.each do |path|
                value = event.get(path)
                event.remove(path) if value.nil? || (value.respond_to?(:empty?) && value.empty?)
            end
            "
    }

```

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 12, 2021, 1:14pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/14 "2021-07-12T13:14:06Z")

</div>

@Badger  
To achieve these, I first added the csv file to the null\_problem index, then I created an index called null\_problem\_finger to organize these duplicate documents with the fingerprint method, but I was unsuccessful.

null\_problem index=\>

```auto
input{
    file { 
      path => ".../null_problem.csv"
      start_position => "beginning"
      sincedb_path => "NUL" 
    }
}
filter{
    csv{        
        autodetect_column_names => "true"
        separator => ","
        skip_header => "true"
        columns => ["name","surname","age","email","phone"]
    }
    mutate { 
        remove_field =>["path", "host", "message", "@version", "@timestamp", "trade_date"]
    }
    ruby {
        code => "
            def walk_hash(parent, path, hash)
                path << parent if parent
                hash.each do |key, value|
                walk_hash(key, path, value) if value.is_a?(Hash)
                @paths << (path + [key]).map {|p| '[' + p + ']' }.join('')
                end
                path.pop
            end
            @paths = []
            walk_hash(nil, [], event.to_hash)
            @paths.each do |path|
                value = event.get(path)
                event.remove(path) if value.nil? || (value.respond_to?(:empty?) && value.empty?)
            end
            "
    }
}
output{
    elasticsearch { 
        hosts => "http://localhost:9200"
        index => "null_problem"
        document_type => "_doc"
    }
    stdout {}
}

```

null\_problem\_fingerprint index =\>

```auto
input {
  elasticsearch {
    hosts => "localhost"
    index => "null_problem"
    query => '{ "sort": ["_doc"] }'
  }
}
filter{  
    fingerprint {
    method => "SHA1"
    source => ["name","surname","age","email","phone"]
    target => "[@metadata][generated_id]"
    concatenate_sources => "true"   
  }
  mutate { 
        remove_field =>["path", "host", "message", "@version", "@timestamp", "trade_date"]
  }
}
output {
    stdout { codec => dots }
    elasticsearch {
        index => "null_problem_fingerprint"
        document_id => "%{[@metadata][generated_id]}"
        doc_as_upsert => "true"
        action => "update"
    }
}

```

I deleted the fields with null values with the code blog in ruby, but after making the fingerprint, I still could not reach the desired output. Please help me!

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 12, 2021, 2:36pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/15 "2021-07-12T14:36:21Z")

</div>

> [@Badger](#):
>
> `if [email] == "NULL" { mutate { remove_field => ["email"] } }`

Actually the csv file is exactly like this:

```auto
My csv file =>
name,surname,age,email,phone
Busra,Duygu,99,,05555555555
Busra,Duygu,,busraduygu@gmail.com,
Busra,Duygu,99,,
Busra,Duygu,,,

```

I wrote null for better understanding

---

<div class="post-metadata">

**Author:** ![stephenb](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/stephenb/32/40856_2.png) [@stephenb](https://discuss.elastic.co/u/stephenb)\
**Post date:** [July 12, 2021, 3:10pm UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/16 "2021-07-12T15:10:43Z")

</div>

This works for me.

The problem is you are trying to use all the fields for the `fingerprint`... but all the fields do not exists on every row, so it does not make sense to try to use all the columns for fingerprint, you should only use the ones available on every row. The only fields that exist on every row are `name` and `surname`. So you might have a collision if users have the same are `name` and `surname`.

```auto
input {
    file { 
        path => "/Users/sbrown/workspace/sample-data/discuss/fingerprint.csv"
        start_position => "beginning"
        sincedb_path => "/dev/null" 
    }
}

filter {
    csv {        
        autodetect_column_names => "true"
        separator => ","
        skip_header => "true"
        columns => ["name","surname","age","email","phone"]
    }
    mutate { 
        remove_field =>["path", "host", "message", "@version", "@timestamp", "trade_date"]
    }

    if ![email] { mutate { remove_field => ["email"] } }
    if ![phone] { mutate { remove_field => ["phone"] } }
    if ![age] { mutate { remove_field => ["age"] } }

    fingerprint {
      method => "SHA1"
      source => ["name","surname"]
      target => "fingerprint"
      concatenate_sources => "true"   
  }

}

output {
    elasticsearch { 
        hosts => "http://localhost:9200"
        index => "null_problem"
        document_type => "_doc"
        document_id => "%{fingerprint}"
        action => 'update'
        doc_as_upsert => true
    }
  stdout {codec => rubydebug}

}

```

```auto

GET null_problem/_search

{
  "took" : 0,
  "timed_out" : false,
  "_shards" : {
    "total" : 1,
    "successful" : 1,
    "skipped" : 0,
    "failed" : 0
  },
  "hits" : {
    "total" : {
      "value" : 1,
      "relation" : "eq"
    },
    "max_score" : 1.0,
    "hits" : [
      {
        "_index" : "null_problem",
        "_type" : "_doc",
        "_id" : "fda96a9008c6ef62e4cf346636cf97e57112519d",
        "_score" : 1.0,
        "_source" : {
          "surname" : "Duygu",
          "fingerprint" : "fda96a9008c6ef62e4cf346636cf97e57112519d",
          "name" : "Busra",
          "age" : "99",
          "phone" : "05555555555",
          "email" : "busraduygu@gmail.com"
        }
      }
    ]
  }
}

```

---

<div class="post-metadata">

**Author:** ![Busra\_Duygu](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/busra_duygu/32/90462_2.png) [@Busra\_Duygu](https://discuss.elastic.co/u/Busra_Duygu)\
**Post date:** [July 13, 2021, 10:33am UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/17 "2021-07-13T10:33:20Z")

</div>

@stephenb Thank you very much for your all help 🤗🤗🤗 , it worked for me too 👌

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 10, 2021, 10:33am UTC](https://discuss.elastic.co/t/how-to-prevent-duplicate-and-has-null-value-documents-with-fingerprint/278150/18 "2021-08-10T10:33:49Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
