# How To: Back-fill Elasticsearch without losing data

**URL:** <https://discuss.elastic.co/t/how-to-back-fill-elasticsearch-without-losing-data/29793>\
**Category:** Logstash\
**Created:** [September 22, 2015, 3:42pm UTC](https://discuss.elastic.co/t/how-to-back-fill-elasticsearch-without-losing-data/29793 "2015-09-22T15:42:48Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![cpattonj](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/cpattonj/32/4365_2.png) [@cpattonj](https://discuss.elastic.co/u/cpattonj)\
**Post date:** [September 22, 2015, 3:42pm UTC](https://discuss.elastic.co/t/how-to-back-fill-elasticsearch-without-losing-data/29793/1 "2015-09-22T15:42:48Z")

</div>

I want to import from my existing log file as a new file input like so to back-fill my old data:

```
input {
  path => ['/path_local_to_es/file']
  start_position => 'beginning'
}

```

AND

Continue to capture the tail of the very same file as it's actively logging through a redis messaging queue...

**On Data Source** :

```
input {
    path => ['/path_local_to_source/file']
    start_position => 'end'
}
filter {
    ...
}
output {
    redis {
         host => 'queue-to-elasticsearch'
         data_type => 'list'
         key => 'logstash'
    }
}

```

**On Elasticsearch**

```
input {
    redis {
        host => '127.0.0.1'
        data_type => 'list'
        key => 'logstash'
    }
}
output {
    elasticsearch {
        host => '127.0.0.1'
    }
}

```

How do I combine the two in a way that lets me back-fill all of the data I want from the file but not duplicate data that's already being ingested by the redis queue? I'm ok with stopping the queue for a moment but even then, how would I prevent missing data between when my local copy of the log file reaches EOF and the redis queue is turned back on?

Thanks.

---

<div class="post-metadata">

**Author:** ![jsvd](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jsvd/32/6203_2.png) [@jsvd](https://discuss.elastic.co/u/jsvd)\
**Post date:** [September 22, 2015, 5:41pm UTC](https://discuss.elastic.co/t/how-to-back-fill-elasticsearch-without-losing-data/29793/2 "2015-09-22T17:41:45Z")

</div>

The easiest way to support frequent backfilling is by controlling the document ids so that each document as a unique id before going to elasticsearch, either extracted from the source or computed using the document data itself (if it is unique enough).

If each document has a unique document id, it can be passed into the elasticsearch output like this: `document_id => "%{[unique_id_field]}"`

By doing this a single document can be index more than once to elasticsearch because it will be overwritten.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 5:28am UTC](https://discuss.elastic.co/t/how-to-back-fill-elasticsearch-without-losing-data/29793/3 "2017-07-06T05:28:19Z")

</div>


