# A working example of how to import a Wikipedia dump with Logstash using es\_bulk

**URL:** <https://discuss.elastic.co/t/a-working-example-of-how-to-import-a-wikipedia-dump-with-logstash-using-es-bulk/253228>\
**Category:** Logstash\
**Created:** [October 25, 2020, 9:46am UTC](https://discuss.elastic.co/t/a-working-example-of-how-to-import-a-wikipedia-dump-with-logstash-using-es-bulk/253228 "2020-10-25T09:46:19Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![Ola\_Gustafsson1](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ola_gustafsson1/32/77331_2.png) [@Ola\_Gustafsson1](https://discuss.elastic.co/u/Ola_Gustafsson1)\
**Post date:** [October 25, 2020, 9:46am UTC](https://discuss.elastic.co/t/a-working-example-of-how-to-import-a-wikipedia-dump-with-logstash-using-es-bulk/253228/1 "2020-10-25T09:46:19Z")

</div>

I'm writing a logstash configuration file for importing a Wikipedia dump, found on [https://dumps.wikimedia.org/other/cirrussearch/current/](https://dumps.wikimedia.org/other/cirrussearch/current/)

The dumps are in the es\_bulk format, ie one line for the action and id of the document and then a line containing the actual JSON data.

I'm changing codecs to make this work and the JSON codec inputs each line as a document. The es\_bulk codec causes a crash and I can't for the life of me understand how a multiline statement would look to make it work.

This is my conf right now:

```auto
input {
    file {
        path => "/home/projects/wiki-load/swwikibooks-20201019-cirrussearch-general.json.gz"
        mode => "read"
        codec => "json"
        start_position => "beginning"
        file_completed_action => "log"
        file_completed_log_path => "/home/projects/wiki-load/log.txt"
    }
}
filter {
    json {
        source => "message"
    }
}
output {
    elasticsearch {
        hosts => ["localhost:9200"]
        index => "svwiki-20201012"
        document_type => "page"
    }
    stdout {
		codec => rubydebug { metadata => false }
	}
}

```

I need to capture the header line and the document line together, keeping the original structure and unique id of the document. Any suggestions here?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 22, 2020, 9:46am UTC](https://discuss.elastic.co/t/a-working-example-of-how-to-import-a-wikipedia-dump-with-logstash-using-es-bulk/253228/2 "2020-11-22T09:46:30Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
