# ElasticSearch Indexing question

**URL:** <https://discuss.elastic.co/t/elasticsearch-indexing-question/35845>\
**Category:** Elasticsearch\
**Created:** [November 29, 2015, 8:32pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845 "2015-11-29T20:32:00Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [November 29, 2015, 8:32pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/1 "2015-11-29T20:32:01Z")

</div>

Hi,

Is there an easy way to index a lot of documents into ES db? i was thinking or using a CMS of sorts that is hooked up to ES.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [November 30, 2015, 9:35pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/2 "2015-11-30T21:35:13Z")

</div>

There are a number of ways. Many people use Logstash to do this.

What sort of data is it? Where is it held?

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 1, 2015, 12:18am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/3 "2015-12-01T00:18:50Z")

</div>

It's a shared folder with many .doc .docx files in various subfolders.  
Isn't logstash just for logs etc?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 1, 2015, 12:24am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/4 "2015-12-01T00:24:25Z")

</div>

Logstash can be used for many things, but not that sort of data.

You could use the mapper attachments plugin - [https://github.com/elastic/elasticsearch-mapper-attachments](https://github.com/elastic/elasticsearch-mapper-attachments) - for this.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 1, 2015, 2:52am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/5 "2015-12-01T02:52:03Z")

</div>

You can give a try to fscrawler. [https://github.com/dadoonet/fscrawler](https://github.com/dadoonet/fscrawler)

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 1, 2015, 8:44pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/6 "2015-12-01T20:44:39Z")

</div>

Thanks!  
I actually found both of these using google but my concerns are the following:

1. mapper-attachments is a plugin that basically converts the documents to text version and allows you to index them if i got it right but i don't see how it automates the process.

2. fscrawler now states it's standalone and not supported by Es so does the results integrate with ES later?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 1, 2015, 9:13pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/7 "2015-12-01T21:13:24Z")

</div>

Fscrawler is not a plugin anymore. It's a standalone app which sends data to elasticsearch.

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 1, 2015, 10:53pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/8 "2015-12-01T22:53:11Z")

</div>

My question was regarding the big warning " Elasticsearch 2.0.0 doesn't support anymore rivers." since rivers were a way for ES to get data from external sources i was asking how FScrawler handles this in 2.0 since it's supported anymore, i mean are the results still compatible and acceptable by ES db?

I didn't notice you are the owner of FScrawler 😊  
I guess it should still workl, i will test it out and report back.  
Right now i will test it on windows, later maybe i"l have a linux machine.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 6:46am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/9 "2015-12-02T06:46:27Z")

</div>

> [@Yorko](#):
>
> Are the results still compatible and acceptable by ES db?

Yes. Fscrawler has been rewritten FOR elasticsearch 2.0. It might work for previous versions as well but untested.

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 13, 2015, 9:53pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/10 "2015-12-13T21:53:26Z")

</div>

Continuing the discussion from [ElasticSearch Indexing question](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/9):

> [@dadoonet](#):
>
> > [@Yorko](#):
> >
> > Are the results still compatible and acceptable by ES db?
> 
> Yes. Fscrawler has been rewritten FOR elasticsearch 2.0. It might work for previous versions as well but untested.

OK i've tried it and have 3 questions:

1. Is there any way to not include the \r\n whitespaces in the \_source, in my docs and docx there are different parts separated by spaces, i would prefer the new lines to be in but not printed...
2. Is there a way to make it recycle memory or will java just keep eating all the memory until there is no more or the scan finishes?
3. I've tried hooking the results into Kibana and i didn't get any results in the discover tab, here is the template i used:

```auto
{
  "name" : "test2",
  "fs" : {
    "url" : "C:/ABP",
    "update_rate" : "15m",
    "includes" : null,
    "excludes" : null,
    "json_support" : false,
    "filename_as_id" : false,
    "add_filesize" : true,
    "remove_deleted" : true,
    "store_source" : false,
    "index_content" : true,
    "indexed_chars" : null
  },
  "server" : null,
  "elasticsearch" : {
    "nodes" : [ {
      "host" : "127.0.0.1",
      "port" : 9200
    } ],
    "index" : "test2",
    "type" : "doc",
    "bulk_size" : 100,
    "flush_interval" : "5s"
  }
}
```

I've used test2 and test2\* as the index name and select date modified (which exists in ES) as the field and still got nothing in the discovery, plus it only shows \_source as the field it searches.

Any ideas?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 15, 2015, 8:36am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/11 "2015-12-15T08:36:46Z")

</div>

> [@Yorko](#):
>
> Is there any way to not include the \r\n whitespaces in the \_source, in my docs and docx there are different parts separated by spaces, i would prefer the new lines to be in but not printed...

No. It's indexed as it is extracted by Tika. But TBH I did not understand what is the problem. May be illustrate with an example what you have now and what you would like to see?

> [@Yorko](#):
>
> Is there a way to make it recycle memory or will java just keep eating all the memory until there is no more or the scan finishes?

May be some enhancements need to be done in fscrawler project. For sure I should support adding easily memory settings to the fscrawler job. For now, you have to hack the script or set `$JAVA_OPTS`.  
I opened [Add FS\_JAVA\_OPTS JVM option · Issue #134 · dadoonet/fscrawler · GitHub](https://github.com/dadoonet/fscrawler/issues/134) for this. Feel free to contribute! 😛

> [@Yorko](#):
>
> I've tried hooking the results into Kibana and i didn't get any results in the discover tab, here is the template i used:

I never tested it with Kibana for now. I'd advice that you first test with simple curl commands that everything has been indexed as expected. Is it the case?

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 16, 2015, 12:01am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/12 "2015-12-16T00:01:47Z")

</div>

Continuing the discussion from [ElasticSearch Indexing question](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/11):

> [@dadoonet](#):
>
> > [@Yorko](#):
> >
> > Is there any way to not include the \r\n whitespaces in the \_source, in my docs and docx there are different parts separated by spaces, i would prefer the new lines to be in but not printed...
> 
> No. It's indexed as it is extracted by Tika. But TBH I did not understand what is the problem. May be illustrate with an example what you have now and what you would like to see?

I have a file that goes:  
Header  
Body  
Footer

So the result is header\nbodynfooter\n so when i view it i want to see it as the original (separated) and not just like one long string.

> [@dadoonet](#):
>
> > [@Yorko](#):
> >
> > Is there a way to make it recycle memory or will java just keep eating all the memory until there is no more or the scan finishes?
> 
> May be some enhancements need to be done in fscrawler project. For sure I should support adding easily memory settings to the fscrawler job. For now, you have to hack the script or set `$JAVA_OPTS`.  
> I opened [Add FS\_JAVA\_OPTS JVM option · Issue #134 · dadoonet/fscrawler · GitHub](https://github.com/dadoonet/fscrawler/issues/134) for this. Feel free to contribute! 😛

Ok nice, i will defenalty try to help 😉

> [@dadoonet](#):
>
> > [@Yorko](#):
> >
> > I've tried hooking the results into Kibana and i didn't get any results in the discover tab, here is the template i used:
> 
> I never tested it with Kibana for now. I'd advice that you first test with simple curl commands that everything has been indexed as expected. Is it the case?

The thing is that at the end i need a dashboard and kibana is an easy choice here so i must have it working, i appreciate the concern and you are correct first try it simple but i also need it to work and if you haven't tested it yet i will gladly volunteer here 🙂

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 18, 2015, 11:34am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/13 "2015-12-18T11:34:32Z")

</div>

> [@Yorko](#):
>
> I have a file that goes:  
> Header  
> Body  
> Footer

could you share somewhere your binary document so I could try some tests on my side?

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 19, 2015, 11:47pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/14 "2015-12-19T23:47:50Z")

</div>

> [@dadoonet](#):
>
> > [@Yorko](#):
> >
> > I have a file that goes:  
> > Header  
> > Body  
> > Footer
> 
> could you share somewhere your binary document so I could try some tests on my side?

First of all it works fine with Kibana i just had the wrong time settings so FYI on that.

Secondly sure here is an example:

```auto
Defect subject: XXXX 
Product: XXXX vX.X 
Severity: XXX 

Description:
First paragraph: A short explanation of the issue. 

Technical Details:
Technical details about how the product was tested. 
1. Example: 
Figure 
2. Example: 
Figure 
 
Recommended Remediation:
Recommendation 
```

Thanks for you help!

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 20, 2015, 12:08am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/15 "2015-12-20T00:08:12Z")

</div>

Is it a TXT file?

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 20, 2015, 7:22am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/16 "2015-12-20T07:22:35Z")

</div>

> [@dadoonet](#):
>
> Is it a TXT file?

MS Word usually .doc or .docx

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 20, 2015, 7:51am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/17 "2015-12-20T07:51:51Z")

</div>

could you share somewhere your binary document?

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 22, 2015, 12:00am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/18 "2015-12-22T00:00:06Z")

</div>

> [@dadoonet](#):
>
> could you share somewhere your binary document?

Here you go, this is just a template but it's the same format just missing real text and images.  
[Template Link](http://www84.zippyshare.com/v/2yRche97/file.html)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 29, 2015, 10:12am UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/19 "2015-12-29T10:12:19Z")

</div>

So I extracted your file with fscrawler and got: `Defect subject: XXXX\nProduct: XXXX vX.X\nSeverity: XXX\nDescription\nFirst paragraph: A short explanation of the issue.\nTechnical Details\nTechnical details about how the product was tested.\nExample:\nFigure\n1. Example:\nFigure\n\nRecommended Remediation\n1. Recommendation\n1. Recommendation\n1. Recommendation\n\n`

I was then able to search for `figure` for example without any issue.

Is there anything wrong with that then?

---

<div class="post-metadata">

**Author:** ![Yorko](https://avatars.discourse-cdn.com/v4/letter/y/58956e/32.png) [@Yorko](https://discuss.elastic.co/u/Yorko)\
**Post date:** [December 29, 2015, 10:05pm UTC](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845/20 "2015-12-29T22:05:30Z")

</div>

> [@dadoonet](#):
>
> So I extracted your file with fscrawler and got: `Defect subject: XXXX\nProduct: XXXX vX.X\nSeverity: XXX\nDescription\nFirst paragraph: A short explanation of the issue.\nTechnical Details\nTechnical details about how the product was tested.\nExample:\nFigure\n1. Example:\nFigure\n\nRecommended Remediation\n1. Recommendation\n1. Recommendation\n1. Recommendation\n\n`
> 
> I was then able to search for `figure` for example without any issue.
> 
> Is there anything wrong with that then?

The problem is with the format, it's not human readable so you can search for words and find them in the whole mess of a text as you showed but if you are only interested in reading a particular section it's hard to find quick where one beings and another ends...

It would be much simpler if the \n characters weren't represented as strings.

[Next page](https://discuss.elastic.co/t/elasticsearch-indexing-question/35845.md?page=2)
