# How to handle case when the INDEX value is not included in the "path" value?

**URL:** <https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510>\
**Category:** Elasticsearch\
**Created:** [December 9, 2016, 4:56am UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510 "2016-12-09T04:56:07Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![chunji08](https://avatars.discourse-cdn.com/v4/letter/c/ee7513/32.png) [@chunji08](https://discuss.elastic.co/u/chunji08)\
**Post date:** [December 9, 2016, 4:56am UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/1 "2016-12-09T04:56:07Z")

</div>

Hi there,  
I am working on a Elastic project that generally of this model:

" Several LogStash Shippers that resides on respective hosts \<==\> one single Redis Server \<==\> multiple LogStash Indexer \<==\> one ElasticSearch Server  
".

As we understand, every piece of log message should come with an "index string", where this message would be registered of this index at ES side. But what happen if such index value is not included in the "path" value. For every Shipper host, this very unique value is within a property file.

How can we pass this value with every line of log message to the ES server ?

Thanks a lot.

Chun

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [December 12, 2016, 4:49am UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/2 "2016-12-12T04:49:58Z")

</div>

> [@chunji08](#):
>
> would be registered of this index at ES side. But what happen if such index value is not included in the "path" value. For every Shipper host, this very unique value is within a property file.

What do you mean with this?

---

<div class="post-metadata">

**Author:** ![chunji08](https://avatars.discourse-cdn.com/v4/letter/c/ee7513/32.png) [@chunji08](https://discuss.elastic.co/u/chunji08)\
**Post date:** [December 12, 2016, 7:36pm UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/3 "2016-12-12T19:36:38Z")

</div>

Hi All,  
Our system is used by multiple users every day to submit different jobs. And when one job is being processed, we would like to have logsatsh to do a real-time analysis if some log message is raised. And if it does, we would have such log message indexed per the current job-id.

As you can see the sample I put at the buttom. This is type of message the "Indexer Logstash" is given, where the "path" does not carry the job-id value. The actual value is kept in a property file. I have followed this link,  
"

> [@How to read the file from linux machine in logstash](https://discuss.elastic.co/t/how-to-read-the-file-from-linux-machine-in-logstash/57691):
>
> For reading the file from window machine I am using the following input filter and able to read the file from remote machine and writing to elasticsearch similarly I want to read the file from linux machine. input { file { type =\> "appname" path =\> "\IP Address/path/csvresultfolder/\*.csv" start\_position =\> "beginning" sincedb\_path =\> "/dev/null" } } Can we refer the hostname specified in properties file in logstash file,Please let me know.for example properties file name is user.prop…

",  
to have the value set in my env, and have the shipper.conf read. And that is why you see the "tag" is giving the unique job-id (36293300) now.  
"  
{  
"path" =\> "/user/logs/work/TOP/SFO/aaa/bbb/logs/server.out",  
"message" =\> " my log messages here ..."  
"type" =\> "sjc",  
"tags" =\> [  
[0] "36293300"  
]  
}  
",

So overall, this problem has been solved and this news group is very helpful for us.

Thanks a lot for the help.

CJ

---

<div class="post-metadata">

**Author:** ![chunji08](https://avatars.discourse-cdn.com/v4/letter/c/ee7513/32.png) [@chunji08](https://discuss.elastic.co/u/chunji08)\
**Post date:** [December 20, 2016, 2:23am UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/4 "2016-12-20T02:23:52Z")

</div>

Hi All,  
I have 2nd question that is related to this. When each job is started by per user, an index per this job is created now. But other than that, I have seen a different index id involved.

here is the example of my index list when I run the index command: " curl -X GET "http://ES\_server:9200/\_cat/indices" | sort ",

And here is the output I got:  
"  
....  
yellow open job-36416492 5 1 379772 0 57.6mb 57.6mb  
yellow open job-36418546 5 1 183099 0 40.7mb 40.7mb  
yellow open job-36418556 5 1 145342 0 35.2mb 35.2mb  
yellow open job-%{jobid} 5 1 8842147 0 1.4gb 1.4gb  
yellow open .kibana 1 1 2 2 10.7kb 10.7kb  
".

I don't know why this entry "  
yellow open job-%{jobid} 5 1 8842147 0 1.4gb 1.4gb  
",  
And it seems to me this entry grows per every job submitted ?

Thanks for the help.

CJI

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 20, 2016, 6:37am UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/5 "2016-12-20T06:37:03Z")

</div>

That index is created for events that does not contain a jobid that can be used by logstash to create the index name. Creating very small indices per jobid is not recommended as each index and shard uses comes with a certain amount of overhead. Having a large number of very small shards in a cluster is very inefficient and this approach will scale badly.

---

<div class="post-metadata">

**Author:** ![chunji08](https://avatars.discourse-cdn.com/v4/letter/c/ee7513/32.png) [@chunji08](https://discuss.elastic.co/u/chunji08)\
**Post date:** [December 20, 2016, 8:14am UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/6 "2016-12-20T08:14:42Z")

</div>

Hi Christian,  
Thanks for the response.  
Here is our scenario, Each of our job contains hundreds of thousands of lines of log messages from different server locations. We would like to have job's logs indexed under one index, and later on we could provide some quick response per client's "job requirements".

Anyway, back to that strange "job-%{jobid} " index I have noticed. I have tried different ways to see how it can be solved, and here is my what I found so far.

In my logstash indexer config file, I have this chunk of setting.  
...  
filter {  
grok {  
match =\> {  
"message" =\> ["%{Customized\_PATCH:some\_pattern1\_here}",  
"%{Custoemized\_FAILURE:some\_pattern2\_here}"]  
"tags" =\> ["%{NUMBER:jobid}"]  
}  
...  
}  
".  
This tag value is created at shipper level and indexer to have it extracted for creating the unique index name.

What I have found is,

1. if the tags part is defined after message part, this extra "job-%{jobid}" index will be generated.
2. if the tags part is defined before message part, there is no more "job-%{jobid}", but customized pattern will not work. it was totally ignored even if there is a pattern match in the log message.

?? Any way to have both working.

Thanks a lot for the help.

Chun

---

<div class="post-metadata">

**Author:** ![chunji08](https://avatars.discourse-cdn.com/v4/letter/c/ee7513/32.png) [@chunji08](https://discuss.elastic.co/u/chunji08)\
**Post date:** [December 20, 2016, 10:19pm UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/7 "2016-12-20T22:19:24Z")

</div>

I have found the answer to this question, that I have to keep them in separate grok and both works fine. Something as:  
"  
filter {  
grok {  
match =\> {  
"message" =\> ["%{Customized\_PATCH:some\_pattern1\_here}",  
"%{Custoemized\_FAILURE:some\_pattern2\_here}"]  
}  
} // end of 1st grok  
grok {  
"tags" =\> ["%{NUMBER:jobid}"]  
} // end of 2nd grok  
} // end of filter.  
"

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 20, 2016, 10:23pm UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/8 "2016-12-20T22:23:11Z")

</div>

How many jobs do you expect to have indexed in the cluster at any point in time?

---

<div class="post-metadata">

**Author:** ![chunji08](https://avatars.discourse-cdn.com/v4/letter/c/ee7513/32.png) [@chunji08](https://discuss.elastic.co/u/chunji08)\
**Post date:** [December 21, 2016, 6:06pm UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/9 "2016-12-21T18:06:44Z")

</div>

It is still in the trial stage. If everything was stable, we are expecting 100 to 150 jobs being indexed every 24 hour. We may also need to have an index de-list logic, as we are only interested of those, that are created for the past 3-4 days. In our current design, every shipper is started in parallel with the running job on its own host.

---

<div class="post-metadata">

**Author:** ![Christian\_Dahlqvist](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/christian_dahlqvist/32/4617_2.png) [@Christian\_Dahlqvist](https://discuss.elastic.co/u/Christian_Dahlqvist)\
**Post date:** [December 21, 2016, 6:15pm UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/10 "2016-12-21T18:15:21Z")

</div>

That will leave you with approximately 500-600 indices in the cluster if I understand you correctly. That may, depending on the size of the nodes and cluster, be OK, but as your indices are so small I would recommend configuring each of them to have 1 shard instead of the default 5.

As said earlier, if you intend to increase this going forward this approach will not scale terribly well.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 18, 2017, 6:15pm UTC](https://discuss.elastic.co/t/how-to-handle-case-when-the-index-value-is-not-included-in-the-path-value/68510/11 "2017-01-18T18:15:39Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
