# Correct way to handle faceted search tokenization issue for dynamic fields?

**URL:** <https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815>\
**Category:** Elasticsearch\
**Created:** [October 1, 2013, 10:26am UTC](https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815 "2013-10-01T10:26:01Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Onur\_Aktas](https://avatars.discourse-cdn.com/v4/letter/o/ac8455/32.png) [@Onur\_Aktas](https://discuss.elastic.co/u/Onur_Aktas)\
**Post date:** [October 1, 2013, 10:26am UTC](https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815/1 "2013-10-01T10:26:01Z")

</div>

Hi all,

I want to index an object which has some static fields and dynamic fields  
(kept in HashMap) holding Product technical feature name and value pair.You  
can guess that there are thousands of different product types from various  
product categories which cause 1000s of different technical  
features/attributes.

Each technical feature value starts with f\_ so I applied a mapping  
something like below.

- dynamic\_templates: [
  - {
    - template\_feature: {
      - mapping: {
        - type: multi\_field
        - fields: {
          - {name}: {
            - type: {dynamic\_type}
            - index: analyzed  
}

          - org: {
            - type: {dynamic\_type}
            - index: not\_analyzed  
}  
}  
}

      - match: f\_\*  
}  
}  
]

However, when I check mapping I see that ElasticSearch creates a mapping  
for each inserted technical feature. So it means when product list grows;  
mapping will significantly grow and there are about total 2000 different  
technical feature value.

- f\_material: {
  - type: multi\_field
  - fields: {
    - f\_material: {
      - type: string  
}

    - org: {
      - type: string
      - index: not\_analyzed
      - omit\_norms: true
      - index\_options: docs
      - include\_in\_all: false  
}  
}  
}

- f\_period\_type: {
  - type: multi\_field
  - fields: {
    - f\_period\_type: {
      - type: string  
}

    - org: {
      - type: string
      - index: not\_analyzed
      - omit\_norms: true
      - index\_options: docs
      - include\_in\_all: false  
}  
}  
}

- f\_production\_type: {
  - type: multi\_field
  - fields: {
    - f\_production\_type: {
      - type: string  
}

    - org: {
      - type: string
      - index: not\_analyzed
      - omit\_norms: true
      - index\_options: docs
      - include\_in\_all: false  
}  
}  
}

- f\_size: {
  - type: multi\_field
  - fields: {
    - f\_size: {
      - type: string  
}

    - org: {
      - type: string
      - index: not\_analyzed
      - omit\_norms: true
      - index\_options: docs
      - include\_in\_all: false  
}  
}  
}

- f\_style: {
  - type: multi\_field
  - fields: {
    - f\_style: {
      - type: string  
}

    - org: {
      - type: string
      - index: not\_analyzed
      - omit\_norms: true
      - index\_options: docs
      - include\_in\_all: false  
}  
}  
}

I applied this mapping just to perform a faceted search without whitespace  
tokenization; but I think it is not the correct way.

Could you please advice the correct way to do this?

KR,  
Onur

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Tchinkatchuk](https://avatars.discourse-cdn.com/v4/letter/t/958977/32.png) [@Tchinkatchuk](https://discuss.elastic.co/u/Tchinkatchuk)\
**Post date:** [November 21, 2013, 10:43am UTC](https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815/2 "2013-11-21T10:43:31Z")

</div>

Hi,

I have exactly the same way to do it and the sema issue.  
COuld you help me if you foudn the answer ?

Thanks.

---

<div class="post-metadata">

**Author:** ![Onur\_Aktas](https://avatars.discourse-cdn.com/v4/letter/o/ac8455/32.png) [@Onur\_Aktas](https://discuss.elastic.co/u/Onur_Aktas)\
**Post date:** [November 21, 2013, 10:58am UTC](https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815/3 "2013-11-21T10:58:18Z")

</div>

On Tuesday, October 1, 2013 1:26:01 PM UTC+3, Onur Aktaş wrote:

> Hi all,
> 
> I want to index an object which has some static fields and dynamic fields  
> (kept in HashMap) holding Product technical feature name and value pair.You  
> can guess that there are thousands of different product types from various  
> product categories which cause 1000s of different technical  
> features/attributes.
> 
> Each technical feature value starts with f\_ so I applied a mapping  
> something like below.
> 
> - dynamic\_templates: [
> - {
> - template\_feature: {
> - mapping: {
> - type: multi\_field
> - fields: {
> - {name}: {
> - type: {dynamic\_type}
> - index: analyzed  
> }
> 
> - org: {
> - type: {dynamic\_type}
> - index: not\_analyzed  
> }  
> }  
> }
> 
> - match: f\_\*  
> }  
> }  
> ]
> 
> However, when I check mapping I see that Elasticsearch creates a mapping  
> for each inserted technical feature. So it means when product list grows;  
> mapping will significantly grow and there are about total 2000 different  
> technical feature value.
> 
> - f\_material: {
> - type: multi\_field
> - fields: {
> - f\_material: {
> - type: string  
> }
> 
> - org: {
> - type: string
> - index: not\_analyzed
> - omit\_norms: true
> - index\_options: docs
> - include\_in\_all: false  
> }  
> }  
> }
> 
> - f\_period\_type: {
> - type: multi\_field
> - fields: {
> - f\_period\_type: {
> - type: string  
> }
> 
> - org: {
> - type: string
> - index: not\_analyzed
> - omit\_norms: true
> - index\_options: docs
> - include\_in\_all: false  
> }  
> }  
> }
> 
> - f\_production\_type: {
> - type: multi\_field
> - fields: {
> - f\_production\_type: {
> - type: string  
> }
> 
> - org: {
> - type: string
> - index: not\_analyzed
> - omit\_norms: true
> - index\_options: docs
> - include\_in\_all: false  
> }  
> }  
> }
> 
> - f\_size: {
> - type: multi\_field
> - fields: {
> - f\_size: {
> - type: string  
> }
> 
> - org: {
> - type: string
> - index: not\_analyzed
> - omit\_norms: true
> - index\_options: docs
> - include\_in\_all: false  
> }  
> }  
> }
> 
> - f\_style: {
> - type: multi\_field
> - fields: {
> - f\_style: {
> - type: string  
> }
> 
> - org: {
> - type: string
> - index: not\_analyzed
> - omit\_norms: true
> - index\_options: docs
> - include\_in\_all: false  
> }  
> }  
> }
> 
> I applied this mapping just to perform a faceted search without whitespace  
> tokenization; but I think it is not the correct way.
> 
> Could you please advice the correct way to do this?
> 
> KR,  
> Onur

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Onur\_Aktas](https://avatars.discourse-cdn.com/v4/letter/o/ac8455/32.png) [@Onur\_Aktas](https://discuss.elastic.co/u/Onur_Aktas)\
**Post date:** [November 21, 2013, 11:01am UTC](https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815/4 "2013-11-21T11:01:35Z")

</div>

Hi Georges,

We decided to give a name to each feature and use it instead of it's title.  
name -\> title where name is feature + (order of the feature)

For example, lets say Category A and B has following technical feature  
fields.

Category A: Material, Size  
Category B: Color, Size

Then we mapped Category A as:  
p01 -\> Material,  
p02 -\> Size

Category B as:  
p01-\> Color  
p02 -\> Size

Then we assumed any category can have max 20 feature values and then mapped  
each feature by its name instead of its title.

Finally we had a mapping something like this:

```
     "p01":{
        "type":"multi_field",
        "fields":{
           "analyzed":{
              "type":"string",
              "index":"analyzed"
           },
           "notanalyzed":{
              "type":"string",
              "index":"not_analyzed"
           }
        }
     },
     "p02":{
        "type":"multi_field",
        "fields":{
           "analyzed":{
              "type":"string",
              "index":"analyzed"
           },
           "notanalyzed":{
              "type":"string",
              "index":"not_analyzed"
           }
        }
     }

```

.. goes up to p20

So products will have a data something like:  
Product A  
p01 -\> Steel  
p02 -\> 15 meters.

_Pros_  
You do not have to create (category count \* unique feature name) mappings.

_Cons_  
You should not rename feature's name; otherwise products will show wrong  
data.

Hope it helps.

KR,  
Onur

On Thursday, November 21, 2013 12:43:32 PM UTC+2, Georges@Bibtol wrote:

> Hi,
> 
> I have exactly the same way to do it and the sema issue.  
> COuld you help me if you foudn the answer ?
> 
> Thanks.
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields-tp4041972p4044698.html](http://elasticsearch-users.115913.n3.nabble.com/Correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields-tp4041972p4044698.html)
> 
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

On Thursday, November 21, 2013 12:43:32 PM UTC+2, Georges@Bibtol wrote:

> Hi,
> 
> I have exactly the same way to do it and the sema issue.  
> COuld you help me if you foudn the answer ?
> 
> Thanks.
> 
> --  
> View this message in context:  
> [http://elasticsearch-users.115913.n3.nabble.com/Correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields-tp4041972p4044698.html](http://elasticsearch-users.115913.n3.nabble.com/Correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields-tp4041972p4044698.html)  
> Sent from the Elasticsearch Users mailing list archive at [Nabble.com](http://Nabble.com).

--  
You received this message because you are subscribed to the Google Groups "elasticsearch" group.  
To unsubscribe from this group and stop receiving emails from it, send an email to [elasticsearch+unsubscribe@googlegroups.com](mailto:elasticsearch+unsubscribe@googlegroups.com).  
For more options, visit [https://groups.google.com/groups/opt\_out](https://groups.google.com/groups/opt_out).

---

<div class="post-metadata">

**Author:** ![Tchinkatchuk](https://avatars.discourse-cdn.com/v4/letter/t/958977/32.png) [@Tchinkatchuk](https://discuss.elastic.co/u/Tchinkatchuk)\
**Post date:** [November 27, 2013, 4:54pm UTC](https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815/5 "2013-11-27T16:54:38Z")

</div>

Thanks for the answer.  
Unfortunately, I do not want to rename all my dynamic attributes.

here's a little mapping configuration I have :

{ "article": { "\_default\_": { "dynamic\_templates": [{ "base": { "match": "\*", "mapping": { "type" : "multi\_field", "fields" : { "{name}" : { "type" : "string", "index" : "analyzed", "store" : "yes", "analyzer" : "my\_string\_analyzer", "search\_analyzer" : "default", "index\_analyzer" : "default\_edge\_n\_grams" }, "raw\_value": {"type": "string", "analyzer": "not\_analyzed"} } } } }] } } }

I want to be able to get facets this way :  
  
GET \_search  
{  
"facets": {  
"brand": {  
"terms": {  
"field" : "brand.raw\_value"  
}  
}  
},  
"query": {  
"filtered" : {  
"query" : {  
"query\_string" : {  
"query" : "\*\*\*"  
}  
}  
}  
}  
}

cause if i do it on brand and not brand.raw\_value, my bran value is tokenized.  
-\> "Elastic Search" wille render 2 facets possibilities "Elastic" & "Value" instead of just one.

Such a shame.  
Do I miss something ?

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:04am UTC](https://discuss.elastic.co/t/correct-way-to-handle-faceted-search-tokenization-issue-for-dynamic-fields/13815/6 "2017-07-06T02:04:30Z")

</div>


