# Analyzer for web page name

**URL:** <https://discuss.elastic.co/t/analyzer-for-web-page-name/68209>\
**Category:** Elasticsearch\
**Created:** [December 6, 2016, 7:14pm UTC](https://discuss.elastic.co/t/analyzer-for-web-page-name/68209 "2016-12-06T19:14:48Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![liorg2](https://avatars.discourse-cdn.com/v4/letter/l/ed8c4c/32.png) [@liorg2](https://discuss.elastic.co/u/liorg2)\
**Post date:** [December 6, 2016, 7:14pm UTC](https://discuss.elastic.co/t/analyzer-for-web-page-name/68209/1 "2016-12-06T19:14:48Z")

</div>

Is there a way to apply the following code to a web page field that i have?  
I have it in many types, so iseally, if i can add analyer, and reindex the data, that will be great..  
Here is a node.js code( language is not important) , that includes the rules i need:

```
exports.cleanPage = function (page) {

if (!page || page.length ==0)
    return page;

if (page == "/") {
    return page;
}

if (!page.startsWith("/")) { // a page must be started with "/"
    page = "/" + page;
}

if (page.endsWith('/')) { // we want to prevent duplication such as en,/en/,en/
    page = page.slice(0, -1);
}

if (page.indexOf("?") > -1) {
    page = page.split("?")[0]; //remove query string from page name
}

return page.toLowerCase();

```

};

I'm using 1.7

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [December 7, 2016, 1:25pm UTC](https://discuss.elastic.co/t/analyzer-for-web-page-name/68209/2 "2016-12-07T13:25:19Z")

</div>

Hey,

there is no direct component to remove query parameters and normalize parts of URLs. There is the [url uax email tokenizer](https://www.elastic.co/guide/en/elasticsearch/reference/5.0/analysis-uaxurlemail-tokenizer.html), that leaves URLs as a single piece. You could however use the [pattern replace tokenfilter](https://www.elastic.co/guide/en/elasticsearch/reference/5.0/analysis-pattern_replace-tokenfilter.html) to use regular expression to trim down your path.

I'd still recommend doing that before indexing data - this might make things much more explanatory than complex regular expressions.

With 5.0 you could take a look at the new ingest node feature assist you with this and change the field to your needs before indexing.

--Alex

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [January 4, 2017, 1:25pm UTC](https://discuss.elastic.co/t/analyzer-for-web-page-name/68209/3 "2017-01-04T13:25:22Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
