# Identifying Significant Words In a Field

**URL:** <https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688>\
**Category:** Kibana\
**Created:** [March 27, 2018, 2:32am UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688 "2018-03-27T02:32:48Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [March 27, 2018, 2:32am UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/1 "2018-03-27T02:32:48Z")

</div>

Working with the Twitter plugin, there are a couple fields that contain the tweet. Is it possible to create a visual that can analyze the text in this field and then create a chart that shows what words are mentioned most?

---

<div class="post-metadata">

**Author:** ![tsullivan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsullivan/32/31077_2.png) [@tsullivan](https://discuss.elastic.co/u/tsullivan)\
**Post date:** [March 27, 2018, 7:42pm UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/2 "2018-03-27T19:42:37Z")

</div>

> Working with the Twitter plugin

I assume you mean the Twitter plugin for Logstash

Yes, this is very possible and I happen to have example that does this. The extra part needed is that the logstash index needs to have custom mappings to store analyzed text versions of the text fields you're interested in.

See: [GitHub - tsullivan/avocado-pipeline](https://github.com/tsullivan/avocado-pipeline/)

The custom mappings are specified in the `avocado-tweets-wildcard.json` file, referenced here: [https://github.com/tsullivan/avocado-pipeline/blob/master/tweet-pipeline.conf#L32](https://github.com/tsullivan/avocado-pipeline/blob/master/tweet-pipeline.conf#L32)

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [March 28, 2018, 5:33am UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/3 "2018-03-28T05:33:01Z")

</div>

Thanks for the info. The stopwords configured at the top, I assume that's where you set words that are not to be analyzed?

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [March 30, 2018, 5:25am UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/4 "2018-03-30T05:25:07Z")

</div>

Finally got around to working on this and it works perfectly. I have only had it running for a short time but is there a need to mutate the field before it's indexed to homogenize the letter case that's used or is Elasticsearch smart enough to know that `This = this = THIS`?

---

<div class="post-metadata">

**Author:** ![tsullivan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsullivan/32/31077_2.png) [@tsullivan](https://discuss.elastic.co/u/tsullivan)\
**Post date:** [March 30, 2018, 5:27pm UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/5 "2018-03-30T17:27:23Z")

</div>

Hm, you add a lowercase filter to the custom analyzer. The documentation on custom analyzers has an example of adding a lowercase filter:

[https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-custom-analyzer.html#\_example\_configuration\_5](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-custom-analyzer.html#_example_configuration_5)

Here's an example I came up with using the `my_stop_analyzer` from the pipeline I shared with you:

## Make an index with the custom analyzer in its settings:

```auto
PUT /cool_example
{
  "settings": {
    "analysis": {
      "analyzer": {
        "my_stop_analyzer": {
          "stopwords": [
            "a", "an", "and", "are", "as", "at", "be", "but", "by", "for", "have", "i", "if", "in", "into", "is", "it", "my", "no", "not", "of", "on", "or", "such", "that", "the", "their", "then", "there", "these", "they", "this", "to", "was", "will", "with", "what", "you", "https", "t", "co", "www", "http", "com"
          ],
          "type": "stop"
        }
      },
      "filter": {
        "lower": {
          "type": "lowercase"
        }
      }
    }
  }
}

```

## Run a test:

```auto
POST cool_example/_analyze
{
  "analyzer": "my_stop_analyzer",
  "text": "Be here on Sunday Sunday SUNDAY"
}

```

## Elasticsearch returns only lowercased non-stopword tokens:

```auto
{
  "tokens": [
    {
      "token": "here",
      "start_offset": 3,
      "end_offset": 7,
      "type": "word",
      "position": 1
    },
    {
      "token": "sunday",
      "start_offset": 11,
      "end_offset": 17,
      "type": "word",
      "position": 3
    },
    {
      "token": "sunday",
      "start_offset": 18,
      "end_offset": 24,
      "type": "word",
      "position": 4
    },
    {
      "token": "sunday",
      "start_offset": 25,
      "end_offset": 31,
      "type": "word",
      "position": 5
    }
  ]
}

```

---

<div class="post-metadata">

**Author:** ![tsullivan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsullivan/32/31077_2.png) [@tsullivan](https://discuss.elastic.co/u/tsullivan)\
**Post date:** [March 30, 2018, 5:28pm UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/6 "2018-03-30T17:28:31Z")

</div>

> [@wwalker](#):
>
> The stopwords configured at the top, I assume that's where you set words that are not to be analyzed?

Sorry, I forgot to reply to this. Yes, the configured stopwords tell Elasticsearch what to drop from analysis: [Stop analyzer | Elasticsearch Guide [8.11] | Elastic](https://www.elastic.co/guide/en/elasticsearch/reference/current/analysis-stop-analyzer.html)

---

<div class="post-metadata">

**Author:** ![wwalker](https://avatars.discourse-cdn.com/v4/letter/w/43a26b/32.png) [@wwalker](https://discuss.elastic.co/u/wwalker)\
**Post date:** [March 31, 2018, 2:50am UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/7 "2018-03-31T02:50:21Z")

</div>

huh....I thought Elasticsearch was just the index and search engine for the data but it looks like it can also do data manipulation as well. So, when Logstash and Elasticsearch can both manipulate the data, in this case the lowercase function, which one is best to use/more efficient, Logstash or Elasticsearch?

---

<div class="post-metadata">

**Author:** ![tsullivan](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tsullivan/32/31077_2.png) [@tsullivan](https://discuss.elastic.co/u/tsullivan)\
**Post date:** [April 3, 2018, 6:04pm UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/8 "2018-04-03T18:04:54Z")

</div>

Personally, I'm not really sure! I'm really more of a Kibana question-answerer.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [May 1, 2018, 6:05pm UTC](https://discuss.elastic.co/t/identifying-significant-words-in-a-field/125688/9 "2018-05-01T18:05:05Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
