# Classification of String Data

**URL:** <https://discuss.elastic.co/t/classification-of-string-data/260194>\
**Category:** Elasticsearch\
**Tags:** elastic-stack-machine-learning\
**Created:** [January 5, 2021, 11:12am UTC](https://discuss.elastic.co/t/classification-of-string-data/260194 "2021-01-05T11:12:05Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Andrey\_Maksimov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrey_maksimov/32/81738_2.png) [@Andrey\_Maksimov](https://discuss.elastic.co/u/Andrey_Maksimov)\
**Post date:** [January 5, 2021, 11:12am UTC](https://discuss.elastic.co/t/classification-of-string-data/260194/1 "2021-01-05T11:12:06Z")

</div>

hey there,

i have a task to classify download stats (basically, that's **URL** + some minor but yet valuable fields like referring host, request country, etc) of open-source products our company provide into metadata like **site, product family, name & component, major, minor & patch version, OS type & version** , etc. and **number of downloads** of course.

at the moment, i have a script which analyses the URL and applies regular expressions to do that job. but number of different files grows, name conventions change over time (yes, i need to process historical data as well) so that's getting really hard to support that script via adding new regex'es.

for me that task looks like a great job for machine learning to classify the string into a set of definite keywords (names) and numbers (versions). luckily, Elastic proposes such functionality as an experimental feature.

by the moment i have some kind of "ground truth": the indices prepared by the script i could use to train a model initially. since Elastic mentions that's a supervised job, i will have a way to train the model further or provide a way for task stakeholders to do so.

but i am unsure if ElasticSearch's experimental ML functionality is the right tool for that task.

1. i've tried to use the feature but when i specify a dependent variable, but it shows  
_Invalid. Field [parsed.component.keyword] must have at most [30] distinct values but there were at least [34]_  
according to [the docs](https://www.elastic.co/guide/en/machine-learning/7.10/dfa-classification.html), dependent values "must contain no more than 30 classes" while i definitely have much more variations of the files to classify.
2. i did not find a way to specify other dependent variables, but the task supposes i need a whole set of different parameters in output.

does the above mean ElasticSearch ML does not fit my needs? am i doing or getting anything wrong?

---

<div class="post-metadata">

**Author:** ![przemekwitek](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/przemekwitek/32/79526_2.png) [@przemekwitek](https://discuss.elastic.co/u/przemekwitek)\
**Post date:** [January 5, 2021, 11:40am UTC](https://discuss.elastic.co/t/classification-of-string-data/260194/2 "2021-01-05T11:40:25Z")

</div>

Hi,

> but i am unsure if Elasticsearch's experimental ML functionality is the right tool for that task.

The purpose of the supervised ML classification job is to classify each document as belonging to one particular class. In other words, as described in [this documentation](https://www.elastic.co/guide/en/machine-learning/current/dfa-classification.html): _"Classification is a machine learning process that enables you to predict the class or category of a data point in your data set"_.

With that in mind, I'd say your problem does not really fit into this category.

> 1. i've tried to use the feature but when i specify a dependent variable, but it shows  
> _Invalid. Field [parsed.component.keyword] must have at most [30] distinct values but there were at least [34]_  
> according to [the docs](https://www.elastic.co/guide/en/machine-learning/7.10/dfa-classification.html), dependent values "must contain no more than 30 classes" while i definitely have much more variations of the files to classify.

As you have noticed, in elasticsearch's supervised ML, there is a restriction on the number of values of dependent variable. Currently it is set to `30`. The actual number is less important. The more important thing here is that the number of categories is (and will be) bounded whereas in your case it is unbounded (e.g.: there can be **many** unique sites or downloads).

> 1. i did not find a way to specify other dependent variables, but the task supposes i need a whole set of different parameters in output.

There can only be one dependent variable, i.e.: the class to which a document belongs.  
So with ML classification, you cannot get all those different parameters in the output.

> does the above mean Elasticsearch ML does not fit my needs? am i doing or getting anything wrong?

To sum up, I don't think the ML classification job can be used for data extraction problem. The main issue is potentially unbounded number of "classes" whereas in ML classification, the set of classes needs to be fixed.

---

<div class="post-metadata">

**Author:** ![Andrey\_Maksimov](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/andrey_maksimov/32/81738_2.png) [@Andrey\_Maksimov](https://discuss.elastic.co/u/Andrey_Maksimov)\
**Post date:** [January 5, 2021, 12:19pm UTC](https://discuss.elastic.co/t/classification-of-string-data/260194/3 "2021-01-05T12:19:53Z")

</div>

thank you much for that detailed explanation, Przemysław. it looks like i need to find a better way. btw, thank you for that great term "data extraction" which explains my task better than tons of words i used.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [February 2, 2021, 12:19pm UTC](https://discuss.elastic.co/t/classification-of-string-data/260194/4 "2021-02-02T12:19:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
