# Ingest pipeline merge subfields to JSON string?

**URL:** <https://discuss.elastic.co/t/ingest-pipeline-merge-subfields-to-json-string/222819>\
**Category:** Elasticsearch\
**Created:** [March 9, 2020, 10:41pm UTC](https://discuss.elastic.co/t/ingest-pipeline-merge-subfields-to-json-string/222819 "2020-03-09T22:41:07Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![jsosic](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jsosic/32/5945_2.png) [@jsosic](https://discuss.elastic.co/u/jsosic)\
**Post date:** [March 9, 2020, 10:41pm UTC](https://discuss.elastic.co/t/ingest-pipeline-merge-subfields-to-json-string/222819/1 "2020-03-09T22:41:07Z")

</div>

Hi guys,

I have my apps set to write JSON logs, and filebeat sends them enriched with host metadata to Elasticsearch cluster.

There they get processed by ingest pipeline which basically runs the following GROK:

```auto
    "processors": [
        {  
            "json": {
                "field": "message",
                "target_field": "apps"
            }  
        }, 

```

Now, the problem that I'm experiencing is that some of the fields in the json go 10+ levels deep and there's a huge amount of them, which causes the mapping explosion very soon after the index creation (within the first ~10k entries).

I know the exact names of the fields that are culprit, for example: `app.payload.params` has a bunch of subfields and levels, which get expanded to something like:

```auto
app.payload.params.001
app.payload.params.002
...
app.payload.params.100
...
app.payload.params.foo.001
...
app.payload.params.foo.002

```

Now, I would like to either limit the `JSON` processor to the depth of processing JSON, but reading the docs that doesn't seem possible. Another option I was thinking about was trying to merge all these fields back into one text field, but it seems Elastic doesn't support `ruby` processor?

So that means I need to run logstash cluster?

Are there any other options?

---

<div class="post-metadata">

**Author:** ![spinscale](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/spinscale/32/25011_2.png) [@spinscale](https://discuss.elastic.co/u/spinscale)\
**Post date:** [March 10, 2020, 1:25pm UTC](https://discuss.elastic.co/t/ingest-pipeline-merge-subfields-to-json-string/222819/2 "2020-03-10T13:25:18Z")

</div>

you could use a `script` processor for that.

Also, you may want to check out the newly added `flattened` datatype, see [https://www.elastic.co/guide/en/elasticsearch/reference/7.6/flattened.html](https://www.elastic.co/guide/en/elasticsearch/reference/7.6/flattened.html)

---

<div class="post-metadata">

**Author:** ![jsosic](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jsosic/32/5945_2.png) [@jsosic](https://discuss.elastic.co/u/jsosic)\
**Post date:** [March 12, 2020, 1:50am UTC](https://discuss.elastic.co/t/ingest-pipeline-merge-subfields-to-json-string/222819/3 "2020-03-12T01:50:24Z")

</div>

I upgraded to 7.6, and just tried flattened. This is the error in filebeat that I get:

```auto
 Connection marked as failed because the onConnect callback failed: error loading template: error creating template from file /etc/filebeat/fields-custom.yml: incorrect type configuration for field 'params': unexpected type 'flattened' for field 'params' 

```

This is the relevant part of the `field-custom.yml`:

```auto
    - name: params
      type: flattened
      description: Some desc.

```

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [April 9, 2020, 1:50am UTC](https://discuss.elastic.co/t/ingest-pipeline-merge-subfields-to-json-string/222819/4 "2020-04-09T01:50:33Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
