# Issue with provision large file to logstash

**URL:** <https://discuss.elastic.co/t/issue-with-provision-large-file-to-logstash/313719>\
**Category:** Logstash\
**Tags:** docker\
**Created:** [September 5, 2022, 10:31pm UTC](https://discuss.elastic.co/t/issue-with-provision-large-file-to-logstash/313719 "2022-09-05T22:31:10Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [September 5, 2022, 10:31pm UTC](https://discuss.elastic.co/t/issue-with-provision-large-file-to-logstash/313719/1 "2022-09-05T22:31:10Z")

</div>

Hi  
I'm facing with case with large file (~800MB) during transfer to logstash  
Indeed this is a case where data doesn't match in order  
This case has begun on ([Csv plugin cooperating with multiplying pattern - #11 by Badger](https://discuss.elastic.co/t/csv-plugin-cooperating-with-multiplying-pattern/309865/11)) @Badger Can You keep an eye once again?

at least I'm using such pipeline with current configuration(there I've decided to use PQ -persitent queue)

```auto
http.host: "0.0.0.0"
pipeline.workers: 1
pipeline.batch.size: 2000
pipeline.batch.delay: 50
pipeline.ordered: true
config.reload.automatic: true
xpack.monitoring.enabled: false
#xpack.management.pipeline.id: ["main"]
pipeline.ecs_compatibility: disabled
log.level: info

```

```auto
- pipeline.id: npdb
  path.config: "/usr/share/logstash/pipeline/pipe1.yml"
  queue.type: persisted
  path.queue: /usr/share/logstash/data/queue/
  queue.max_bytes: 2000mb

```

as test I'm uploading file on tcp port

```auto
input {
  tcp { port => 12367
        codec => multiline { pattern => "^#" negate => true what => "previous" multiline_tag => "" }

      }
	
}

filter {
    if [message] =~ "# 20" { drop{ } }
    if [message] =~ "table" { drop{ } }
    if [message] =~ "# number of Blocks" { drop{ } }
    mutate { remove_field => ["[event]", "log" ] }
    if "# snapshot" in [message] {
	 dissect {
            mapping => {
                "[message]" => "# %{activity},%{val},%{time}"
            }
            remove_field => ["[message]"]
        }
        date {
                match => ["time", "yyyyMMddHHmmss"]
                timezone => "Europe/Paris"
            }
	ruby { code => '@@metadata = event.get("@timestamp")' }
         # mutate { add_field => { "eventType" => "Header" } }
	drop {}
    } else if "# Network Entities" in [message] {
        mutate { add_field => { "eventType" => "Network Entities" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["ID","Type","PCType","PC","GC","RI","SSN","CCGT","NTT","NNAI","NNP","DA","SRFIMSI"] 
		}
	}
	ruby { code => 'event.set("@timestamp", @@metadata)' }

    } else if "# DNs" in [message] {
        mutate { add_field => { "eventType" => "DNs" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["DN","IMS","PT","SP","RN","VMS","GRN","ASD","ST","NSDN","CGBL","CDBL"] 
		}
        }
	ruby { code => 'event.set("@timestamp", @@metadata)' }
    } else if "# DN Blocks" in [message] {
        mutate { add_field => { "eventType" => "DN Blocks" } }
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["BDN","EDN","PT","SP","RN","VMS","GRN","ASD","ST","NSDN","CGBL","CDBL"] 
		}
	}
	ruby { code => 'event.set("@timestamp", @@metadata)' }
    } 
    
    else {
        mutate { add_field => {"eventType" => "Blocs"}}
        split { field => "message" }
        if [message] !~ /^#/ {
            csv { columns => ["IM","SVN","WHITE","GRAY","BLACK"] 
          
		} 
        } 
	ruby { code => 'event.set("@timestamp", @@metadata)' }
    }
mutate {
        remove_field => ["host", "count", "fields", "@version", "input_type", "source", "tags", "type", "time", "path", "activity", "val", "message", "port"]
        }
}

```

what is strange that the file consists a huge of data rows  
but as I observed it was processes only 501 of hints for eventType -\> DNs or Network Entities. whether I shorten the log or not. At least I've concluded that this pipeline diverges with data matching at some point, as if it gets lost after a certain number of processed records.  
this is sample data [https://filetransfer.io/data-package/SkuqBzeX#link](https://filetransfer.io/data-package/SkuqBzeX#link)  
Thanks for Your insight.

---

<div class="post-metadata">

**Author:** ![INS](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ins/32/92827_2.png) [@INS](https://discuss.elastic.co/u/INS)\
**Post date:** [September 6, 2022, 7:32am UTC](https://discuss.elastic.co/t/issue-with-provision-large-file-to-logstash/313719/2 "2022-09-06T07:32:03Z")

</div>

I have some doubts that multiline\_codec\_max\_lines\_reached ... ?  
I think that multiline codec is not the solution for large input of row.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 4, 2022, 7:33am UTC](https://discuss.elastic.co/t/issue-with-provision-large-file-to-logstash/313719/3 "2022-10-04T07:33:00Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
