# How to upload a file into ElastiSearch

**URL:** <https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491>\
**Category:** Elasticsearch\
**Created:** [August 3, 2020, 2:37am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491 "2020-08-03T02:37:09Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![noor\_basha](https://avatars.discourse-cdn.com/v4/letter/n/8c91f0/32.png) [@noor\_basha](https://discuss.elastic.co/u/noor_basha)\
**Post date:** [August 3, 2020, 2:37am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/1 "2020-08-03T02:37:09Z")

</div>

Hi I am new to ELK,  
I need to upload different types of files into ELASTIC SEARCH, is it possible to do that. if it is anyone can please help me how to do with sample code. As of now, I am doing with reading the file and sending the data. out of 6k files of one zip file, I am able to send only around 3k. But I need to attach a file directly without reading its content.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 3, 2020, 4:34am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/2 "2020-08-03T04:34:52Z")

</div>

Welcome to our community! 😃  
What sort of files are they?

---

<div class="post-metadata">

**Author:** ![noor\_basha](https://avatars.discourse-cdn.com/v4/letter/n/8c91f0/32.png) [@noor\_basha](https://discuss.elastic.co/u/noor_basha)\
**Post date:** [August 3, 2020, 4:44am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/3 "2020-08-03T04:44:23Z")

</div>

Hi warkolm,  
Thanks for the replay.  
It is a LINUX SOS REPORT zip file consist of so many files nearly (11k files) with different type of extensions like txt, conf, rules, xml, File, CRON ,CNF File, CF File, REPO File, ALLOW File, CFG File etc..

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 3, 2020, 4:45am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/4 "2020-08-03T04:45:03Z")

</div>

You could use Filebeat for most of it, but it would require a bit of configuration so that it'd extract the right patterns.

---

<div class="post-metadata">

**Author:** ![noor\_basha](https://avatars.discourse-cdn.com/v4/letter/n/8c91f0/32.png) [@noor\_basha](https://discuss.elastic.co/u/noor_basha)\
**Post date:** [August 3, 2020, 4:48am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/5 "2020-08-03T04:48:11Z")

</div>

Sorry i don't know much about ELK just few days back only i moved to this project, can't we strore it as a blob ?. I need to store all 11k files and again i need to get it back.

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 3, 2020, 4:51am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/6 "2020-08-03T04:51:04Z")

</div>

No, Elasticsearch is not a binary store. Everything is converted to json.

---

<div class="post-metadata">

**Author:** ![noor\_basha](https://avatars.discourse-cdn.com/v4/letter/n/8c91f0/32.png) [@noor\_basha](https://discuss.elastic.co/u/noor_basha)\
**Post date:** [August 3, 2020, 4:52am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/7 "2020-08-03T04:52:53Z")

</div>

Ok, so Filebeat is the only way to store all these files?

---

<div class="post-metadata">

**Author:** ![warkolm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/warkolm/32/39224_2.png) [@warkolm](https://discuss.elastic.co/u/warkolm)\
**Post date:** [August 3, 2020, 4:53am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/8 "2020-08-03T04:53:23Z")

</div>

It sends the data to Elasticsearch.

Check out [https://www.elastic.co/products/](https://www.elastic.co/products/) for a bit more info on everything.

---

<div class="post-metadata">

**Author:** ![noor\_basha](https://avatars.discourse-cdn.com/v4/letter/n/8c91f0/32.png) [@noor\_basha](https://discuss.elastic.co/u/noor_basha)\
**Post date:** [August 3, 2020, 4:54am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/9 "2020-08-03T04:54:01Z")

</div>

Ok thanks warkolm for your time.

---

<div class="post-metadata">

**Author:** ![noor\_basha](https://avatars.discourse-cdn.com/v4/letter/n/8c91f0/32.png) [@noor\_basha](https://discuss.elastic.co/u/noor_basha)\
**Post date:** [August 3, 2020, 11:56am UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/10 "2020-08-03T11:56:15Z")

</div>

HI warkolm

I am iterating all files and reading each content and creating a document in elastic search. out of 6k files i am able to read and upload only 3k for first time. each and every run the no:of uploading files decreases. in different different platforms it is inserting different number of files.  
here is my sample code:

def upload\_file\_to\_es(name='',path=''):  
if(name is not ''):  
try:  
try:  
with open(os.path.join(path, name),'r') as file\_obj:  
log\_file\_string = file\_obj.read()  
except:  
with open(os.path.join(path, name),'rb') as file\_obj:  
log\_file\_string = file\_obj.read())  
file\_path = os.path.join(path.replace("C:\Users\Manikindi\_shaik\_Noor\Music\100.64.24.138\_2020\_Apr\_15\_07\_10\nts\_sles12\_base\_200415\_0259",""), name)  
upload\_data={}  
upload\_data["path"]=file\_path  
upload\_data["file"]=name  
upload\_data["content"]=str(log\_file\_string)#json.dumps(log\_file\_string)  
upload\_data["ip"]="0.0.0.3"  
upload\_data["exec"]="execution-ghi-013"  
upload\_data["timestamp"]=datetime.datetime.now().strftime("%Y-%m-%dT%H:%M:%S")

```
        try:
            resp_elastic = requests.post(
                elastic_url,
                headers=headers,
                data=json.dumps(upload_data),
                verify=False
            )
            print("ES Entry completed for {}".format(file_path))
            count21+=1
            print("file number %s"%count21)
        except Exception as e:
            print("ERROR: %s"%e)
            print("ERROR: FILE %s"%file_path)
            
    except Exception as e:
        print("ERROR: %s"%str(e))
        print("ERROR: FILE %s"%file_path) 

```

each and everytime the no:of files entry is different to ELK any suggestions please

---

<div class="post-metadata">

**Author:** ![tomrade](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/tomrade/32/26626_2.png) [@tomrade](https://discuss.elastic.co/u/tomrade)\
**Post date:** [August 9, 2020, 6:54pm UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/11 "2020-08-09T18:54:22Z")

</div>

I see you are using requests there , I would recommend the python library (and in particular the streaming bulk helper [https://elasticsearch-py.readthedocs.io/en/master/helpers.html#example](https://elasticsearch-py.readthedocs.io/en/master/helpers.html#example))

As stated elasticsearch will not do well with the file contents you are likely finding indexing errors on files as elasticsearch has guessed the field mapping (based upon the first doc) and some file contents will not be valid for that field mapping, You can ingest it as a blob(base64 it in your python first) with a non analysed field mapping if you really need to store them (you cannot search non analysed fields however).

Easiest way is using an index mapping template  
[https://www.elastic.co/guide/en/elasticsearch/reference/current/index-templates.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/index-templates.html)

You need to enabled to false as so  
[https://www.elastic.co/guide/en/elasticsearch/reference/current/enabled.html](https://www.elastic.co/guide/en/elasticsearch/reference/current/enabled.html)

This is a bit of hacky workaround, I really recommend storing the data outside elasticsearch though and just ingesting a link to the file ie in s3/unc or something like that. Its not much more code to upload the file in python to something in then ingest the link for search purposes.

Offtopic  
the "Correct" way to do this is to use something like tika [https://tika.apache.org/](https://tika.apache.org/) to parse metadata from the contents and then ingest that in a uniform manner.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 9, 2020, 9:35pm UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/12 "2020-08-09T21:35:39Z")

</div>

> [@tomrade](#):
>
> the "Correct" way to do this is to use something like tika [https://tika.apache.org/](https://tika.apache.org/) to parse metadata from the contents and then ingest that in a uniform manner.

You can use the [ingest attachment plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html).

There an example here: [https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html](https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html)

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }
  ]
}
PUT my_index/_doc/my_id?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}
GET my_index/_doc/my_id

```

The `data` field is basically the BASE64 representation of your binary file.

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 6, 2020, 9:35pm UTC](https://discuss.elastic.co/t/how-to-upload-a-file-into-elastisearch/243491/13 "2020-09-06T21:35:42Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
