# Code indexation mail to elasticsearch

**URL:** <https://discuss.elastic.co/t/code-indexation-mail-to-elasticsearch/130846>\
**Category:** Logstash\
**Created:** [May 7, 2018, 2:21pm UTC](https://discuss.elastic.co/t/code-indexation-mail-to-elasticsearch/130846 "2018-05-07T14:21:00Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![rouchad\_rouchad](https://avatars.discourse-cdn.com/v4/letter/r/4491bb/32.png) [@rouchad\_rouchad](https://discuss.elastic.co/u/rouchad_rouchad)\
**Post date:** [May 7, 2018, 2:21pm UTC](https://discuss.elastic.co/t/code-indexation-mail-to-elasticsearch/130846/1 "2018-05-07T14:21:00Z")

</div>

hello, I want to make a filter to pull some information from each email like (sender , @IP , text , name of the company etc ... I heard about grok filter but I can not use it well but since we can work with any computer language I chose PYTHON and I found a code that allows me to pull everything I want from each email. the problem what I want applies this code in logstash to find the same results on the interface kibana  
how to do it ?  
thank you

there is the code In PYTHON :  
import sys, os, re, StringIO  
import email, mimetypes

invalid\_chars\_in\_filename='\<\>:"/\|?\*%''+reduce(lambda x,y:x+chr(y), range(32), '')  
invalid\_windows\_name='CON PRN AUX NUL COM1 COM2 COM3 COM4 COM5 COM6 COM7 COM8 COM9 LPT1 LPT2 LPT3 LPT4 LPT5 LPT6 LPT7 LPT8 LPT9'.split()

# email address REGEX matching the RFC 2822 spec from perlfaq9

# my $atom = qr{[a-zA-Z0-9\_!#$%&'\*+/=?^`{}~|-]+};

# my $dot\_atom = qr{$atom(?:.$atom)\*};

# my $quoted = qr{"(?:\[^\r\n]|[^\"])\*"};

# my $local = qr{(?:$dot\_atom|$quoted)};

# my $domain\_lit = qr{[(?:\\S|[\x21-\x5a\x5e-\x7e])\*]};

# my $domain = qr{(?:$dot\_atom|$domain\_lit)};

# my $addr\_spec = qr{$local@$domain};

# 

# Python's translation

atom\_rfc2822=r"[a-zA-Z0-9\_!#$%&'_+/=?^`{}~|\-]+" atom_posfix_restricted=r"[a-zA-Z0-9_#\$&'*+/=?\^`{}~|-]+" # without '!' and '%'  
atom=atom\_rfc2822  
dot\_atom=atom + r"(?:." + atom + ")_"  
quoted=r'"(?:\[^\r\n]|[^\"])_"'  
local="(?:" + dot\_atom + "|" + quoted + ")"  
domain\_lit=r"[(?:\\S|[\x21-\x5a\x5e-\x7e])_]"  
domain="(?:" + dot\_atom + "|" + domain\_lit + ")"  
addr\_spec=local + "@" + domain

email\_address\_re=re.compile('^'+addr\_spec+'$')

class Attachment:  
def **init** (self, part, filename=None, type=None, payload=None, charset=None, content\_id=None, description=None, disposition=None, sanitized\_filename=None, is\_body=None):  
self.part=part # original python part  
self.filename=filename # filename in unicode (if any)  
self.type=type # the mime-type  
self.payload=payload # the MIME decoded content  
self.charset=charset # the charset (if any)  
self.description=description # if any  
self.disposition=disposition # 'inline', 'attachment' or None  
self.sanitized\_filename=sanitized\_filename # cleanup your filename here (TODO)  
self.is\_body=is\_body # usually in (None, 'text/plain' or 'text/html')  
self.content\_id=content\_id # if any  
if self.content\_id:  
# strip '\<\>' to ease searche and replace in "root" content (TODO)  
if self.content\_id.startswith('\<') and self.content\_id.endswith('\>'):  
self.content\_id=self.content\_id[1:-1]

def getmailheader(header\_text, default="ascii"):  
"""Decode header\_text if needed"""  
try:  
headers=email.Header.decode\_header(header\_text)  
except email.Errors.HeaderParseError:  
# This already append in email.base64mime.decode()  
# instead return a sanitized ascii string  
# this faile '=?UTF-8?B?15HXmdeh15jXqNeVINeY15DXpteUINeTJ9eV16jXlSDXkdeg15XXldeUINem15PXpywg15TXptei16bXldei15nXnSDXqdecINek15zXmdeZ?==?UTF-8?B?157XldeR15nXnCwg157Xldek16Ig157Xl9eV15wg15HXodeV15bXnyDXk9ec15DXnCDXldeh15gg157Xl9eR16rXldeqINep15wg15HXmdeQ?==?UTF-8?B?15zXmNeZ?='  
return header\_text.encode('ascii', 'replace').decode('ascii')  
else:  
for i, (text, charset) in enumerate(headers):  
try:  
headers[i]=unicode(text, charset or default, errors='replace')  
except LookupError:  
# if the charset is unknown, force default  
headers[i]=unicode(text, default, errors='replace')  
return u"".join(headers)

def getmailaddresses(msg, name):  
"""retrieve addresses from header, 'name' supposed to be from, to, ..."""  
addrs=email.utils.getaddresses(msg.get\_all(name, []))  
for i, (name, addr) in enumerate(addrs):  
if not name and addr:  
# only one string! Is it the address or is it the name ?  
# use the same for both and see later  
name=addr

```
    try:
        # address must be ascii only
        addr=addr.encode('ascii')
    except UnicodeError:
        addr=''
    else:
        # address must match address regex
        if not email_address_re.match(addr):
            addr=''
    addrs[i]=(getmailheader(name), addr)
return addrs

```

* * *

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 4, 2018, 2:21pm UTC](https://discuss.elastic.co/t/code-indexation-mail-to-elasticsearch/130846/2 "2018-06-04T14:21:01Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
