# Unable to "import" json file into ES2.1

**URL:** <https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857>\
**Category:** Elasticsearch\
**Created:** [November 30, 2015, 3:15am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857 "2015-11-30T03:15:50Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [November 30, 2015, 3:15am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/1 "2015-11-30T03:15:50Z")

</div>

Why does this not work in ES2.1?:

curl -XPOST "[http://localhos:9200/\_bulk](http://localhos:9200/_bulk)" --data-binary @I:\ES\flu\_tweet\_file.json {"create": {"\_index": "flu", "\_type": "tweets", "\_id": 1}} \n {"title": "flu tweets"} \n

(note: the json file is the direct result from a search of twitter, so it is genuine.)

The errors I got:

type: illegal argument exception / malformed action/metadata line 1; expected START\_object or END\_object but found [VALUE\_STRING], STATUS 400  
curl: (3) [globbing] unmatched brace in column 1  
curl: (6) could not resolve host: flu,  
curl: (6) could not resolve host: \_type  
curl: (6) could not resolve host: tweets,  
curl: (6) could not resolve host: \_id  
curl: (3) [globbing] unmatched close brace/bracket in column 2  
curl: (6) could not resolve host: \n  
curl: (3) [globbing] unmatched close brace in column 1  
curl: (3) [globbing] unmatched close brace/bracket in column 11  
curl: (6) could not resolve host: \n

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [November 30, 2015, 6:21am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/2 "2015-11-30T06:21:24Z")

</div>

Why do you add a json content after the filename?

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 2, 2015, 4:52am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/3 "2015-12-02T04:52:14Z")

</div>

This curl:  
c:\>curl -X POST "[http://localhost:9200/\_bulk](http://localhost:9200/_bulk)" -d @I:\ES\flu\_tweet\_file.json

results in this error:

{"error":{"root\_cause":[{"type":"action\_request\_validation\_exception","reason":"  
Validation Failed: 1: no requests added;"}],"type":"action\_request\_validation\_ex  
ception","reason":"Validation Failed: 1: no requests added;"},"status":400}

and when I try this curl:  
c:\>curl -XPOST "[http://localhost:9200/\_bulk](http://localhost:9200/_bulk)" --data-binary @I:\ES\flu\_tweet\_file.json

I get this error:

{"error":{"root\_cause":[{"type":"illegal\_argument\_exception","reason":"Malformed action/metadata line [1], expected START\_OBJECT or END\_OBJECT but found [VALUE\_STRING]"}],"type":"illegal\_argument\_exception","reason":"Malformed action/metadata line [1], expected START\_OBJECT or END\_OBJECT but found [VALUE\_STRING]"},"status":400}

In short, I still can't get ES to "import" (index) a (bulk) json file.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 5:56am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/4 "2015-12-02T05:56:46Z")

</div>

What is in flu\_tweet\_file.json?

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 2, 2015, 6:37am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/5 "2015-12-02T06:37:30Z")

</div>

Twitter API search result for "flu" resulted in json file format

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 2, 2015, 6:42am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/6 "2015-12-02T06:42:16Z")

</div>

Excerpt:

{"contributors": null, "truncated": false, "text": "Vaccinate your child against flu. More at: https://t.co/3gHB3j5r0a #staywellthiswinter https://t.co/3XDXiAvFHU", "is\_quote\_status": false, "in\_reply\_to\_status\_id": null, "id": 669420613761671168, "favorite\_count": 0, "source": "\<a href="http://www.socialsignin.co.uk" rel="nofollow"\>SocialSignIn Application", "retweeted": false, "coordinates": null, "entities": {"symbols": [], "user\_mentions": [], "hashtags": [{"indices": [68, 87], "text": "staywellthiswinter"}], "urls": [{"url": "[https://t.co/3gHB3j5r0a](https://t.co/3gHB3j5r0a)", "indices": [44, 67], "expanded\_url": "[http://socsi.in/EBI6q](http://socsi.in/EBI6q)", "display\_url": "socsi.in/EBI6q"}], "media": [{"expanded\_url": "[http://twitter.com/west\_lei\_ccg/status/669420613761671168/photo/1](http://twitter.com/west_lei_ccg/status/669420613761671168/photo/1)", "display\_url": "[pic.twitter.com/3XDXiAvFHU](http://pic.twitter.com/3XDXiAvFHU)", "url": "[https://t.co/3XDXiAvFHU](https://t.co/3XDXiAvFHU)", "media\_url\_https": "[https://pbs.twimg.com/media/CUpCgCuWUAAf0By.jpg](https://pbs.twimg.com/media/CUpCgCuWUAAf0By.jpg)", "id\_str": "669420612872458240", "sizes": {"small": {"h": 226, "resize":

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 6:54am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/7 "2015-12-02T06:54:27Z")

</div>

That's not an expected bulk format. Look at the docs.

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 2, 2015, 2:02pm UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/8 "2015-12-02T14:02:53Z")

</div>

Following the docs,  
this curl gives this error:  
curl "[http://localhost:9200/fluproj/tweets/4901](http://localhost:9200/fluproj/tweets/4901)" -d @I:\ES\flu\_tweet\_file.json

{"error":{"root\_cause":[{"type":"mapper\_parsing\_exception","reason":"failed to parse"}],"type":"mapper\_parsing\_exception","reason":"failed to parse","caused\_by":{"type":"illegal\_argument\_exception","reason":"Malformed content, found extra data after parsing: START\_OBJECT"}},"status":400}

and this one:  
curl -XPOST "[http://localhost:9200/fluproj/tweets/4901](http://localhost:9200/fluproj/tweets/4901)" --data-binary @I:\ES\flu\_tweet\_file.json

{"error":{"root\_cause":[{"type":"mapper\_parsing\_exception","reason":"failed to parse"}],"type":"mapper\_parsing\_exception","reason":"failed to parse","caused\_by":{"type":"illegal\_argument\_exception","reason":"Malformed content, found extra data after parsing: START\_OBJECT"}},"status":400}

Please point me to the appropriate docs you are referring to. Thanks.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 3:15pm UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/9 "2015-12-02T15:15:48Z")

</div>

Did you search for "bulk format" in docs?

If you did, you should have read [https://www.elastic.co/guide/en/elasticsearch/guide/current/bulk.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/bulk.html) which is clear enough I think.

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 2, 2015, 4:56pm UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/10 "2015-12-02T16:56:01Z")

</div>

Yes, I did read that document. Several times. Please look at my posts again: I have made attempts to follow the document. It's when the attempts fail that I try other things.

Here are the issues as I see / understand them:

1. When I use Kibana -Sense (per ES2.1 / Kibana 4.3), I get the same errors as when I use command-line CURL
2. When I use CURL, I have not found a way to add a new line and complete the {action: statements}.

Based on the posts above and the documents you referred me to which should be simple enough but haven't worked out that way for me:  
a) What am I missing?  
b) Is my json file in the wrong format (is it pretty-printed -- I didn't design it that way; 'just used the search result as is.  
c) Do you need to see the Sense statements and results?  
d) What else can I do?  
Thanks.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 2, 2015, 5:21pm UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/11 "2015-12-02T17:21:35Z")

</div>

How can I know what you are doing without a concrete example?

> curl -XPOST "[http://localhos:9200/\_bulk](http://localhos:9200/_bulk)" --data-binary @I:\ES\flu\_tweet\_file.json {"create": {"index": "flu", "type": "tweets", "\_id": 1}} \n {"title": "flu tweets"} \n

This is the first thing you posted and this is wrong.

> {"contributors": null, "truncated": false, "text": "Vaccinate your child against flu. More at: [http://socsi.in/EBI6q](https://t.co/3gHB3j5r0a) #staywellthiswinter [https://twitter.com/west\_lei\_ccg/status/669420613761671168/photo/1](https://t.co/3XDXiAvFHU)", "is\_quote\_status": false, "in\_reply\_to\_status\_id": null, "id": 669420613761671168, "favorite\_count": 0, "source": "SocialSignIn Application", "retweeted": false, "coordinates": null, "entities": {"symbols": , "user\_mentions": , "hashtags": [{"indices": [68, 87], "text": "staywellthiswinter"}], "urls": [{"url": "[http://socsi.in/EBI6q](https://t.co/3gHB3j5r0a)", "indices": [44, 67], "expanded\_url": "[http://socsi.in/EBI6q](http://socsi.in/EBI6q)", "display\_url": "socsi.in/EBI6q"}], "media": [{"expanded\_url": "[http://twitter.com/west\_lei\_ccg/status/669420613761671168/photo/1](http://twitter.com/west_lei_ccg/status/669420613761671168/photo/1)", "display\_url": "[https://twitter.com/west\_lei\_ccg/status/669420613761671168/photo/1](http://pic.twitter.com/3XDXiAvFHU)", "url": "[https://twitter.com/west\_lei\_ccg/status/669420613761671168/photo/1](https://t.co/3XDXiAvFHU)", "media\_url\_https": "[https://pbs.twimg.com/media/CUpCgCuWUAAf0By.jpg](https://pbs.twimg.com/media/CUpCgCuWUAAf0By.jpg)", "id\_str": "669420612872458240", "sizes": {"small": {"h": 226, "resize":

This is also wrong.

This is the only thing that I can tell based on what you provided so far.

So please reproduce your error with a script you can share and share this here or on [gist.github.com](http://gist.github.com). Read [https://www.elastic.co/help/](https://www.elastic.co/help/) if you need details.

If we can't reproduce your error, we can't help.

Also, beware of curl on windows. It really sucks.  
Yes you can consider SENSE. It supports bulk format.

> Is my json file in the wrong format (is it pretty-printed)

Most likely yes. The BULK format states:

> The lines cannot contain unescaped newline characters, as they would interfere with parsing. This means that the JSON must not be pretty-printed.

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 4, 2015, 7:29am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/12 "2015-12-04T07:29:46Z")

</div>

Thank you for your ongoing interest in helping me out.  
Task: generate a json file to be indexed later in ES from searching a topic in Twitter. The Twitter search is performed using Python. The resulting json file is "new\_tweet\_file.json"

Python code is posted below. Search term is "H3N2"

I am open to other ways of performing the same task so long as it yields the correct-fomat json file "importable" to ES for indexing.  
Thanks in advance.

```auto
from __future__ import division, print_function 
import twitter # work with Twitter APIs
import json # methods for working with JSON data
windows_system = True # set to True if this is a Windows computer
if windows_system:
    line_termination = '\r\n' # Windows line termination
if (windows_system == False):
    line_termination = '\n' # Unix/Linus/Mac line termination

 json_filename = 'new_tweet_file.json'  

full_text_filename = 'new_tweet_review_file.txt'  

partial_text_filename = 'new_tweet_text_file.txt'  

def oauth_login():

    CONSUMER_KEY = ''
    CONSUMER_SECRET = ''
    OAUTH_TOKEN = ''
    OAUTH_TOKEN_SECRET = ''
    
    auth = twitter.oauth.OAuth(OAUTH_TOKEN, OAUTH_TOKEN_SECRET,
                               CONSUMER_KEY, CONSUMER_SECRET)
    
    twitter_api = twitter.Twitter(auth=auth)
    return twitter_api

def twitter_search(twitter_api, q, max_results=200, **kw):
  
    search_results = twitter_api.search.tweets(q=q, count=100, **kw)
    
    statuses = search_results['statuses']
    
     max_results = min(1000, max_results)
    
    for _ in range(10): # 10*100 = 1000
        try:
            next_results = search_results['search_metadata']['next_results']
        except KeyError, e: # No more results when next_results doesn't exist
            break        
  
     kwargs = dict([ kv.split('=') 
                        for kv in next_results[1:].split("&") ])
        
        search_results = twitter_api.search.tweets(**kwargs)
        statuses += search_results['statuses']
        
        if len(statuses) > max_results: 
            break
            
    return statuses

twitter_api = oauth_login()   
print(twitter_api) # verify the connection

q = "*H3N2*" # one of many possible search strings
results = twitter_search(twitter_api, q, max_results = 200) # limit to 200 tweets

print('\n\ntype of results:', type(results)) 
print('\nnumber of results:', len(results)) 
print('\ntype of results elements:', type(results[0]))

item_count = 0 # initialize count of objects dumped to file
with open(json_filename, 'w') as outfile:
    for dict_item in results:
        json.dump(dict_item, outfile, encoding = 'utf-8')
        item_count = item_count + 1
        if item_count < len(results):
             outfile.write(line_termination) # new line between items
                     
item_count = 0 # initialize count of objects dumped to file
with open(full_text_filename, 'w') as outfile:
    for dict_item in results:
        outfile.write('Item index: ' + str(item_count) +\
             ' -----------------------------------------' + line_termination)
        # indent for pretty printing
        outfile.write(json.dumps(dict_item, indent = 4))  
        item_count = item_count + 1
        if item_count < len(results):
             outfile.write(line_termination) # new line between items  
        
item_count = 0 # initialize count of objects dumped to file
with open(partial_text_filename, 'w') as outfile:
    for dict_item in results:
        outfile.write(json.dumps(dict_item['text']))
        item_count = item_count + 1
        if item_count < len(results):
             outfile.write(line_termination) # new line between text items  

```

Next step is to index the result in json file format in ES:

```
curl -XPOST "http://localhost:9200/fluproj/_bulk" -d @I:\ES\new_tweet_file.json

```

The errors and difficulties I have been posting come from using the bulk method as above.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 4, 2015, 7:56am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/13 "2015-12-04T07:56:40Z")

</div>

What is in `new_tweet_file.json`?

> I am open to other ways of performing the same task so long as it yields the correct-fomat json file "importable" to ES for indexing.

Use Logstash and its twitter input. You won't need to write and debug your own code...

I wrote a blog post about it: [https://david.pilato.fr/blog/2015-06-01-indexing-twitter-with-logstash-and-elasticsearch/](http://david.pilato.fr/blog/2015/06/01/indexing-twitter-with-logstash-and-elasticsearch/)

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 4, 2015, 8:52am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/14 "2015-12-04T08:52:17Z")

</div>

new\_tweet\_file.json is the output / result of the Twitter search per Python code above.  
Thanks for the new link. I am looking into it.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 4, 2015, 9:23am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/15 "2015-12-04T09:23:34Z")

</div>

Yeah I mean that I'm not going to execute your code so I was asking if you can share this file on [gist.github.com](http://gist.github.com) for example.

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 4, 2015, 10:01am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/16 "2015-12-04T10:01:41Z")

</div>

Got the 627Kb file ready. Not sure how to share it on [gist.github.com](http://gist.github.com); getting "server not found" browser error. How else can I upload it?  
Thanks.

In the mean time, I need to read up on Logstash. I'd like to use the script you supplied.

BTW: Where can I buy the book which I am certain you have published on ES?

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 4, 2015, 10:24am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/17 "2015-12-04T10:24:35Z")

</div>

[https://gist.github.com/](https://gist.github.com/)

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 4, 2015, 10:25am UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/18 "2015-12-04T10:25:26Z")

</div>

> [@oguchi](#):
>
> BTW: Where can I buy the book which I am certain you have published on ES?

> **[Elasticsearch: The Definitive Guide](https://www.oreilly.com/library/view/elasticsearch-the-definitive/9781449358532/)**
>
> Whether you need full-text search or real-time analytics of structured data—or both—the Elasticsearch distributed search engine is an ideal way to put your data to work. This practical guide not … - Selection from Elasticsearch: The Definitive Guide...

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [December 4, 2015, 3:33pm UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/20 "2015-12-04T15:33:30Z")

</div>

It's incorrect and does not respect the bulk format.

Again, read the documentation: [https://www.elastic.co/guide/en/elasticsearch/guide/current/bulk.html](https://www.elastic.co/guide/en/elasticsearch/guide/current/bulk.html)

---

<div class="post-metadata">

**Author:** ![oguchi](https://avatars.discourse-cdn.com/v4/letter/o/e95f7d/32.png) [@oguchi](https://discuss.elastic.co/u/oguchi)\
**Post date:** [December 4, 2015, 3:58pm UTC](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857/21 "2015-12-04T15:58:56Z")

</div>

Great reference; I've been using it. Ordered another one on the same subject not yet released (Dec 8, I'm told by Amazon). But, I think it is for For Developers.  
Need one for real beginners and non-developers who may be data scientists or data enthusiasts.  
Analogy: teaching what a driver can do on the freeway or on a race course or obstacle course (a lot of spectacular things) versus teaching or including how to get on the freeway / race course / obstacle course in the first place.  
Thanks for your patience. If you have other books, I'm certain to get them.

[Next page](https://discuss.elastic.co/t/unable-to-import-json-file-into-es2-1/35857.md?page=2)
