# How to load a json file that contains special characters

**URL:** <https://discuss.elastic.co/t/how-to-load-a-json-file-that-contains-special-characters/67147>\
**Category:** Elasticsearch\
**Created:** [November 25, 2016, 2:47am UTC](https://discuss.elastic.co/t/how-to-load-a-json-file-that-contains-special-characters/67147 "2016-11-25T02:47:21Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![AussiePete2015](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aussiepete2015/32/13407_2.png) [@AussiePete2015](https://discuss.elastic.co/u/AussiePete2015)\
**Post date:** [November 25, 2016, 2:47am UTC](https://discuss.elastic.co/t/how-to-load-a-json-file-that-contains-special-characters/67147/1 "2016-11-25T02:47:21Z")

</div>

Hi all,

I have created a json file using Talend to load the text from sas code and transform it into json,  
I have created an index however, the import fails because the sas code contains many different symbols

e.g.  
{"index":{"\_index":"sascode\_idx", "\_type":"content", "\_id": "1"}}  
{"BuildAllTriangles":[{"content":"﻿/\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\r\n\* PROGRAM NAME : BuildAllTriangles.sas\r\n\* PROGRAMMER : Peter Birk\r\n\* DATE WRITTEN : 20120912\r\n\* DESCRIPTION : \r\n\r\n\tMake lots of liability triangles in an improbably short amount of\r\n\tdevelopment time available.\r\n\r\n\* DEPENDENCIES :\r\n\r\n\tRawFiles\Reference\CC\Reserving Triangle Delivery.xls\r\n\r\n\tMacro variables from SplitTransByReservingClass:\r\n\r\n\t&&SplitData&n..\r\n\t&SplitDataCount.\r\n\r\n\* OUTPUTS"}]}

As you can see this is a json array which I've verified via [http://jsonviewer.stack.hu/](http://jsonviewer.stack.hu/)

I can load this json file into MongoDB but obviously there is an issue with elasticsearch and the characters in the content.

How can I modify the content to be accepted into Elasticsearch without altering the content too drastically?

Cheers

---

<div class="post-metadata">

**Author:** ![guilherme\_maranhao](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/guilherme_maranhao/32/101483_2.png) [@guilherme\_maranhao](https://discuss.elastic.co/u/guilherme_maranhao)\
**Post date:** [November 30, 2016, 10:25am UTC](https://discuss.elastic.co/t/how-to-load-a-json-file-that-contains-special-characters/67147/2 "2016-11-30T10:25:36Z")

</div>

Hi Aussie,

We've faced a similar issue in our indexing process. What we've done was removing all the special characters with a gsub method (our interface language to Elasticsearch is Ruby):

`content = content.gsub(/[\“\”\"\'\\\']/m, ' ').gsub(/[\n\t\r]/m, ' ').gsub(/\s+/m, ' ').strip`

Hope it works for you!

Guilherme

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [December 28, 2016, 10:25am UTC](https://discuss.elastic.co/t/how-to-load-a-json-file-that-contains-special-characters/67147/3 "2016-12-28T10:25:37Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
