# Formatted text mapping for elasticsearch

**URL:** <https://discuss.elastic.co/t/formatted-text-mapping-for-elasticsearch/10198>\
**Category:** Elasticsearch\
**Created:** [December 29, 2012, 12:34am UTC](https://discuss.elastic.co/t/formatted-text-mapping-for-elasticsearch/10198 "2012-12-29T00:34:12Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Emery\_2](https://avatars.discourse-cdn.com/v4/letter/e/ce7236/32.png) [@Emery\_2](https://discuss.elastic.co/u/Emery_2)\
**Post date:** [December 29, 2012, 12:34am UTC](https://discuss.elastic.co/t/formatted-text-mapping-for-elasticsearch/10198/1 "2012-12-29T00:34:12Z")

</div>

Hello I have a text content in XML format with formatting tags inside text:  
 ... water chemical formula is H2O and the energy is  
E=MC2 formula.  
How to properly convert this example to JSON format for elastic search, and  
to keep the search features and highlighting consistent?

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 29, 2012, 7:06pm UTC](https://discuss.elastic.co/t/formatted-text-mapping-for-elasticsearch/10198/2 "2012-12-29T19:06:52Z")

</div>

Your challenges are:

- the "abstract" XML element is semantic markup, it means "here comes an  
abstract"

- "sub" /"sup" elements are (X)HTML markup and they mean "display me in  
superscript/subscript style on your favorite output device"

The "abstract" element need to be parsed, and you need to decide how to  
index abstracts in ES. The "sub"/sup" elements need to be dropped. Display  
markup in your index mixed up with your textual content will render your  
index unusable.

That's the reason why the Lucene community use HTML strip filter in the  
analysis phase, and so does  
ES: [http://www.elasticsearch.org/guide/reference/index-modules/analysis/htmlstrip-charfilter.html](http://www.elasticsearch.org/guide/reference/index-modules/analysis/htmlstrip-charfilter.html)

The ES highlighting uses HTML-like tags, but this is just for convenience,  
it could also be other pre\_tags/post\_tags,  
also non-XML: [http://www.elasticsearch.org/guide/reference/api/search/highlighting.html](http://www.elasticsearch.org/guide/reference/api/search/highlighting.html)

An alternative to stripping HTML tags is to convert the text to Markdown  
(or another tag-less markup language) before indexing

An XSL styleheet is here  
[http://getsymphony.com/download/xslt-utilities/view/20573/](http://getsymphony.com/download/xslt-utilities/view/20573/)

Assuming the markdown control characters are not interfering with your  
Lucene analysis and word search, you could even add a Markdown formatter to  
present your ES docs / snippets.

Jörg

--

---

<div class="post-metadata">

**Author:** ![jprante](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jprante/32/44941_2.png) [@jprante](https://discuss.elastic.co/u/jprante)\
**Post date:** [December 29, 2012, 7:32pm UTC](https://discuss.elastic.co/t/formatted-text-mapping-for-elasticsearch/10198/3 "2012-12-29T19:32:06Z")

</div>

Another note, there are Unicode characters SUPERSCRIPT TWO U+00B2 und  
SUBSCRIPT TWO U+2082 which may help as a replacement before indexing.

Jörg

--

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2017, 2:58am UTC](https://discuss.elastic.co/t/formatted-text-mapping-for-elasticsearch/10198/4 "2017-07-06T02:58:12Z")

</div>


