# Short field type taking 4 bytes instead of 2

**URL:** <https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809>\
**Category:** Elasticsearch\
**Created:** [July 3, 2020, 1:11pm UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809 "2020-07-03T13:11:11Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![yfful](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yfful/32/88024_2.png) [@yfful](https://discuss.elastic.co/u/yfful)\
**Post date:** [July 3, 2020, 1:11pm UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809/1 "2020-07-03T13:11:11Z")

</div>

Hello,

Looking at the `short` field type, I see that the `Field`s created are based on the `integer` implementation:

> <https://github.com/elastic/elasticsearch/blob/4366360895dbcd28bf993000b80c95f83ecb79a5/server/src/main/java/org/elasticsearch/index/mapper/NumberFieldMapper.java#L573>

However, this means that a `IntPoint` is used when a `short` field is indexed. This leads to the `short` point using 4 bytes per dimension (with only 1 dimension), from what I understand:

> <https://github.com/apache/lucene-solr/blob/05324e7b1813c43084fbce7f3e6305db0ac94c32/lucene/core/src/java/org/apache/lucene/document/IntPoint.java#L49>

There is probably something I am missing, but does that mean that a `short` field indexes values using 4 bytes per dimension, rather than 2 ? Can you shed some light on this ?

Thanks!

---

<div class="post-metadata">

**Author:** ![ywelsch](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ywelsch/32/7751_2.png) [@ywelsch](https://discuss.elastic.co/u/ywelsch)\
**Post date:** [July 3, 2020, 1:34pm UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809/2 "2020-07-03T13:34:27Z")

</div>

The [docs](https://www.elastic.co/guide/en/elasticsearch/reference/current/number.html#_which_type_should_i_use) state:

> As far as integer types ( `byte` , `short` , `integer` and `long` ) are concerned, you should pick the smallest type which is enough for your use-case. This will help indexing and searching be more efficient. Note however that storage is optimized based on the actual values that are stored, so picking one type over another one will have no impact on storage requirements.

---

<div class="post-metadata">

**Author:** ![yfful](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yfful/32/88024_2.png) [@yfful](https://discuss.elastic.co/u/yfful)\
**Post date:** [July 3, 2020, 1:42pm UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809/3 "2020-07-03T13:42:58Z")

</div>

Thanks for pointing this out, i've read it but didn't reflect on it.

> Note however that storage is optimized based on the actual values that are stored, so picking one type over another one will have no impact on storage requirements.

Ok so this means that how a field is stored to disk is handled internally, regardless of the actual type.

> This will help indexing and searching be more efficient.

What is the benefit of using `short` over `integer` at index/search time, if a `short` is represented internally as an `integer` (from what i've understood of the code) ?

---

<div class="post-metadata">

**Author:** ![ywelsch](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/ywelsch/32/7751_2.png) [@ywelsch](https://discuss.elastic.co/u/ywelsch)\
**Post date:** [July 3, 2020, 2:06pm UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809/4 "2020-07-03T14:06:33Z")

</div>

> [@yfful](#):
>
> Ok so this means that how a field is stored to disk is handled internally, regardless of the actual type.

correct. See Lucene's `packed` package: [org.apache.lucene.util.packed (Lucene 8.5.2 API)](https://lucene.apache.org/core/8_5_2/core/org/apache/lucene/util/packed/package-summary.html)

> [@yfful](#):
>
> What is the benefit of using `short` over `integer` at index/search time, if a `short` is represented internally as an `integer` (from what i've understood of the code) ?

I'm not exactly sure why the docs are stated that way. The main benefit I see is that the document's data is validated at index time to fit into the given value range (byte, short, ...), which helps the underlying "packing" techniques to work well if the range of actual values is limited (i.e. no outliers).

@jpountz might have more insights here.

---

<div class="post-metadata">

**Author:** ![yfful](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/yfful/32/88024_2.png) [@yfful](https://discuss.elastic.co/u/yfful)\
**Post date:** [July 3, 2020, 3:31pm UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809/5 "2020-07-03T15:31:15Z")

</div>

@ywelsch Thanks for the reference for the packing, that's interesting.

I've omitted by mistake in my post's description a link to `BKDWriter`, where the `bytesPerDim` field gets its value from a `FieldInfo`. I believe that field ultimately comes from `IntPoint#getType()` which I have referenced earlier. From what I can tell, the value `bytesPerDim` has an impact on the allocated byte arrays in `BKDWriter`.

> <https://github.com/apache/lucene-solr/blob/05324e7b1813c43084fbce7f3e6305db0ac94c32/lucene/core/src/java/org/apache/lucene/util/bkd/BKDWriter.java#L165>

---

<div class="post-metadata">

**Author:** ![jpountz](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/jpountz/32/45836_2.png) [@jpountz](https://discuss.elastic.co/u/jpountz)\
**Post date:** [July 6, 2020, 8:04am UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809/6 "2020-07-06T08:04:04Z")

</div>

@ywelsch is right, the only benefit of `short` over `integer` is input validation.

Having native support for shorts would **not** help reduce disk space or search-time memory usage. The main benefit is that it would help save some memory in the IndexWriter buffer and thus create new segments a bit less frequently.

I'm rarely seeing `short`s in mappings, so given that there are only tiny benefits over `integer`, I think it's the right trade-off to implement them as `integer`s under the hood. I understand why the documentation might sound surprising once you're familiar with this implementation detail, but in my opinion documenting it would be even more confusing, so I don't dislike it that way.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [August 3, 2020, 8:04am UTC](https://discuss.elastic.co/t/short-field-type-taking-4-bytes-instead-of-2/239809/7 "2020-08-03T08:04:18Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
