# Indexing Subtitles and Maintaining Timestamps

**URL:** <https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708>\
**Category:** Elasticsearch\
**Created:** [September 25, 2023, 6:36am UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708 "2023-09-25T06:36:22Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![ADarkDividedGem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adarkdividedgem/32/125799_2.png) [@ADarkDividedGem](https://discuss.elastic.co/u/ADarkDividedGem)\
**Post date:** [September 25, 2023, 6:36am UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/1 "2023-09-25T06:36:22Z")

</div>

I am wanting to index subtitles and also maintain the timestamp data.

My initial thought was to make each line of text a document with the start and end timestamps stored as fields for that document. For example the following two lines will be stored as two documents (e.g. start, end, text, title, date):

```auto
01:33:22.285 --> 01:33:24.365
Lorem Ipsum is simply dummy

01:33:24.365 --> 01:33:27.485
text of the printing typesetting

```

However, if I was to search for `text : "simply dummy text"` no single document contains that exact phrase so no results are returned. If I wanted the searches to be useful I would have to index every line of text as one document but I would loose the timestamp data.

Does anyone know how I can store the timestamp data since some subtitle files contain 3 hours worth of text and any questions answered by the search results will be directly followed by when exactly did this speech occur.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 25, 2023, 7:08am UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/2 "2023-09-25T07:08:38Z")

</div>

Could you share a typical document you are indexing so far?

---

<div class="post-metadata">

**Author:** ![ADarkDividedGem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adarkdividedgem/32/125799_2.png) [@ADarkDividedGem](https://discuss.elastic.co/u/ADarkDividedGem)\
**Post date:** [September 25, 2023, 8:20am UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/3 "2023-09-25T08:20:32Z")

</div>

Just your usual WebVTT file. For example I have one file for a 2 hour show that is 229KB and contains 3,730 captions each containing a start time, end time and one line of text no longer than 33 characters.

This means my mappings are as follows:

```auto
mappings = {
    "properties": {
        "episode_id": {"type": "integer"},
        "episode_title": {"type": "text", "analyzer": "english", "fielddata": "true"},
        "episode_start": {"type": "date", "format": "date_time_no_millis"},
        "subtitle_start": {"type": "date", "format": "date_time"},
        "subtitle_end": {"type": "date", "format": "date_time"},
        "subtitle_text": {"type": "text", "analyzer": "english", "fielddata": "true"},
    }
}

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 25, 2023, 9:04am UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/4 "2023-09-25T09:04:44Z")

</div>

In that case, you would index every single caption as a document, right?

Something like:

```auto
POST /captions/_doc
{
  "episode_id": 1,
  "episode_title": "Foo Bar Baz",
  "episode_start": "2023-09-25",
  "subtitle_start": 5602285,
  "subtitle_end": 5604365,
  "subtitle_text": "Lorem Ipsum is simply dummy"
}
POST /captions/_doc
{
  "episode_id": 1,
  "episode_title": "Foo Bar Baz",
  "episode_start": "2023-09-25",
  "subtitle_start": 5604365,
  "subtitle_end": 5607485,
  "subtitle_text": "text of the printing typesetting"
}

```

Is that correct? If so, what is wrong with that?

---

<div class="post-metadata">

**Author:** ![ADarkDividedGem](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/adarkdividedgem/32/125799_2.png) [@ADarkDividedGem](https://discuss.elastic.co/u/ADarkDividedGem)\
**Post date:** [September 25, 2023, 9:11am UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/5 "2023-09-25T09:11:52Z")

</div>

Yes that is correct but as I mentioned in my original post searching for the exact phrase "simply dummy text" returns no results. This is because that exact phrase spans across two captions and therefore two documents.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 25, 2023, 9:15am UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/6 "2023-09-25T09:15:42Z")

</div>

I understand now the problem...  
Thanks.

May be you need to do something "smarter", like adding text before to your subtitles. Ideally full sentences but that'd mean that you have to detect what a sentence is...  
So may be index something like:

```auto
POST /captions/_doc
{
  "episode_id": 1,
  "episode_title": "Foo Bar Baz",
  "episode_start": "2023-09-25",
  "subtitle_start": 5604365,
  "subtitle_end": 5607485,
  "subtitle_text": "text of the printing typesetting",
  "subtitle_text_with_previous": "Lorem Ipsum is simply dummy text of the printing typesetting"
}

```

I don't see a way to do that automatically in Elasticsearch.

---

<div class="post-metadata">

**Author:** ![aaron\_ximm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aaron_ximm/32/61229_2.png) [@aaron\_ximm](https://discuss.elastic.co/u/aaron_ximm)\
**Post date:** [September 26, 2023, 5:28pm UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/7 "2023-09-26T17:28:09Z")

</div>

AFAI know, the short answer is, this requires application-side document processing. (Unless you make one document per line, which is likely too much overhead.)

I can tell you that this is what we do in several applications:

- index the text, for matching purposes, with all subtitles (and eg. soft line breaks) removed
- for each hit to be displayed/used, we dereference by scanning the original text to find the highlighted snippet

This is obviously quite expensive and only viable for small result sets; it also requires storing the original time-code decorated document.

I have done this with this exact ASR format ftr.

There are some corner cases to be aware of, e.g. multiple occurrences of the same text in the document. In the worst cases you may not have sufficient context to disambiguate.

---

<div class="post-metadata">

**Author:** ![aaron\_ximm](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aaron_ximm/32/61229_2.png) [@aaron\_ximm](https://discuss.elastic.co/u/aaron_ximm)\
**Post date:** [September 26, 2023, 5:32pm UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/8 "2023-09-26T17:32:00Z")

</div>

Idle idea, if performance is critical you could perhaps store in your search documents an inverted index of some kind. I can imagine tree data structures that would be quite expensive to generate but very fast to consult to get back the timecode offset at a word resolution.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 24, 2023, 5:32pm UTC](https://discuss.elastic.co/t/indexing-subtitles-and-maintaining-timestamps/343708/9 "2023-10-24T17:32:53Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
