# Dec 7th, 2023: \[EN\] Find book about Christmas without searching for Christmas

**URL:** <https://discuss.elastic.co/t/dec-7th-2023-en-find-book-about-christmas-without-searching-for-christmas/348129>\
**Category:** Advent Calendar\
**Created:** [November 28, 2023, 11:12am UTC](https://discuss.elastic.co/t/dec-7th-2023-en-find-book-about-christmas-without-searching-for-christmas/348129 "2023-11-28T11:12:58Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![lio](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/lio/32/87276_2.png) [@lio](https://discuss.elastic.co/u/lio)\
**Post date:** [November 28, 2023, 11:12am UTC](https://discuss.elastic.co/t/dec-7th-2023-en-find-book-about-christmas-without-searching-for-christmas/348129/1 "2023-11-28T11:12:58Z")

</div>

As we’re getting closer to the holiday season, I’m looking forward to getting cozy, picking up a new book and having a relaxing time.

But book discovery online using a search bar is not as easy as it seems.... Most retail search engines rely solely on keyword searches, which is fine when we know exactly what title we're looking for, but it becomes more challenging when we only have a vague idea of the theme.

So, for this quick post, I decided to explore how I could leverage [Elasticsearch’s support for semantic search](https://www.elastic.co/guide/en/elasticsearch/reference/current/semantic-search.html) to help people who want to find a book about Christmas… without using the word “Christmas”.

For our example, we’ll be using a [dataset containing book summaries](https://raw.githubusercontent.com/elastic/elasticsearch-labs/main/datasets/book_summaries_1000_chunked.json). To follow along, you will need an Elasticsearch cluster up and running with the [ELSER model downloaded](https://www.elastic.co/guide/en/machine-learning/8.11/ml-nlp-elser.html#download-deploy-elser),

First, let’s configure an ingest pipeline to generate the sparse vectors for each book synopsis.

```auto
# Init Elasticsearch connection
es = Elasticsearch(
 cloud_id=ELASTIC_CLOUD_ID,
 api_key=ELASTIC_API_KEY,
 request_timeout=600
)

# ingest pipeline definition
PIPELINE_ID="vectorize_books_elser"

es.ingest.put_pipeline(id=PIPELINE_ID, processors=[{
     "foreach": {
         "field": "synopsis_passages",
         "processor": {
           "inference": {
             "field_map": {
               "_ingest._value.text": "text_field"
             },
             "model_id": ".elser_model_2_linux-x86_64",
             "target_field": "_ingest._value.vector",
             "on_failure": [
               {
                 "append": {
                   "field": "_source._ingest.inference_errors",
                   "value": [
                     {
                       "message": "Processor 'inference' in pipeline 'ml-inference-title-vector' failed with message '{{ _ingest.on_failure_message }}'",
                       "pipeline": "ml-inference-title-vector",
                       "timestamp": "{{{ _ingest.timestamp }}}"
                     }
                   ]
                 }
               }
             ]
           }
         }
       }
}])

```

Then create the books index where we will index our documents.

```auto
# Define the mapping
mappings = {
   "properties": {
       "title": {"type": "text"},
       "published_date": {"type": "text"},
       "synopsis": {"type": "text"},
       "synopsis_passages": {
         "type": "nested",
         "properties": {
             "vector": {
               "properties": {
                 "is_truncated": {
                   "type": "boolean"
                 },
                 "model_id": {
                   "type": "text",
                   "fields": {
                     "keyword": {
                       "type": "keyword",
                       "ignore_above": 256
                     }
                   }
                 },
                 "predicted_value": {
                   "type": "sparse_vector
"
                 }
            }
         }
     }
   }
}
}

# Create the index (deleting any previously existing index)
es.indices.delete(index="books", ignore_unavailable=True)
es.indices.create(index="books", mappings=mappings)

```

Now we can use the bulk API to ingest our documents. Note that we pass the pipeline name created previously to enrich documents using our ELSER ML model.

```auto
url = "https://raw.githubusercontent.com/elastic/elasticsearch-labs/main/datasets/book_summaries_1000_chunked.json"
response = urlopen(url)
books = json.loads(response.read())

from elasticsearch.helpers import streaming_bulk
count = 0
def generate_actions(books):

 for book in books:
   doc = {}
   doc["_index"] = "books"
   doc["pipeline"] = "vectorize_books_elser"
   doc["_source"] = book
   yield doc

for ok, info in streaming_bulk(client=es, index="books", actions=generate_actions(books), chunk_size=50):
 if not ok:
   print(f"Unable to index {info['index']['_id']}: {info['index']['error']}")

```

We’re now ready for the interesting part: Testing some queries to see the results we’re getting. One great thing here is that Elasticsearch supports keyword search and semantic search with the same index, as long as the data has been indexed correctly, which is the case here. We have indexed the synopsis as `text` and also an array of `sparse vectors`.

Here we will try to find books about Christmas using the following queries:

- “Story with Santa Claus”
- “Xmas stories”
- “Gift receiving and festive season”

The query to search using keyword search (BM25) is the following:

```auto
POST books/_search
{
  "_source": ["title"], 
  "query": {
    "match": {
      "synopsis": "Xmas stories"
    }
  }
}

```

And the query to search using semantic search is this one:

```auto
POST books/_search
{
  "_source": [
    "title"
  ],
  "query": {
    "nested": {
      "path": "synopsis_passages",
      "query": {
        "text_expansion": {
          "synopsis_passages.vector.predicted_value": {
            "model_id": ".elser_model_2_linux-x86_64",
            "model_text": "Xmas stories"
          }
        }
      }
    }
  }
}

```

Because we’re not using the keyword Christmas, semantic search outperforms lexical search in this instance.

Look at the result for the first query: “Story with Santa Claus”. The semantic search looks way more relevant.

 ![advent](https://us1.discourse-cdn.com/elastic/original/3X/c/d/cd340a266270d6440b9cc8c46988cb474a3588fa.jpeg)

For the other two test queries, we're getting the following results:

- “Xmas stories”

- “Gift receiving and festive season”

I’ll let you read the books to see which ones are the most related to the Christmas celebrations 🙂

Have a great holiday!
