# What should fscrawler mapping look like to index each pdf document as a single unit of text?

**URL:** <https://discuss.elastic.co/t/what-should-fscrawler-mapping-look-like-to-index-each-pdf-document-as-a-single-unit-of-text/198660>\
**Category:** Elasticsearch\
**Created:** [September 9, 2019, 10:50am UTC](https://discuss.elastic.co/t/what-should-fscrawler-mapping-look-like-to-index-each-pdf-document-as-a-single-unit-of-text/198660 "2019-09-09T10:50:04Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![neergttocsdivad](https://avatars.discourse-cdn.com/v4/letter/n/43a26b/32.png) [@neergttocsdivad](https://discuss.elastic.co/u/neergttocsdivad)\
**Post date:** [September 9, 2019, 10:50am UTC](https://discuss.elastic.co/t/what-should-fscrawler-mapping-look-like-to-index-each-pdf-document-as-a-single-unit-of-text/198660/1 "2019-09-09T10:50:05Z")

</div>

ref.: [https://fscrawler.readthedocs.io/en/fscrawler-2.5/admin/fs/elasticsearch.html#creating-your-own-mapping-analyzers](https://fscrawler.readthedocs.io/en/fscrawler-2.5/admin/fs/elasticsearch.html#creating-your-own-mapping-analyzers)

I have read through the above page but still don't know what to do.

Below is the auto-generated mapping.  
I guess most of this can be removed, but I don't know what can be subtracted without killing it.

```auto
{
  "mapping": {
    "dynamic_templates": [
      {
        "raw_as_text": {
          "path_match": "meta.raw.*",
          "mapping": {
            "fields": {
              "keyword": {
                "ignore_above": 256,
                "type": "keyword"
              }
            },
            "type": "text"
          }
        }
      }
    ],
    "properties": {
      "attachment": {
        "type": "binary"
      },
      "attributes": {
        "properties": {
          "group": {
            "type": "keyword"
          },
          "owner": {
            "type": "keyword"
          }
        }
      },
      "content": {
        "type": "text"
      },
      "file": {
        "properties": {
          "checksum": {
            "type": "keyword"
          },
          "content_type": {
            "type": "keyword"
          },
          "created": {
            "type": "date",
            "format": "date_optional_time"
          },
          "extension": {
            "type": "keyword"
          },
          "filename": {
            "type": "keyword",
            "store": true
          },
          "filesize": {
            "type": "long"
          },
          "indexed_chars": {
            "type": "long"
          },
          "indexing_date": {
            "type": "date",
            "format": "date_optional_time"
          },
          "last_accessed": {
            "type": "date",
            "format": "date_optional_time"
          },
          "last_modified": {
            "type": "date",
            "format": "date_optional_time"
          },
          "url": {
            "type": "keyword",
            "index": false
          }
        }
      },
      "meta": {
        "properties": {
          "altitude": {
            "type": "text"
          },
          "author": {
            "type": "text"
          },
          "comments": {
            "type": "text"
          },
          "contributor": {
            "type": "text"
          },
          "coverage": {
            "type": "text"
          },
          "created": {
            "type": "date",
            "format": "date_optional_time"
          },
          "creator_tool": {
            "type": "keyword"
          },
          "date": {
            "type": "date",
            "format": "date_optional_time"
          },
          "description": {
            "type": "text"
          },
          "format": {
            "type": "text"
          },
          "identifier": {
            "type": "text"
          },
          "keywords": {
            "type": "text"
          },
          "language": {
            "type": "keyword"
          },
          "latitude": {
            "type": "text"
          },
          "longitude": {
            "type": "text"
          },
          "metadata_date": {
            "type": "date",
            "format": "date_optional_time"
          },
          "modifier": {
            "type": "text"
          },
          "print_date": {
            "type": "date",
            "format": "date_optional_time"
          },
          "publisher": {
            "type": "text"
          },
          "rating": {
            "type": "byte"
          },
          "relation": {
            "type": "text"
          },
          "rights": {
            "type": "text"
          },
          "source": {
            "type": "text"
          },
          "title": {
            "type": "text"
          },
          "type": {
            "type": "text"
          }
        }
      },
      "path": {
        "properties": {
          "real": {
            "type": "keyword",
            "fields": {
              "fulltext": {
                "type": "text"
              },
              "tree": {
                "type": "text",
                "analyzer": "fscrawler_path",
                "fielddata": true
              }
            }
          },
          "root": {
            "type": "keyword"
          },
          "virtual": {
            "type": "keyword",
            "fields": {
              "fulltext": {
                "type": "text"
              },
              "tree": {
                "type": "text",
                "analyzer": "fscrawler_path",
                "fielddata": true
              }
            }
          }
        }
      }
    }
  }
}

```

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [September 10, 2019, 11:27am UTC](https://discuss.elastic.co/t/what-should-fscrawler-mapping-look-like-to-index-each-pdf-document-as-a-single-unit-of-text/198660/2 "2019-09-10T11:27:19Z")

</div>

The question is:

> What should fscrawler mapping look like to index each pdf document as a single unit of text?

That's the default behavior of FSCrawler. It's not related to the mapping.  
What is the exact problem you're facing?

---

<div class="post-metadata">

**Author:** ![neergttocsdivad](https://avatars.discourse-cdn.com/v4/letter/n/43a26b/32.png) [@neergttocsdivad](https://discuss.elastic.co/u/neergttocsdivad)\
**Post date:** [September 11, 2019, 7:03pm UTC](https://discuss.elastic.co/t/what-should-fscrawler-mapping-look-like-to-index-each-pdf-document-as-a-single-unit-of-text/198660/3 "2019-09-11T19:03:03Z")

</div>

Hi  
It isn't really a problem  
Since I am interested in performing corpus-linguistic analyses of the text, I don't need all the metadata and I imagined that I could save indexing-time and drive-space by removing lines from the \_mappings

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [October 9, 2019, 7:03pm UTC](https://discuss.elastic.co/t/what-should-fscrawler-mapping-look-like-to-index-each-pdf-document-as-a-single-unit-of-text/198660/4 "2019-10-09T19:03:14Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
