# Index pdf files to AWS Elasticsearch service using Elasticsearch File System Crawler

**URL:** <https://discuss.elastic.co/t/index-pdf-files-to-aws-elasticsearch-service-using-elasticsearch-file-system-crawler/132666>\
**Category:** Elasticsearch\
**Created:** [May 21, 2018, 2:28pm UTC](https://discuss.elastic.co/t/index-pdf-files-to-aws-elasticsearch-service-using-elasticsearch-file-system-crawler/132666 "2018-05-21T14:28:27Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![Fish](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/fish/32/19628_2.png) [@Fish](https://discuss.elastic.co/u/Fish)\
**Post date:** [May 21, 2018, 2:28pm UTC](https://discuss.elastic.co/t/index-pdf-files-to-aws-elasticsearch-service-using-elasticsearch-file-system-crawler/132666/1 "2018-05-21T14:28:27Z")

</div>

I can index pdf files to a local Elasticsearch using Elasticsearch File System Crawler. The default, fscrawler setting has port, host and scheme parameters as shown below.

```auto
{
  "name" : "job_name2",
  "fs" : {
    "url" : "/tmp/es",
    "update_rate" : "15m",
    "excludes" : ["~*"],
    "json_support" : false,
    "filename_as_id" : false,
    "add_filesize" : true,
    "remove_deleted" : true,
    "add_as_inner_object" : false,
    "store_source" : false,
    "index_content" : true,
    "attributes_support" : false,
    "raw_metadata" : true,
    "xml_support" : false,
    "index_folders" : true,
    "lang_detect" : false,
    "continue_on_error" : false,
    "pdf_ocr" : true,
    "ocr" : {
      "language" : "eng"
    }
  },
  "elasticsearch" : {
    "nodes" : [ {
      "host" : "127.0.0.1",
      "port" : 9200,
      "scheme" : "HTTP"
    } ],
    "bulk_size" : 100,
    "flush_interval" : "5s"
  },
  "rest" : {
    "scheme" : "HTTP",
    "host" : "127.0.0.1",
    "port" : 8080,
    "endpoint" : "fscrawler"
  }
}

```

However, I have difficulty using it to index to AWS elasticsearch service because to index to AWS elasticsearch, I have to provide the AWS\_ACCESS\_KEY, AWS\_SECRET\_KEY, region, and service as documented [here](https://docs.aws.amazon.com/elasticsearch-service/latest/developerguide/es-indexing-programmatic.html). Any help on how to index pdf files to AWS elasticsearch service is highly appreciated.

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [May 28, 2018, 8:39am UTC](https://discuss.elastic.co/t/index-pdf-files-to-aws-elasticsearch-service-using-elasticsearch-file-system-crawler/132666/2 "2018-05-28T08:39:20Z")

</div>

I know I tested it with the official cloud by elastic offer (see below).  
I don't know how AWS service works exactly but I guess you can have a username and password?  
In which case you can define them in the `elasticsearch.nodes` setting?

See [https://github.com/dadoonet/fscrawler#elasticsearch-settings](https://github.com/dadoonet/fscrawler#elasticsearch-settings) for more details.

* * *

BTW did you look at [https://www.elastic.co/cloud](https://www.elastic.co/cloud) and [https://aws.amazon.com/marketplace/pp/B01N6YCISK](https://aws.amazon.com/marketplace/pp/B01N6YCISK) ?

Cloud by elastic is the only way to have access to X-Pack. Think about what is there yet like Security, Monitoring, Reporting and what is coming like Canvas, SQL...

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [June 25, 2018, 8:39am UTC](https://discuss.elastic.co/t/index-pdf-files-to-aws-elasticsearch-service-using-elasticsearch-file-system-crawler/132666/3 "2018-06-25T08:39:26Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
