# Recommended workflow for indexing many binary docs

**URL:** <https://discuss.elastic.co/t/recommended-workflow-for-indexing-many-binary-docs/275142>\
**Category:** Elasticsearch\
**Created:** [June 7, 2021, 11:49am UTC](https://discuss.elastic.co/t/recommended-workflow-for-indexing-many-binary-docs/275142 "2021-06-07T11:49:11Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![ottadini](https://avatars.discourse-cdn.com/v4/letter/o/b19c9b/32.png) [@ottadini](https://discuss.elastic.co/u/ottadini)\
**Post date:** [June 7, 2021, 11:49am UTC](https://discuss.elastic.co/t/recommended-workflow-for-indexing-many-binary-docs/275142/1 "2021-06-07T11:49:11Z")

</div>

I have two options in front of me -

1. use fscrawler to crawl my directory, use the setting `fs:store_source` (Base64-encoded document) to add the file as a field to the message, send that message to a pipeline on ES that has the ingest-attachment plugin. If it's a static set of files, then I may not even need fscrawler.... just some simple code to crawl the directory and encode the files before sending them to ES.
2. use fscrawler to do something... but I'm not sure what it is. It's hard for me to understand it. Does fscrawler somehow replace the ingest-attachment plugin? The docs seem to suggest that it doesn't need ingest-attachment plugin.

ben

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 7, 2021, 1:58pm UTC](https://discuss.elastic.co/t/recommended-workflow-for-indexing-many-binary-docs/275142/2 "2021-06-07T13:58:53Z")

</div>

You can use the [ingest attachment plugin](https://www.elastic.co/guide/en/elasticsearch/plugins/current/ingest-attachment.html).

There an example here: [Using the Attachment Processor in a Pipeline | Elasticsearch Plugins and Integrations [7.13] | Elastic](https://www.elastic.co/guide/en/elasticsearch/plugins/current/using-ingest-attachment.html)

```auto
PUT _ingest/pipeline/attachment
{
  "description" : "Extract attachment information",
  "processors" : [
    {
      "attachment" : {
        "field" : "data"
      }
    }
  ]
}
PUT my_index/_doc/my_id?pipeline=attachment
{
  "data": "e1xydGYxXGFuc2kNCkxvcmVtIGlwc3VtIGRvbG9yIHNpdCBhbWV0DQpccGFyIH0="
}
GET my_index/_doc/my_id

```

The `data` field is basically the BASE64 representation of your binary file.

You can use [FSCrawler](https://fscrawler.readthedocs.io). There's [a tutorial](https://fscrawler.readthedocs.io/en/latest/user/tutorial.html) to help you getting started.

Indeed FSCrawler does not need a pipeline to work as it does the extraction in FSCrawler instead of elasticsearch.

---

<div class="post-metadata">

**Author:** ![ottadini](https://avatars.discourse-cdn.com/v4/letter/o/b19c9b/32.png) [@ottadini](https://discuss.elastic.co/u/ottadini)\
**Post date:** [June 8, 2021, 12:31am UTC](https://discuss.elastic.co/t/recommended-workflow-for-indexing-many-binary-docs/275142/3 "2021-06-08T00:31:50Z")

</div>

Thank you David!  
I am wondering then which would be the better approach for my use case, where I have thousands of binary files to index, many many gigabytes. My current fscrawler settings file uses `store_source`, and then sends each file to an ES pipeline. The ES pipeline uses the ingest-attachment plugin. I don't use the REST server.

Would it be better to use fscrawler to extract the content from the binary docs, then send that content to a bare (no ingest-attachment plugin) ES index? And should that use the REST server, or some other method?

Here's my current settings file:

```auto
name: "job"
fs:
  url: "/projects/stock"
  continue_on_error: true
  index_content: "false"
  indexed_chars: "5000"
  ignore_above: "50mb"
  store_source: true
elasticsearch:
  nodes:
  - url: "http://localhost:9200"
  pipeline: "office-docs"
  index: "docs"

```

The current ES pipeline looks like this:

```auto
PUT _ingest/pipeline/office-docs
{
    "description": "Extract attachment information",
    "processors": [
      {
        "attachment": {
            "field": "attachment",
            "indexed_chars": -1
        }
      },
      {
        "remove": {
          "field": "attachment"
        }
      }
    ]
}

```

Using the fscrawler-only method like this:

```auto
name: "test"
fs:
  url: "/projects/stock"
  excludes:
  - "*/~*"
  continue_on_error: true
  index_folders: false
elasticsearch:
  nodes:
  - url: "http://localhost:9200"
  index: "test"

```

Is this more performant? Will it handle many gigabytes of data better than the ingest-attachment pipeline method?

Ben

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [June 8, 2021, 6:50am UTC](https://discuss.elastic.co/t/recommended-workflow-for-indexing-many-binary-docs/275142/4 "2021-06-08T06:50:30Z")

</div>

> [@ottadini](#):
>
> My current fscrawler settings file uses `store_source` , and then sends each file to an ES pipeline. The ES pipeline uses the ingest-attachment plugin. I don't use the REST server.

I'd not do that. I'd not use `store_source` option.

> [@ottadini](#):
>
> Would it be better to use fscrawler to extract the content from the binary docs, then send that content to a bare (no ingest-attachment plugin) ES index?

Yes. That'd remove a lot of memory pressure on elasticsearch nodes.

> [@ottadini](#):
>
> And should that use the REST server, or some other method?

No. This is not needed.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [July 6, 2021, 6:51am UTC](https://discuss.elastic.co/t/recommended-workflow-for-indexing-many-binary-docs/275142/5 "2021-07-06T06:51:12Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
