# Fscrawler only indexed 59 of a 2000 page pdf

**URL:** <https://discuss.elastic.co/t/fscrawler-only-indexed-59-of-a-2000-page-pdf/312262>\
**Category:** Elasticsearch\
**Created:** [August 17, 2022, 9:31am UTC](https://discuss.elastic.co/t/fscrawler-only-indexed-59-of-a-2000-page-pdf/312262 "2022-08-17T09:31:02Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![defalt](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/defalt/32/71379_2.png) [@defalt](https://discuss.elastic.co/u/defalt)\
**Post date:** [August 17, 2022, 9:31am UTC](https://discuss.elastic.co/t/fscrawler-only-indexed-59-of-a-2000-page-pdf/312262/1 "2022-08-17T09:31:03Z")

</div>

Hello 👋

i guess @dadoonet would know best about this but if someone else can answer this it would be great.

I am currently trying to index pdfs into elasticsearch. I installed the 'Ingest Attachment Processor Plugin' and downloaded fscrawler.zip. I unpacked it, ran `bin/fscrawler testjob`, edited the created \_settings.yaml to the right url for my pdfs and restarted the job.  
This was the output:

```auto
10:48:02,119 INFO [f.p.e.c.f.c.BootstrapChecks] Memory [Free/Total=Percent]: HEAP [235.9mb/3.8gb=6.01%], RAM [223.6mb/15.3gb=1.43%], Swap [940.8mb/1.9gb=45.94%].
10:48:02,316 INFO [f.p.e.c.f.FsCrawlerImpl] Starting FS crawler
10:48:02,316 INFO [f.p.e.c.f.FsCrawlerImpl] FS crawler started in watch mode. It will run unless you stop it with CTRL+C.
10:48:02,662 INFO [f.p.e.c.f.c.ElasticsearchClient] Elasticsearch Client connected to a node running version 7.17.3
10:48:02,706 INFO [f.p.e.c.f.c.ElasticsearchClient] Elasticsearch Client connected to a node running version 7.17.3
10:48:05,030 INFO [f.p.e.c.f.FsParserAbstract] FS crawler started for [testjob] for [/home/administrator/Downloads/Pdfs] every [15m]
10:48:05,344 INFO [f.p.e.c.f.t.TikaInstance] OCR is disabled.
10:48:06,529 WARN [o.a.p.p.f.FileSystemFontProvider] New fonts found, font cache will be re-built
10:48:06,529 WARN [o.a.p.p.f.FileSystemFontProvider] Building on-disk font cache, this may take a while
10:48:11,649 WARN [o.a.p.p.f.FileSystemFontProvider] Finished building on-disk font cache, found 288 fonts
10:48:11,776 WARN [o.a.p.p.f.PDType1Font] Using fallback font LiberationSans for base font Symbol
10:48:11,777 WARN [o.a.p.p.f.PDType1Font] Using fallback font LiberationSans for base font ZapfDingbats

```

It indexed the documents and so far so good. I see all 3 pdfs in the index and also most of its content. Now I encountered a problem, it only indexed 59 of the 2435 pages (I tried it with a pdf of the bible just for testing).

I dont know what the limiting factor is. Is it elasticsearch only allowing so many charecters or do I have to change some fscrawler setting?

Thanks for any help 🙂

---

<div class="post-metadata">

**Author:** ![dadoonet](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dadoonet/32/137187_2.png) [@dadoonet](https://discuss.elastic.co/u/dadoonet)\
**Post date:** [August 17, 2022, 1:06pm UTC](https://discuss.elastic.co/t/fscrawler-only-indexed-59-of-a-2000-page-pdf/312262/2 "2022-08-17T13:06:10Z")

</div>

Have a look at this setting : [Local FS settings — FSCrawler 2.10-SNAPSHOT documentation](https://fscrawler.readthedocs.io/en/latest/admin/fs/local-fs.html#extracted-characters)

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 14, 2022, 1:06pm UTC](https://discuss.elastic.co/t/fscrawler-only-indexed-59-of-a-2000-page-pdf/312262/3 "2022-09-14T13:06:32Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
