# Pulling more than 2000 records from ServiceNow api using http\_poller

**URL:** <https://discuss.elastic.co/t/pulling-more-than-2000-records-from-servicenow-api-using-http-poller/280840>\
**Category:** Logstash\
**Created:** [August 9, 2021, 4:50pm UTC](https://discuss.elastic.co/t/pulling-more-than-2000-records-from-servicenow-api-using-http-poller/280840 "2021-08-09T16:50:37Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Purushottam22](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/purushottam22/32/92715_2.png) [@Purushottam22](https://discuss.elastic.co/u/Purushottam22)\
**Post date:** [August 9, 2021, 4:50pm UTC](https://discuss.elastic.co/t/pulling-more-than-2000-records-from-servicenow-api-using-http-poller/280840/1 "2021-08-09T16:50:37Z")

</div>

I have created the pipeline using http\_poller to pull the serviceNow CI records.  
To pull more than 2000 records, i am hitting the ServiceNow api multiple time with setting up start and count values using http\_poller in pipeline.

However currently i am receiving the duplicates records also, why i am receiving those duplicate records and how can i overcome from its.

below is my pipeline config file.

input {  
http\_poller {  
urls =\> {  
records2000 =\> {  
# Supports all options supported by ruby's Manticore HTTP client  
method =\> get  
url =\> "[https://service-now.com/api/bebup/config/ci/query/list?encoded\_query=install\_status!%3D104%26start%3D1%26count%3D2000%26use\_display\_value%3DTRUE](https://service-now.com/api/bebup/config/ci/query/list?encoded_query=install_status%21%3D104%26start%3D1%26count%3D2000%26use_display_value%3DTRUE)"  
headers =\> {  
Accept =\> "application/json"  
}  
}  
records4000 =\> {  
# Supports all options supported by ruby's Manticore HTTP client  
method =\> get  
url =\> "[https://service-now.com/api/bebup/config/ci/query/list?encoded\_query=install\_status!%3D104%26start%3D2001%26count%3D4000%26use\_display\_value%3DTRUE](https://service-now.com/api/bebup/config/ci/query/list?encoded_query=install_status%21%3D104%26start%3D2001%26count%3D4000%26use_display_value%3DTRUE)"  
headers =\> {  
Accept =\> "application/json"  
}  
}  
}  
request\_timeout =\> 90

```
    schedule => { cron => "0 0 * * *"}
    socket_timeout => 90
    codec => "json"
    # A hash of request metadata info (timing, response headers, etc.) will be sent here
    # metadata_target => "http_poller_metadata"
}

```

}

---

<div class="post-metadata">

**Author:** ![aaron-nimocks](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/aaron-nimocks/32/73965_2.png) [@aaron-nimocks](https://discuss.elastic.co/u/aaron-nimocks)\
**Post date:** [August 9, 2021, 10:49pm UTC](https://discuss.elastic.co/t/pulling-more-than-2000-records-from-servicenow-api-using-http-poller/280840/2 "2021-08-09T22:49:53Z")

</div>

I would browse to both of those URL's or do a curl to get the results. Compare both of those and see if you are getting duplicates between those. If so the issue is with the source or query.

I can't think of any reasons this pipeline would create duplicate records so I would test that first.

---

<div class="post-metadata">

**Author:** ![Badger](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/badger/32/25190_2.png) [@Badger](https://discuss.elastic.co/u/Badger)\
**Post date:** [August 10, 2021, 3:08am UTC](https://discuss.elastic.co/t/pulling-more-than-2000-records-from-servicenow-api-using-http-poller/280840/3 "2021-08-10T03:08:46Z")

</div>

I cannot find any documentation of this API on the Internet so I would firstly ask if it returns records in order? Do you need to supply a sort option?

Also, in the second one, you use count=4000. Should that be 2000?

You do not say what your output is but if it is elasticsearch then you may be able to handle duplicates by setting the document id to a hash of identifying fields from the events using a fingerprint filter.

Note that if the results are not sorted, then although using a fingerprint filter will avoid duplicates, you still may never get the complete result set. A randomly selected subset from a large group can result in some records never being selected.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [September 7, 2021, 3:09am UTC](https://discuss.elastic.co/t/pulling-more-than-2000-records-from-servicenow-api-using-http-poller/280840/4 "2021-09-07T03:09:43Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
