# Rally failed for downloading data file from aws S3

**URL:** <https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856>\
**Category:** Elasticsearch\
**Tags:** rally\
**Created:** [October 31, 2018, 2:02pm UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856 "2018-10-31T14:02:58Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![mahesh\_varak89](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mahesh_varak89/32/37160_2.png) [@mahesh\_varak89](https://discuss.elastic.co/u/mahesh_varak89)\
**Post date:** [October 31, 2018, 2:02pm UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856/1 "2018-10-31T14:02:58Z")

</div>

It reports an error as following :

* * *

\*\*\*\*\*\* Use this pipeline only if you are aware of the tradeoffs. \*\*\*\*\*\*  
\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\* Watch your step! \*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*\*

* * *

[INFO] Racing on track [company], challenge [index-and-query] and car ['external'] with version [6.2.3].

[WARNING] indexing\_total\_time is 24166 ms indicating that the cluster is not in a defined clean state. Recorded index time metrics may be misleading.  
[WARNING] refresh\_total\_time is 5519 ms indicating that the cluster is not in a defined clean state. Recorded index time metrics may be misleading.

[ERROR] Cannot race. Error in track preparator (('Could not download [[https://s3.amazonaws.com/FOLDER\_PATH/company\_data.json.bz2](https://s3.amazonaws.com/FOLDER_PATH/company_data.json.bz2)] to [/home/ec2-user/.rally/benchmarks/data/rally-tutorial/company\_data.json.bz2] (HTTP status: 403)', None))

Getting further help:

* * *

- Check the log files in /home/ec2-user/.rally/logs for errors.
- Read the documentation at [https://esrally.readthedocs.io/en/1.0.1/](https://esrally.readthedocs.io/en/1.0.1/)
- Ask a question on the forum at [https://discuss.elastic.co/c/elasticsearch/rally](https://discuss.elastic.co/c/elasticsearch/rally)
- Raise an issue at [https://github.com/elastic/rally/issues](https://github.com/elastic/rally/issues) and include the log files in /home/ec2-user/.rally/logs.

* * *

## [INFO] FAILURE (took 1 second)

However, I am able to download data file through CLI (manually).

I am using AWS elasticsearch domain to test the performance and created my own track. It is working fine when I download the data file manually and place it to company folder. But when I use "base-url" property to download the data file automatically from rally it fails to download the data file. The query I am using is as follows:  
CMD\>\> **esrally --pipeline=benchmark-only --track-path=/home/ec2-user/.rally/benchmarks/rally-track/company --target-hosts=https**  
**://vpc-test-rally-es-XXXXXX.us-east-1.es.amazonaws.com**

Here is the log snapshot:

 ![rally_7_1](https://us1.discourse-cdn.com/elastic/original/3X/5/4/549895619e96872e8885acc352e8f9ec9d807561.png)

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [November 1, 2018, 8:31am UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856/2 "2018-11-01T08:31:33Z")

</div>

Hello,

You mentioned:

> It is working fine when I download the data file manually and place it to company folder.

Did you try to download the file manually (e.g. using `curl` or `wget`) in the **same instance** where you are running Rally?

Can you please execute:

```auto
curl -O http://benchmarks.elasticsearch.org.s3.amazonaws.com/corpora/geonames/documents-2.json.bz2

```

on the Rally instance and report back if this is working?

For the record, Rally [invokes a normal urllib3 GET request](https://github.com/elastic/rally/blob/fedf8b12ffba28e455dfb06e44493f626b678fc4/esrally/utils/net.py#L58-L59) to download the track, so there's really no magic involved here; it will honor the `http_proxy` env var, if defined.

Dimitris

---

<div class="post-metadata">

**Author:** ![mahesh\_varak89](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mahesh_varak89/32/37160_2.png) [@mahesh\_varak89](https://discuss.elastic.co/u/mahesh_varak89)\
**Post date:** [November 1, 2018, 8:41am UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856/3 "2018-11-01T08:41:55Z")

</div>

> [@dliappis](#):
>
> curl -O [http://benchmarks.elasticsearch.org.s3.amazonaws.com/corpora/geonames/documents-2.json.bz2](http://benchmarks.elasticsearch.org.s3.amazonaws.com/corpora/geonames/documents-2.json.bz2)

I am able to download the documents-2.json.bz2. Following is the output :

**-rw-rw-r-- 1 ec2-user ec2-user 264698741 Nov 1 08:35 documents-2.json.bz2**

And I am able to download the data file manually by AWS CLI. Following is the sample command:

**aws s3 cp s3://BUCKET\_NAME/FOLDER\_PATH/company\_data.json.bz2 company\_data.json.bz2**

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [November 1, 2018, 9:11am UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856/4 "2018-11-01T09:11:40Z")

</div>

Hi,

I observed that the error you are receiving (in the initial comment) is:

```auto
Racing on track [company], challenge [index-and-query] and car ['external'] with version [6.2.3].

```

i.e. Rally fails to download your own company track (sorry, I was under the impression Rally failed to download the default, geonames, track, that's why I asked you to curl exactly that).

So I presume that your company track resides in an S3 bucket of your own.  
Rally tries to download things using an http URL, which is what you should be defining as base\_url; is your bucket (or object) publicly accessible thought? i.e. if you try to `curl -O https://<url_of_your_s3_bucket/...` (you'll find the URL in the S3 properties) does this get you the company\_data file?

The aws cli tools (`aws s3 cp` etc.) use the `s3://bucket-name/path` schema and IAM Roles and Instance Profiles can be adjusted to grant access to a bucket from within an instance, which may explain why you are able to grab the file using `aws s3 cp`; this doesn't automatically grant access to the object via http, though.

Regards,  
Dimitris

---

<div class="post-metadata">

**Author:** ![mahesh\_varak89](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/mahesh_varak89/32/37160_2.png) [@mahesh\_varak89](https://discuss.elastic.co/u/mahesh_varak89)\
**Post date:** [November 1, 2018, 9:32am UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856/5 "2018-11-01T09:32:20Z")

</div>

> [@dliappis](#):
>
> erved that the error you a

when I am trying to download dta file with **curl** command it through following error :  
e.g. \> **curl -O [https://s3.amazonaws.com/BUCKET\_NAME/FOLDER\_PATH/company\_data.json](https://s3.amazonaws.com/BUCKET_NAME/FOLDER_PATH/company_data.json)**

Error :  
**\<?xml version="1.0" encoding="UTF-8"?\>**  
**`AccessDenied`Access Denied7867ADC755A3F2ECVG4Lhh2B4veFd8bhtUk1E8Ew/Kg/CCPCSGDdydMhm5ArinXBBaGnY9bLXAWFXs5ZgEsqPPWOJlQ=**

According to above error do I need to make my object public?

---

<div class="post-metadata">

**Author:** ![dliappis](https://sea2.discourse-cdn.com/elastic/user_avatar/discuss.elastic.co/dliappis/32/56174_2.png) [@dliappis](https://discuss.elastic.co/u/dliappis)\
**Post date:** [November 1, 2018, 9:57am UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856/6 "2018-11-01T09:57:36Z")

</div>

> curl -O [https://s3.amazonaws.com/BUCKET\_NAME/FOLDER\_PATH/company\_data.json](https://s3.amazonaws.com/BUCKET_NAME/FOLDER_PATH/company_data.json)

This doesn't look right a correct http URL for an s3 bucket (the s3:// schema is used by the aws cli command only).

Please refer to: [Buckets overview - Amazon Simple Storage Service](https://docs.aws.amazon.com/AmazonS3/latest/dev/UsingBucket.html#access-bucket-intro)

to see how to get the URL of the bucket+object and how to make it public.

---

<div class="post-metadata">

**Author:** ![system](https://us1.discourse-cdn.com/elastic/original/3X/1/a/1ac57faf039f6b580b3f104ef42a2a89e41014de.png) [@system](https://discuss.elastic.co/u/system)\
**Post date:** [November 29, 2018, 9:57am UTC](https://discuss.elastic.co/t/rally-failed-for-downloading-data-file-from-aws-s3/154856/7 "2018-11-29T09:57:38Z")

</div>

This topic was automatically closed 28 days after the last reply. New replies are no longer allowed.
