# How to download the data with "wget or curl"?

**URL:** <https://forum.depmap.org/t/how-to-download-the-data-with-wget-or-curl/1642>\
**Category:** Q&A\
**Created:** [August 18, 2022, 7:08pm UTC](https://forum.depmap.org/t/how-to-download-the-data-with-wget-or-curl/1642 "2022-08-18T19:08:35Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![Shicheng\_Guo](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.depmap.org/shicheng_guo/32/1121_2.png) [@Shicheng\_Guo](https://forum.depmap.org/u/Shicheng_Guo)\
**Post date:** [August 18, 2022, 7:08pm UTC](https://forum.depmap.org/t/how-to-download-the-data-with-wget-or-curl/1642/1 "2022-08-18T19:08:35Z")

</div>

Dear Team,

I am trying to download the full dataset with some command line like “wget or curl”. Is there anyway to do it rather than manfully download the file one-by-one and then upload the some Linux servers?

Thanks.

Shicheng

CRISPR\_gene\_effect.csv  
CRISPR\_gene\_dependency.csv  
CCLE\_expression.csv  
CCLE\_gene\_cn.csv  
CCLE\_wes\_gene\_cn.csv  
CCLE\_mutations.csv  
sample\_info.csv  
Achilles\_metadata.csv  
Achilles\_gene\_effect.csv  
Achilles\_gene\_effect\_uncorrected.csv  
Achilles\_gene\_dependency.csv  
Achilles\_common\_essentials.csv  
Achilles\_guide\_efficacy.csv  
Achilles\_cell\_line\_efficacy.csv  
Achilles\_cell\_line\_growth\_rate.csv  
CRISPR\_dataset\_sources.csv  
CRISPR\_gene\_effect.csv  
CRISPR\_gene\_dependency.csv  
CRISPR\_common\_essentials.csv  
common\_essentials.csv  
nonessentials.csv  
Achilles\_raw\_readcounts.csv  
Achilles\_raw\_readcounts\_failures.csv  
Achilles\_logfold\_change.csv  
Achilles\_logfold\_change\_failures.csv  
Achilles\_guide\_map.csv  
Achilles\_replicate\_map.csv  
Achilles\_replicate\_QC\_report\_failing.csv  
Achilles\_dropped\_guides.csv  
Achilles\_high\_variance\_genes.csv  
CCLE\_RNAseq\_reads.csv  
CCLE\_expression\_full.csv  
CCLE\_expression.csv  
CCLE\_expression\_transcripts\_expected\_count.csv  
CCLE\_expression\_proteincoding\_genes\_expected\_count.csv  
CCLE\_RNAseq\_transcripts.csv  
CCLE\_segment\_cn.csv  
CCLE\_wes\_segment\_cn.csv  
CCLE\_gene\_cn.csv  
CCLE\_wes\_gene\_cn.csv  
CCLE\_fusions.csv  
CCLE\_fusions\_unfiltered.csv  
CCLE\_mutations.csv  
CCLE\_mutations\_bool\_hotspot.csv  
CCLE\_mutations\_bool\_damaging.csv  
CCLE\_mutations\_bool\_nonconserving.csv  
CCLE\_mutations\_bool\_otherconserving.csv

---

<div class="post-metadata">

**Author:** ![james-synthego](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.depmap.org/james-synthego/32/1027_2.png) [@james-synthego](https://forum.depmap.org/u/james-synthego)\
**Post date:** [February 3, 2024, 12:03am UTC](https://forum.depmap.org/t/how-to-download-the-data-with-wget-or-curl/1642/2 "2024-02-03T00:03:21Z")

</div>

Yes it’s possible, one can figure it out using Google Chrome’s Developer Tools’ Network tab.

Here is some Python 3.10 code (made on 2/2/2024) that shows how to download a few files:

```python
import asyncio
import os
import urllib.parse

import httpx # ==0.26.0

THIS_DIR = os.path.dirname( __file__ )
DEPMAP_FIGSHARE_FILES_BASE_URL = (
    "https://ndownloader.figshare.com/files/"
)
DEPMAP_FIGSHARE_DATASET_IDS: dict[str, int] = {
    "Model_v2.csv": 43746708
}
DEPMAP_API_BASE_URL = "https://depmap.org/portal/api"
DEPMAP_CUSTOM_DATASET_IDS: dict[str, str] = {
    "Expression_Public_23Q4": "expression",
    "Proteomics": "proteomics",
}

def _dump_to_file(
    response: httpx.Response, target_file: str | os.PathLike
) -> None:
    with open(target_file, "wb") as f:
        for chunk in response.iter_bytes(chunk_size=8192):
            f.write(chunk)

async def download_file(
    depmap_id: int | str,
    target_file: str | os.PathLike,
    download_timeout: float | None = 60 * 10.0,
) -> None:
    """Download a DepMap file given its ID to a local file."""
    async with httpx.AsyncClient() as client:
        response = await client.get(
            f"{DEPMAP_FIGSHARE_FILES_BASE_URL}/{depmap_id}",
            follow_redirects=True,
            timeout=download_timeout,
        )
        response.raise_for_status()
    await asyncio.to_thread(_dump_to_file, response, target_file)

async def download_custom_dataset(
    depmap_id: str,
    target_dir: str | os.PathLike = THIS_DIR,
    add_cell_line_metadata: bool = False,
    preparation_wait_interval: float = 10.0,
    download_timeout: float | None = 60 * 10.0,
) -> None:
    """
    Download a DepMap custom dataset given its ID to a local directory.

    Args:
        depmap_id: DepMap ID for the custom dataset to download.
        target_dir: Target directory to download the file to.
        add_cell_line_metadata: Set True to include cell line metadata
            columns in the CSV (e.g. DepMap ID, cell line display name).
        preparation_wait_interval: Sleep interval (sec) while DepMap
            prepares the download.
        download_timeout: Optional timeout (sec) on the file download.
    """
    async with httpx.AsyncClient() as client:
        response = await client.post(
            f"{DEPMAP_API_BASE_URL}/download/custom",
            json={
                "datasetId": depmap_id,
                "dropEmpty": True,
                "addCellLineMetadata": add_cell_line_metadata,
            },
            timeout=15.0,
        )
        response.raise_for_status()
        task_id: str = response.json()["id"]
        while True:
            response = await client.get(
                f"{DEPMAP_API_BASE_URL}/task/{task_id}", timeout=15.0
            )
            response.raise_for_status()
            response_data = response.json()
            match response_data["state"]:
                case "PROGRESS":
                    await asyncio.sleep(preparation_wait_interval)
                case "SUCCESS":
                    download_url: str = response_data["result"][
                        "downloadUrl"
                    ]
                    break
                case _:
                    raise NotImplementedError(
                        f"Unexpected state {response_data['state']} for"
                        " task {task_id}."
                    )
        query_params: dict[str, str] = dict(
            q.split("=", maxsplit=1)
            for q in urllib.parse.urlparse(download_url).query.split(
                "&"
            )
        )
        target_file = os.path.join(
            target_dir, query_params.pop("name")
        )
        response = await client.get(
            download_url,
            timeout=download_timeout,
            params=query_params,
        )
        response.raise_for_status()
    await asyncio.to_thread(_dump_to_file, response, target_file)

```

---

<div class="post-metadata">

**Author:** ![pmontgom](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.depmap.org/pmontgom/32/4_2.png) [@pmontgom](https://forum.depmap.org/u/pmontgom)\
**Post date:** [March 20, 2024, 6:57pm UTC](https://forum.depmap.org/t/how-to-download-the-data-with-wget-or-curl/1642/3 "2024-03-20T18:57:37Z")

</div>

Another approach, which also requires a little scripting, is to download files directly from figshare.

To find a release on Figshare, you can click “View full release details” and you should be shown a screen where we list how to cite the dataset:

 ![image](https://canada1.discourse-cdn.com/flex035/uploads/depmap/original/2X/0/007772f128356e37f75a7334037e59c4b3b65d24.png)

Clicking the link will take you to the dataset released on Figshare, and figshare [has a nice API in addition to it’s UI for downloading files](https://docs.figshare.com/old_docs/api/articles/#download-public-files).

So, after navigating to figshare, we can see the article ID for the DepMap 23Q4 release is “24667905” and we’re currently on version “2” by looking at the url you’re directed to:

```auto
https://plus.figshare.com/articles/dataset/DepMap_23Q4_Public/24667905/2

```

Now we can use that information to request all URLs for all files:

```auto
$ curl https://api.figshare.com/v2/articles/24667905/versions/2

{"url_public_html": 
   "https://plus.figshare.com/articles/dataset/DepMap_23Q4_Public/24667905/2", 
   "files": [
      {"id": 43347678, 
       "name": "README.txt", 
       "size": 29103, 
       "is_link_only": false, 
       "download_url": "https://ndownloader.figshare.com/files/43347678", ...

```

You can then use this response by looping through the `files` field, and download each `download_url` to a file with the given `name` field.

Here’s an example in python doing this:

```auto
import requests
import subprocess

article_id="24667905"
version="2"

article = requests.get(f"https://api.figshare.com/v2/articles/{article_id}/versions/{version}").json()
for file in article["files"]:
    print(f'downloading {file["name"]} from {file["download_url"]}')
    subprocess.run(["curl", file["download_url"], "-o", file["name"]], check=True)

```

---

<div class="post-metadata">

**Author:** ![Shicheng\_Guo](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.depmap.org/shicheng_guo/32/1121_2.png) [@Shicheng\_Guo](https://forum.depmap.org/u/Shicheng_Guo)\
**Post date:** [March 20, 2024, 7:33pm UTC](https://forum.depmap.org/t/how-to-download-the-data-with-wget-or-curl/1642/4 "2024-03-20T19:33:17Z")

</div>

> [@pmontgom](#):
>
> ```auto
> https://plus.figshare.com/articles/dataset/DepMap_23Q4_Public/24667905/2
> 
> ```

Thank you!! I think it exactly solve my problem as below!

 ![image](https://canada1.discourse-cdn.com/flex035/uploads/depmap/original/2X/0/09a090f0440c8cd7f83d6e8dac28cbf1a0d15022.png)

---

<div class="post-metadata">

**Author:** ![pmontgom](https://yyz1.discourse-cdn.com/flex035/user_avatar/forum.depmap.org/pmontgom/32/4_2.png) [@pmontgom](https://forum.depmap.org/u/pmontgom)\
**Post date:** [March 20, 2024, 8:33pm UTC](https://forum.depmap.org/t/how-to-download-the-data-with-wget-or-curl/1642/5 "2024-03-20T20:33:56Z")

</div>

Yes, I forgot you can also just click the “download all” button once you’re at figshare. 🙂
