| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| cluster-file_transfer [2026/09/03 10:28] – gabriele | cluster-file_transfer [2026/09/03 12:58] (current) – gabriele |
|---|
| |
| ===== 2. Transferring data externally ===== | ===== 2. Transferring data externally ===== |
| To transfer data from outside RHUL, you'd need to first move the data to a separate ''graphics'' server before moving them to ''/MRIWork''. This is for security reasons. Data will scanned by an antivirus as you move them. From this ''graphics'' server, you can then easily copy them in your folder within ''/MRIWork''. | To transfer data from outside RHUL, you'd need to first move the data to a separate "graphics" server before moving them to ''/MRIWork''. This is for security reasons. Data will scanned by an antivirus as you move them. From this "graphics" server, you can then easily copy them in your folder within ''/MRIWork''. |
| |
| This ''graphics'' server has IP: 134.219.34.200 and is called ''graphics'' because it has a graphics card that allows users to connect via any remote desktop app of their choice (e.g., [[https://apps.microsoft.com/detail/9n1f85v9t8bn?hl=en-US&gl=GB|Windows App]]). The ''graphics'' server is managed by Jonas Larsson (jonas.larsson@rhul.ac.uk). Please contact Jonas to have him create an account for you on this server. | This "graphics" server has IP: 134.219.34.200 and is called ''graphics'' because it has a graphics card that allows users to connect via any remote desktop app of their choice (e.g., [[https://apps.microsoft.com/detail/9n1f85v9t8bn?hl=en-US&gl=GB|Windows App]]). The "graphics" server is managed by Jonas Larsson (jonas.larsson@rhul.ac.uk). Please contact Jonas to have him create an account for you on this server. |
| |
| |
| ==== Examples: ==== | ==== Example 1: ==== |
| |
| Suppose you need to download a dataset from a website that does not require a data sharing agreement and is a simple browser-based download, like [[https://datadryad.org/dataset/doi:10.5061/dryad.2v6wwpzt3|UCL - Release of cognitive and multimodal MRI data including real-world tasks and hippocampal subfield segmentations]]. | Suppose you need to download a dataset from a website that does not require a data sharing agreement and is a simple browser-based download, like [[https://datadryad.org/dataset/doi:10.5061/dryad.2v6wwpzt3|UCL - Release of cognitive and multimodal MRI data including real-world tasks and hippocampal subfield segmentations]]. |
| - Enter your username: gabriele (for your it's the name Jonas gave you, most likely first letter of your name + your surname) | - Enter your username: gabriele (for your it's the name Jonas gave you, most likely first letter of your name + your surname) |
| - Enter your password: **** (the one Jonas gave you to connect the first time, which need to be changed at first login) | - Enter your password: **** (the one Jonas gave you to connect the first time, which need to be changed at first login) |
| - Once connected, click on the World on the bar at the bottom of the screen. This will open a Firefox webpage | - Once connected, click on the World icon on the bar at the bottom of the screen. This will open a Firefox webpage |
| - Search your dataset page like you'd do it normally on your browser | - Search your dataset page like you'd do it normally on your browser |
| - Locate your dataset. | - Locate your dataset. |
| - Choose your local folder where to extract the files. I suggest you create here the folder structure you'd like your data to have in ''/MRIWork'' for your future analyses. | - Choose your local folder where to extract the files. I suggest you create here the folder structure you'd like your data to have in ''/MRIWork'' for your future analyses. |
| |
| | From the graphics server you can't log into psychp01 directly, but the same storage is mounted read-only. You'll see a ''/MRIWork'' folder with exactly the same contents. It isn't a copy, it's the same files, accessed over the network. Because the mount is read-only, you can't create, edit, or delete anything on psychp01 from the "graphics" server. You can only read from it and copy data across to the "graphics" server's own storage. This is convenient, as you can create and test your environment here before moving your data onto psychp01. You can also test-run some analyses, as you'll have the same software (Jonas is taking care that the "graphics" server has all the software psychp01 has), but please be aware of the warning below. |
| Now, from the ''graphics'' server, you don't have access to your '/MRIWork' folder on psychp01 (that's way is "safer"), but all your data and folder on pyschp01 are "mounted" (kind of copied) on the ''graphics'' server. This means that the environment within it will look exactly the same as your psychp01 environment. That is, you'll have also here a folder called "MRIWork" and you'll see whatever you have on pyschp01. This is convenient, as you can create a test your environment here before moving your data onto psychp01. You can also test-run some analyses, as you'll have the same software (Jonas is taking care that the ''graphics'' server has all the software psychp01 has). | |
| |
| Once you have downloaded all your data, created your folder structure and potentially tested some analyses, you can pull the data from psychp01 for storage and additional analyses. | Once you have downloaded all your data, created your folder structure and potentially tested some analyses, you can pull the data from psychp01 for storage and additional analyses. |
| |
| To move the data to psychp01. Suppose you have saved all your data on the ''graphics'' server in a folder with path: ''/home/gabriele/Downloads/new_dataset'' and you want to move the all ''new_dataset'' folder to '/MRIWork/MRIWork25/gb/gabriele_bellucci/' on psychp01. | == To move the data from the "graphics" server onto psychp01. == |
| - Establish an ssh connection with pyshcp01 via the terminal (see). | Suppose you have saved all your data on the "graphics" server in a folder with path: ''/home/gabriele/Downloads/new_dataset'' and you want to move the ''new_dataset'' folder to '/MRIWork/MRIWork25/gb/gabriele_bellucci/' on psychp01. |
| - Establish an sftp connection with the ''graphics'' server: | - Establish an ssh connection with pyshcp01 via the terminal (see [[cluster-access#log_in_with_ssh|SSH]]). |
| | - Establish an sftp connection with the "graphics" server (see [[cluster-file_transfer#sftp|SFTP]] below): |
| sftp gabriele@134.219.34.200 | sftp gabriele@134.219.34.200 |
| - You'll be prompted to enter your password. | - You'll be prompted to enter your password. |
| |
| |
| Remember: the ''graphics'' server is simply a computer sitting in Wolfson, which is NOT backed up and should NOT be used for long-term storage or important data analyses. If there's an outage (which unfortunately happens pretty often), the ''graphics'' server will be severely impacted and you may use all your data. | **WARNING** |
| | Remember: the "graphics" server is simply a computer sitting in Wolfson, which is NOT backed up and should NOT be used for long-term storage or important data analyses. If there's an outage (which unfortunately happens pretty often), the "graphics" server will be severely impacted and you may use all your data. |
| | |
| | |
| | |
| | ==== Example 2: ==== |
| | |
| | Suppose you need to download a dataset that **cannot** be downloaded by simply clicking a link in the browser, because it is hosted on a repository that uses ''git-annex''. A good example is the [[https://doi.gin.g-node.org/10.12751/g-node.5mv3bf/|Welsh Advanced Neuroimaging Database (WAND)]], hosted on [[https://gin.g-node.org/CUBRIC/WAND|GIN]] by Cardiff University: 170 participants, 8 imaging sessions, roughly **1.95 TB** in total. |
| | |
| | Repositories like GIN do not store the data in one downloadable archive. Instead, you first download a small "skeleton" of the dataset (the folder structure and the file names, a few hundred MB), and then ask for the content of only the files you actually need. This is very convenient: you can look at the whole dataset, decide what you want, and download only that. |
| | |
| | Everything below is done **on the "graphics" server** (see above for how to connect), following the same logic as Example 1: external data must land on "graphics" first, and can then be moved to ''/MRIWork'' on psychp01. |
| | |
| | == Before you start == |
| | |
| | Three things to know, because none of them are obvious and none of them are in GIN's own documentation: |
| | |
| | * You **cannot** download this dataset over HTTPS. GIN serves the folder structure over HTTPS, but the actual imaging files only over ''ssh''. You therefore need a (free) GIN account with your own ''ssh'' key. There is no way around this. |
| | * The ''ssh'' address needs a **leading slash** after the colon. Without it, the server replies ''GIN: Invalid repository path''. |
| | * ''git'', ''git-annex'' and the ''gin'' client are already installed on the "graphics" server for all users. You do not need to install anything. |
| | |
| | == 1. Set your git identity == |
| | |
| | ''git-annex'' writes small commits as it works and refuses to run without a name and an email address: |
| | |
| | <code> |
| | git config --global user.name "Your Name" |
| | git config --global user.email "your.name@rhul.ac.uk" |
| | </code> |
| | |
| | These are only labels written into your local copy. They are not checked against anything, and they do not give anyone access to anything. |
| | |
| | == 2. Register a GIN account == |
| | |
| | This step is done in the browser (click on the World icon at the bottom of the screen to open Firefox). |
| | |
| | - Go to [[https://gin.g-node.org|https://gin.g-node.org]] |
| | - Click ''Register'' (top right) |
| | - Choose a username, enter your RHUL email and a password |
| | - Confirm your account by clicking the link in the email you receive |
| | |
| | Registration is free and immediate; there is no approval process and no data sharing agreement for this dataset. |
| | |
| | == 3. Create an ''ssh'' key and upload it to GIN == |
| | |
| | An ''ssh'' key comes in two halves: a **private** key that never leaves the "graphics" server, and a **public** key that you give to GIN so it can recognise you. |
| | |
| | First check whether you already have one: |
| | |
| | <code> |
| | ls ~/.ssh/id_*.pub |
| | </code> |
| | |
| | If nothing is listed, create one: |
| | |
| | <code> |
| | ssh-keygen -t ed25519 -C "your.name@rhul.ac.uk" |
| | </code> |
| | |
| | Press Enter to accept the default file name. You can leave the passphrase empty; if you set one, you will be asked for it every time, which is inconvenient for downloads that run for hours. |
| | |
| | Now print the **public** key (note the ''.pub'' — never share the file without it): |
| | |
| | <code> |
| | cat ~/.ssh/id_ed25519.pub |
| | </code> |
| | |
| | You will see a single long line starting with ''ssh-ed25519''. Select and copy the **whole** line, including the email at the end. |
| | |
| | Then, in the browser, logged into GIN: |
| | |
| | - Click your avatar (top right) and choose ''Your Settings'' |
| | - Choose ''SSH Keys'' in the menu on the left |
| | - Click ''Add Key'' |
| | - Give it a name, e.g. ''graphics server'' |
| | - Paste the line into the ''Content'' box and save |
| | |
| | Check that it worked: |
| | |
| | <code> |
| | ssh -T git@gin.g-node.org |
| | </code> |
| | |
| | The first time, you will be asked whether you trust the server. Type ''yes''. You should then see: |
| | |
| | <code> |
| | Hi there, You've successfully authenticated, but GIN does not provide shell access. |
| | </code> |
| | |
| | This message means **success**. GIN only allows ''git'' operations, never an interactive login, so "no shell access" is the expected and correct answer. |
| | |
| | == 4. Choose where the data will go == |
| | |
| | **Do not download into your home directory.** Use ''Storage 1'' or ''Storage 2'' on the "graphics" server (roughly 10 TB each), exactly as in Example 1. |
| | |
| | Check the free space before you start: |
| | |
| | <code> |
| | df -h /path/to/storage |
| | cd /path/to/storage |
| | </code> |
| | |
| | Note that ''/MRIWork'' is mounted **read-only** on the "graphics" server, so it cannot be used as the download destination. The data go into ''Storage 1'' or ''Storage 2'' first, and are moved to ''/MRIWork'' afterwards (see below). |
| | |
| | == 5. Download the dataset skeleton == |
| | |
| | <code> |
| | git clone git@gin.g-node.org:/CUBRIC/WAND.git |
| | cd WAND |
| | git annex init "graphics server" |
| | </code> |
| | |
| | **Note the slash** immediately after the colon, before ''CUBRIC''. This is the single most common cause of failure and it is not documented by GIN. |
| | |
| | This step is quick and small. You now have the complete folder structure with the real file names, all the metadata files (''.json'', ''.tsv'', the README, ''participants.tsv''), and //placeholders// where the imaging files will go. |
| | |
| | You can check the overall picture with: |
| | |
| | <code> |
| | git annex info |
| | </code> |
| | |
| | which reports the total size of the dataset, how much you currently have locally (zero at this point), and how much disk space is available. |
| | |
| | == 6. Check how big your request is, before downloading == |
| | |
| | ''git-annex'' knows the size of every file without downloading anything. This lets you find out in advance whether your selection will fit. For example, to add up sessions 02, 03 and 06 across all participants: |
| | |
| | <code> |
| | for s in 02 03 06; do |
| | printf "ses-%s: " "$s" |
| | git annex find sub-*/ses-$s --format='${bytesize}\n' 2>/dev/null \ |
| | | awk '{t+=$1} END {printf "%.1f GB\n", t/1e9}' |
| | done |
| | </code> |
| | |
| | Change the list ''02 03 06'' to the sessions you need. Compare the result with the output of ''df -h .'' **before** starting a transfer, not halfway through one. |
| | |
| | == 7. Download the data you need == |
| | |
| | Always test with a single participant first, so you can see how much one subject costs: |
| | |
| | <code> |
| | git annex get sub-00395/ses-03/anat |
| | du -sh . |
| | </code> |
| | |
| | Then start the real download. Use ''screen'' so that the transfer survives a lost connection or a closed remote desktop session: |
| | |
| | <code> |
| | screen -S wand |
| | git annex get -J4 sub-*/ses-0{2,3,6} |
| | </code> |
| | |
| | ''-J4'' runs four downloads in parallel. Detach from the ''screen'' session with ''Ctrl-A'' then ''D'', and come back to it later with ''screen -r wand''. |
| | |
| | If the transfer is interrupted, simply run the same command again: ''git-annex'' keeps track of what it already has and continues where it stopped. |
| | |
| | You will see error messages for participants who do not have a given session. This is normal: not every volunteer took part in every session (the 7 T and TMS sessions in particular had much smaller sub-samples). |
| | |
| | To download only some participants, put their IDs in a text file, one per line: |
| | |
| | <code> |
| | while read s; do |
| | git annex get -J4 "$s"/ses-0{2,3,6} |
| | done < subjects.txt |
| | </code> |
| | |
| | == 8. Move the data onto psychp01 == |
| | |
| | Once the download is finished and you are happy with your folder structure, move the data to ''/MRIWork'' exactly as described above in [[cluster-file_transfer#to_move_the_data_from_the_graphics_server_onto_psychp01|To move the data from the "graphics" server onto psychp01]]. |
| | |
| | Note that the downloaded dataset contains a hidden ''.git'' folder holding the ''git-annex'' machinery. If you copy the whole ''WAND'' folder, it comes along and roughly doubles the space used. If you only want the imaging files on psychp01, copy the ''sub-*'' folders and the metadata files, and leave the repository behind on the "graphics" server. |
| | |
| | == Structure of the WAND dataset == |
| | |
| | The data follow the [[https://bids.neuroimaging.io/|BIDS]] standard: first participant, then session, then data type, i.e. ''sub-<ID>/ses-<NN>/<datatype>/''. |
| | |
| | ^ Session ^ Content ^ |
| | | ses-01 | MEG (CTF ''.ds'' folders) | |
| | | ses-02 | Connectom 3 T, ultra-strong gradients: diffusion and quantitative MRI | |
| | | ses-03 | Prisma 3 T: structural, functional, perfusion | |
| | | ses-04 | 7 T spectroscopy | |
| | | ses-05 | 3 T GABA-edited spectroscopy (MEGA-PRESS) | |
| | | ses-06 | 7 T structural and functional | |
| | | ses-07 | 3 T metabolic (subset, about 39 participants) | |
| | | ses-08 | TMS (subset, about 40 participants) | |
| | |
| | ''ses-02'' and ''ses-06'' are by far the largest. |
| | |
| | == Common pitfalls == |
| | |
| | * **The files look like they are already there, but they are not.** After cloning, you can see and browse every file name, including the MEG ''.ds'' folders. Those are placeholders until you run ''git annex get''. Analysis software pointed at a dataset that has not been downloaded will usually report //corrupt data// rather than //missing files//, which is confusing. To list what is still missing in the current folder: ''git annex find . %%--%%not %%--%%in here'' |
| | * **Never run ''git annex get'' without a path.** With no path it means "download everything", i.e. 1.95 TB. Do that only if you want to download the full dataset. The same applies to ''gin sync %%--%%content''. |
| | * **The download links on the GIN wiki are dead.** The ''gin'' client is now distributed through [[https://github.com/G-Node/gin-cli/releases|GitHub releases]]. It is already installed on the "graphics" server, and in practice you can do everything with plain ''git'' and ''git annex'' anyway. |
| | * **Remember the warning above**: the "graphics" server is not backed up. Do not leave the only copy of anything there. |
| |
| |