User Tools

Site Tools


howto:bioinformatics

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revision Previous revision
Next revision
Previous revision
howto:bioinformatics [2025/05/01 23:10]
nmfuphi [Software Environments]
howto:bioinformatics [2025/05/21 10:50] (current)
nmfuphi [Singularity]
Line 194: Line 194:
 </code> </code>
  
-==== Software Environments ====+===== Software Environments =====
  
-=== Singularity ===+==== Conda ==== 
 +Many scientific software tools rely on specific versions of libraries, compilers, and dependencies that often conflict with each other or with system-wide installations. **Conda** is a powerful, language-agnostic environment and package manager that helps solve this problem by allowing users to manage **Python**, **R**, **C/C++**, **FORTRAN**, and other language ecosystems in **isolated environments**. 
 + 
 +=== Shared Conda Environments === 
 + 
 +For most use cases, especially in bioinformatics, CHPC provides pre-built, **shared Conda environments** installed under: 
 + 
 +  '/apps/chpc/bio/anaconda3-2020.02/envs' 
 + 
 +These environments are curated by CHPC staff to include commonly used tools in genomics, transcriptomics, and other domains. 
 + 
 +=== Step-by-step usage === 
 + 
 +=== Step 1: Load required modules === 
 +To access Conda functionality, first load the required modules: 
 +<code bash> 
 +module load chpc/BIOMODULES 
 +module load conda_init 
 +</code> 
 + 
 +The second module updates your .bashrc file by adding necessary shell variables. 
 +To apply these changes, you can either log out and log back in, or run: <code bash>source ~/.bashrc. </code> 
 +After this setup, you won’t need to load additional modules for your jobs—only the eval and conda activate steps are required. 
 + 
 +=== Step 2: Initialize Conda in your shell === 
 +Activate Conda shell integration: 
 +<code bash> 
 +eval "$(conda shell.bash hook)" 
 +</code> 
 + 
 +This command sets up your shell environment to recognize Conda commands like `conda activate`. 
 + 
 +=== Step 3: List available environments === 
 +<code bash> 
 +conda info --envs 
 +</code> 
 + 
 +This will display all available shared Conda environments and their paths. 
 + 
 +=== Step 4: Activate a shared environment === 
 +<code bash> 
 +conda activate nameOfTheEnv 
 +</code> 
 + 
 +Replace ''nameOfTheEnv'' with the name of an environment from the previous step. 
 + 
 +> **Tip:** If you're unsure which environment to use, contact CHPC support or explore the environment's contents with `conda list`. 
 + 
 +> **Note:** You do **not** need to install anything when using shared environments. 
 + 
 +=== Creating Private Conda Environments === 
 + 
 +If you need software that is not included in the shared environments, you may create your own **private Conda environment**. This gives you full control over the software stack and package versions. 
 + 
 +> **Important:** Do **not** install environments in your home directory (''/home/<username>'') -use your Lustre project storage instead. 
 + 
 +=== Step-by-step setup === 
 +ssh to username@scp.chpc.ac.za, the password is the same as the one you use on lengau 
 +=== Step 1: Load Conda === 
 +<code bash> 
 +module load chpc/BIOMODULES 
 +module load conda_init 
 +eval "$(conda shell.bash hook)" 
 +</code> 
 + 
 +=== Step 2: Create a new environment === 
 +<code bash> 
 +conda create --prefix /mnt/lustre/<username>/myenv python=3.10 
 +</code> 
 + 
 +This will create a Conda environment at the specified path with Python 3.10 installed. You can replace the Python version or leave it out if not needed. 
 + 
 +=== Step 3: Activate your environment === 
 +<code bash> 
 +conda activate /mnt/lustre/<username>/myenv 
 +</code> 
 + 
 +After activation, you can install any packages you need. 
 + 
 +=== Step 4: (Optional) Install Mamba for faster package management === 
 +<code bash> 
 +conda install mamba -n base -c conda-forge 
 +</code> 
 + 
 +> **Tip:** Mamba is a drop-in replacement for Conda that uses a faster dependency solver written in C++. Once installed, you can use `mamba` instead of `conda` for installing packages: 
 +<code bash> 
 +mamba install numpy pandas 
 +</code> 
 + 
 +This significantly speeds up installations and environment solves, especially when working with large scientific packages. 
 + 
 + 
 +=== Step 5: Install packages === 
 +<code bash> 
 +conda install numpy pandas matplotlib 
 +</code> 
 + 
 +You can install packages one by one, or include them during environment creation: 
 +<code bash> 
 +conda create --prefix /mnt/lustre/<username>/myenv python=3.10 numpy pandas 
 +</code> 
 + 
 +=== Step 6: Remove unused environments === 
 +Old or unused environments can be removed to free up space: 
 +<code bash> 
 +conda remove --prefix /mnt/lustre/<username>/myenv --all 
 +</code> 
 + 
 +=== Best Practices === 
 +  * ✅ Use **shared environments** whenever possible for consistency and faster setup. 
 +  * 📁 Create private environments **only in Lustre** directories, such as ''/mnt/lustre/<username>''
 +  * ⚠️ Do **not** use Conda in your ''$HOME'' directory, it may lead to quota issues or slow performance. 
 +  * 📌 Use the ''--prefix'' flag to create environments with absolute paths, especially on clusters where ''--name'' may default to ''$HOME''
 +  * 🧼 Periodically clean up unused environments with `conda remove --all`. 
 +  * 🔁 Reuse environment definitions by exporting and sharing them with others or for reproducibility. 
 + 
 +=== Troubleshooting === 
 +  * ❓ **Conda not recognized?** Make sure you loaded `conda_init` and ran `eval "$(conda shell.bash hook)"`. 
 +  * 🚫 **Permission denied?** You might be trying to write to a restricted directory like ''/apps'' or ''$HOME''
 +  * 🔄 **Environment behaving unexpectedly?** Try deactivating (`conda deactivate`) and reactivating, or recreate the environment. 
 +  * 🧪 **Conflicts during install?** Use `conda clean --all` to clear caches and retry with a minimal environment. 
 + 
 +==== Singularity ====
  
 Singularity is an open-source, cross-platform container platform specifically designed for scientific and high-performance computing environments. It prioritizes reproducibility, portability, and security, all essential for scientific workflows. Singularity enables users to package entire workflows, including software, libraries, and environment settings, into a single container image. This ensures consistent application execution across various systems without modification. This capability simplifies the migration of complex computational environments and supports reproducible research practices. [[https://docs.sylabs.io/guides/3.7/user-guide/|A detailed Singularity user guide is available here]]. Singularity is an open-source, cross-platform container platform specifically designed for scientific and high-performance computing environments. It prioritizes reproducibility, portability, and security, all essential for scientific workflows. Singularity enables users to package entire workflows, including software, libraries, and environment settings, into a single container image. This ensures consistent application execution across various systems without modification. This capability simplifies the migration of complex computational environments and supports reproducible research practices. [[https://docs.sylabs.io/guides/3.7/user-guide/|A detailed Singularity user guide is available here]].
Line 233: Line 355:
 ⚠️ **Important**: Singularity image files can be very large and may consume significant storage in your Lustre or project directory. Please remove any images you no longer need to help conserve shared storage resources. ⚠️ **Important**: Singularity image files can be very large and may consume significant storage in your Lustre or project directory. Please remove any images you no longer need to help conserve shared storage resources.
  
-If you're unsure or need assistance, log a support ticket here:   
-👉 https://users.chpc.ac.za/helpdesk/tickets/submit/ 
  
 **To pull Singularity images from public container registries (like DockerHub), follow these steps:** **To pull Singularity images from public container registries (like DockerHub), follow these steps:**
Line 342: Line 462:
 #PBS -N singularity_job #PBS -N singularity_job
 #PBS -q normal #PBS -q normal
-#PBS -l select=1:ncpus=8:mem=32gb+#PBS -l select=1:ncpus=24
 #PBS -l walltime=12:00:00 #PBS -l walltime=12:00:00
 #PBS -o singularity_output.log #PBS -o singularity_output.log
Line 359: Line 479:
 </code> </code>
  
 +==== Nextflow ====
  
-== Need Help? ==+Nextflow is a free and open-source workflow management system that enables the development and execution of data analysis pipelines. It simplifies complex computational workflows and ensures reproducibility, scalability, and portability—whether you’re working on a laptop, HPC cluster, or in the cloud. Nextflow workflows are written using DSL2, allowing modular code design and seamless integration with container technologies like Docker, Singularity, Conda, or manual installations. 
 +[[https://www.nextflow.io/docs/latest/index.html|Official documentation is available here]].
  
-Encountering errors? Submit a ticket:+=== Running Nextflow on the CHPC Cluster ===
  
-👉 https://users.chpc.ac.za/helpdesk/tickets/submit/+CHPC supports Nextflow workflows through Singularity containersSince compute nodes have no internet access, all dependencies must be downloaded in advance on the login node.
  
-**Include:**+=== 1. Connect to the Login Node === 
 + 
 +Log into the CHPC login node using your Lengau credentials: 
 + 
 +<code bash> ssh username@scp.chpc.ac.za </code> 
 + 
 +Use this session to prepare your workflow and submit jobs. 
 + 
 +=== 2. Load the Nextflow Module === 
 + 
 +Load the necessary environment modules: 
 + 
 +<code bash>module load chpc/BIOMODULES nextflow 
 +module load chpc/singularity 
 +</code> 
 + 
 +Note: Modules must be reloaded in every new session unless added to your ~/.bashrc. 
 + 
 +=== 3. Pull Workflow and Dependencies === 
 + 
 +Pull your workflow and dependencies on the login1 node: 
 + 
 +Pull the workflow: 
 +<code bash> nextflow pull nf-core/rnaseq </code> 
 + 
 +Run a test execution: 
 +<code bash> nextflow run nf-core/rnaseq -profile test </code> 
 + 
 +This will: 
 + 
 +Cache the workflow in ~/.nextflow/assets/ 
 + 
 +Download containers (if configured) 
 + 
 +Retrieve auxiliary files and dependencies 
 + 
 +=== Cached Files and Workflow Structure === 
 + 
 +Workflow code is stored in: 
 +<code>~/.nextflow/assets</code> 
 + 
 +Container images are stored in: 
 +<code>~/.singularity</code> 
 + 
 +**🗂 Finding the nextflow.config File** 
 + 
 +After pulling a workflow, you’ll typically find the nextflow.config file in its root directory. 
 + 
 +Example: 
 +<code bash> 
 +cd ~/.nextflow/assets/nf-core/rnaseq/ 
 +ls 
 +</code> 
 + 
 +Look for: 
 +<code>nextflow.config</code> 
 +If missing, config files may reside in the conf/ directory or be fetched remotely. You can always override settings by creating your own nextflow.config. 
 + 
 +🚫 No Manual PBS Scripts Needed 
 +Nextflow automatically generates and submits PBS scripts. You only define resources in nextflow.config. 
 + 
 +=== ⚙️ Configuration with nextflow.config === 
 + 
 +== 🔧 1. Global Resource Settings == 
 + 
 +Set default resource usage for all workflow processes: 
 + 
 +<code nextflow> process { 
 +    executor = 'pbs' 
 + 
 +    withLabel: big_job { 
 +        queue = 'smp' 
 +        cpus = 24 
 +        memory = '120 GB' 
 +        time = '24h' 
 +    } 
 +
 + </code> 
 +== 🏷️ 2. Custom Resource Labels == 
 + 
 +Customize resources for specific process groups using labels: 
 + 
 +<code nextflow> 
 +process { 
 +    executor = 'pbs' 
 + 
 +    withLabel: big_job { 
 +        cpus = 16 
 +        memory = '64 GB' 
 +        time = '12h' 
 +        queue = 'smp' 
 +    } 
 + 
 +    withLabel: short_job { 
 +        cpus = 1 
 +        memory = '1 GB' 
 +        time = '15m' 
 +        queue = 'smp' 
 +    } 
 +
 + 
 +</code> 
 + 
 +Use the label in your pipeline: 
 +<code nextflow> 
 +process bigTask { 
 +label 'big_job' 
 +... 
 +
 +</code> 
 + 
 +== 📦 3. Singularity Integration == 
 + 
 +Enable Singularity support: 
 + 
 +<code nextflow> singularity.enabled = true  
 +singularity.autoMounts = true </code> 
 + 
 +Specify containers: 
 + 
 +From Docker Hub: 
 +<code nextflow> 
 +process.container = 'docker://biocontainers/fastqc:v0.11.9_cv8' 
 +</code>     
 +From local image: 
 +<code nextflow> 
 +process.container = '/path/to/image.sif' 
 +</code> 
 + 
 +Set cache directory to avoid re-downloads: 
 + 
 +<code nextflow> singularity.cacheDir = '/path/to/.singularity' </code> 
 +== 🖥️ 4. PBS Executor Settings == 
 + 
 +Customize PBS job submission: 
 + 
 + 
 +<code nextflow>  
 +executor { 
 +  name = 'pbs' 
 +  queueSize = 20 
 +  submitOptions = '-V -m abe -M your@email.com' 
 +}  
 +</code> 
 + 
 +== 📂 5. Using Profiles == 
 + 
 +Profiles let you switch configurations easily: 
 + 
 +<code nextflow> profiles { 
 +  standard { 
 +    process.executor = 'pbs' 
 +    process.queue = 'smp' 
 +  } 
 + 
 +  local { 
 +    process.executor = 'local' 
 +    docker.enabled = false 
 +  } 
 + 
 +  cluster_singularity { 
 +    process.executor = 'pbs' 
 +    singularity.enabled = true 
 +    process.container = 'file:///path/to/container.sif' 
 +  } 
 +
 + </code> 
 +Run a profile: 
 + 
 +<code bash> nextflow run main.nf -profile cluster_singularity </code> 
 + 
 +== 🚫 Offline Mode == 
 + 
 +Run jobs on compute nodes without internet access: 
 + 
 +<code bash> nextflow run ~/.nextflow/assets/nf-core/rnaseq \  
 +-profile singularity -offline </code> 
 +⚠️ Always include -offline on compute nodes to prevent online fetching. 
 + 
 +=== 🧭 Debugging and Logs === 
 + 
 +Each Nextflow process generates a unique work directory (work/ab/xyz123), containing: 
 + 
 +.command.run — generated PBS job script 
 + 
 +.command.sh — wrapped shell script 
 + 
 +.command.log — job output 
 + 
 +.exitcode — exit status 
 + 
 +To inspect a failed job: 
 + 
 +<code bash> cd work/ab/xyz123/ 
 +less .command.log </code> 
 + 
 +=== ✅ Summary === 
 + 
 +nextflow.config centralizes all pipeline settings. 
 + 
 +No need to write PBS scripts manually. 
 + 
 +Resources, container usage, and submission options are all configurable. 
 + 
 +Profiles improve portability and reproducibility. 
 + 
 +Offline mode is essential for CHPC compute node compatibility. 
 + 
 +🧠 Tip: For workflows requiring reference data, bind directories just like with containers: 
 + 
 +<code bash> nextflow run /path/to/my_pipeline -profile singularity -offline \  
 +--input /data/input.fastq \  
 +--genomeDir /mnt/lustre/bsp/DB/genomes 
 +</code> 
 + 
 +🧹 Clean Up: 
 +Nextflow stores all its cache files in your home directory, so it's important to clean up these files once you're finished using a workflow to avoid running out of space. 
 +<code bash> 
 +rm -rf ~/.nextflow/assets/  
 +rm -rf ~/.nextflow/tmp  
 +rm -rf ~/.singularity 
 +</code>
  
-1. The commands you ran   +=== Need Help? === 
-2. The Singularity image used   +If you encounter issues or need a specific tool installed contact the CHPC support team at: 
-3Paths and error logs+  * 📧 help@chpc.ac.za 
 +  * https://users.chpc.ac.za/helpdesk/tickets/submit/ 
 +Include your job script and all errors encountered.
 ===== Basic examples ===== ===== Basic examples =====
  
/app/dokuwiki/data/attic/howto/bioinformatics.1746133847.txt.gz · Last modified: 2025/05/01 23:10 by nmfuphi