The CLUMP app identifies approximately independent genetic association signals within selected GWAS studies. It filters variants by statistical significance, then uses linkage disequilibrium (LD) to retain representative lead variants rather than multiple highly correlated variants from the same locus.
This is useful when preparing a set of independent genetic instruments for downstream analyses, including Mendelian randomisation.
Before you begin
Select the studies you want to analyse from the Apps page, then select Launch on the CLUMP app.
[Screenshot 1: Use the Apps page. Highlight the CLUMP app card and its “Launch” button.]
The CLUMP launch form will open.
Enter a job name
Enter a short, descriptive name in the Prefix field. This is required.
[Screenshot 2: Red box around the Prefix field.]
The prefix helps you identify your job and output files later. For example:
BMI_clumpingLDL_instrumentsCAD_EUR_clump
Set the variant significance filter
Use filter to define which variants are eligible for clumping. The default value, pval < 5e-8, retains genome-wide significant variants.
[Screenshot 3: Red box around the filter field.]
You may adjust this threshold according to your analysis plan. A more stringent threshold returns fewer variants; a less stringent threshold returns more variants.
Optionally select genomic regions
Use regions to restrict clumping to specific genomic regions.
[Screenshot 4: Red box around the regions dropdown.]
Leave this field blank to perform clumping across all available genomic regions in the selected studies.
Configure LD clumping settings
The following parameters determine how variants are considered independent.
[Screenshot 5: Red boxes around the r2, padding(LD), MAF(LD filter), and LD Ref fields.]
- r2 — The LD threshold used to determine whether variants are correlated.
A lower value is more stringent and retains more independent variants. The default of0.001is commonly used when selecting independent instruments. - padding(LD) — The genomic distance, in base pairs, used when assessing LD between variants. The default is
10000. - MAF(LD filter) — The minimum minor allele frequency used in LD calculations. The default is
0.01. - LD Ref — The ancestry-specific LD reference panel used for clumping. Select a reference population that is appropriate for your study population whenever possible. For example, choose EUR for analyses based primarily on European-ancestry data.
Choose output columns
Use columns_to_include to specify the study-level fields that should appear in the output.
[Screenshot 6: Red box around the columns_to_include field.]
The default fields are:
abbreviationstudy_phenotypestudy_id
These help identify the study and phenotype associated with each retained variant.
Select a project
Choose a destination in the project_id field. This is required.
[Screenshot 7: Red box around the project_id dropdown and the note directing users to Research Hub → Projects.]
If you do not yet have a suitable project, create one in Research Hub → Projects before launching the app.
Gene padding
Use padding(gene) to include an additional genomic distance around genes when relevant to your workflow. The default is 0.
[Screenshot 8: Red box around the padding(gene) field.]
Leave the default value unless you need to extend the region considered around genes.
Launch the app
Review the selected studies and settings, then select Launch App.
[Screenshot 9: Red box and arrow pointing to “Launch App”.]
CLUMP runs across the selected studies. Processing typically takes approximately one minute per study.
Retrieve the output
When the job is complete, download the output archive from the relevant project or job results area. The archive contains the clumped results for the selected studies.
Tips
- Use a clear job prefix so that you can easily distinguish repeated analyses.
- Match the LD reference population to the ancestry represented in your GWAS data where possible.
- The default
r2value of0.001is intentionally stringent and is appropriate when you need highly independent variants. - Restricting to regions can reduce processing time when you only need to analyse a specific locus or set of loci.