LRRtransfer is a gene family annotation transfer pipeline designed for plant LRR (Leucine-Rich Repeat) containing receptors.
The pipeline first identifies candidate loci in a closely related target genome using BLAST-based similarity searches, then transfers reference gene models using complementary nucleotide- and protein-based alignment strategies.
LRRtransfer is mainly composed of Bash and Python scripts orchestrated with Snakemake and is designed for local and HPC execution.
Clone the repository:
git clone https://github.com/ranwez/GeneModelTransfer.git
cd GeneModelTransferLRRtransfer requires:
- Snakemake 7 (tested with 7.32.4)
- Apptainer or Singularity
All workflow runtime dependencies are provided through a versioned container image. No separate installation of BLAST, Exonerate, samtools or the other bioinformatics dependencies is required.
The default dependency image is:
docker://ghcr.io/johgi/genemodeltransfer-deps:0.4.0
The image definition and locked dependency environment are available in container/.
The default image can be overridden through the singularity_image configuration parameter, for example to use a local SIF file on a cluster without access to GHCR.
A typical HPC execution is:
snakemake \
--snakefile SNAKEMAKE/LRR_transfer.smk \
--configfile path/to/config.yaml \
--profile SNAKEMAKE/PROFILES/<profile> \
--cluster-config path/to/cluster_config.yamlFor a local execution:
snakemake \
--snakefile SNAKEMAKE/LRR_transfer.smk \
--configfile path/to/config.yaml \
--cores 4 \
--use-singularityWorkflow parameters are defined in a YAML configuration file and validated against SNAKEMAKE/schemas/config.schema.yaml.
Example configuration files are available under:
example/*/CONFIG/config.yaml
| Parameter | Required | Default | Description |
|---|---|---|---|
target_genome |
yes | - | Target genome FASTA to annotate. |
ref_gff |
yes | - | Reference GFF containing the gene models to transfer. |
ref_genome |
conditional | null |
Reference genome FASTA. Required when an existing lrrome is not provided. |
lrrome |
conditional | null |
Prebuilt LRRome directory. Required when ref_genome is not provided. |
ref_locus_info |
no | null |
TSV containing reference gene identifier, family and class information. |
tblastn_results |
no | null |
Precomputed TBLASTN results. Generated by the workflow when omitted. |
blastn_results |
no | null |
Precomputed BLASTN results. Generated by the workflow when omitted. |
CL_min_sim |
no | 0.25 |
Minimum similarity used when selecting candidate loci. Must be between 0 and 1. |
ignore_exonerate_errors |
no | false |
Continue after unexpected Exonerate errors instead of stopping the workflow. |
out_dirname |
no | results |
Output directory. |
out_feature_id_prefix |
no | LRRt |
Prefix assigned to transferred feature identifiers. |
singularity_image |
no | GHCR dependency image | Override the default container with another GHCR image or a local SIF file. |
Exactly one of ref_genome and lrrome must be provided.
tblastn_results and blastn_results are independent. Either, both or neither can be supplied. Missing search results are computed automatically by the workflow.
A minimal configuration using a reference genome is therefore:
target_genome: "/path/to/target.fasta"
ref_gff: "/path/to/reference.gff"
ref_genome: "/path/to/reference.fasta"To reuse an existing LRRome instead:
target_genome: "/path/to/target.fasta"
ref_gff: "/path/to/reference.gff"
lrrome: "/path/to/LRRome"Snakemake profiles define execution settings specific to the computing environment, such as the cluster submission command, maximum number of concurrent jobs, latency settings and container execution.
Example profiles are provided in:
SNAKEMAKE/PROFILES/
├── sge_migale/
├── slurm_io/
└── slurm_muse/
Copy the closest profile and adapt its config.yaml to your HPC scheduler.
Cluster-specific resource values such as partition, account, walltime, memory and log paths are defined separately through --cluster-config.
Example cluster configuration files are available under:
example/*/CONFIG/slurm_config.yaml
example/*/CONFIG/sge_config.yaml
Input files are expected to be mutually consistent. In particular, identifiers used in the reference GFF must match the corresponding reference data and, when an LRRome is provided, its reference models. Precomputed BLAST results must refer to valid reference and target sequence identifiers.
LRRtransfer requires a reference GFF and one source of reference sequences:
ref_genome: LRRtransfer builds the LRRome from the reference genome and GFF.lrrome: an existing LRRome is validated and reused, and the reference genome is not required.
The reference GFF must be a valid GFF3 file.
In addition, LRRtransfer requires that:
- each gene contains exactly one mRNA;
- each mRNA contains at least one CDS.
A prebuilt LRRome must have the following structure:
LRRome/
├── REF_proteins.fasta
├── REF_loci.fasta
├── REF_PEP/
├── REF_EXONS/
├── REF_cDNA/
├── REF_LOCI/
└── REF_LOCI_GFF/
It contains the reference sequences and gene-model information required by LRRtransfer.
REF_proteins.fasta: protein sequences for all reference models.REF_loci.fasta: genomic locus sequences for all reference models.REF_PEP/: one protein FASTA file per reference model.REF_EXONS/: individual exon FASTA files for each reference model.REF_cDNA/: one cDNA FASTA file per reference model.REF_LOCI/: one genomic locus FASTA file per reference model.REF_LOCI_GFF/: one GFF file per reference model describing its gene structure.
ref_locus_info is an optional headerless tab-separated file containing exactly three columns:
| Column | Description |
|---|---|
gene_id |
Reference gene identifier. It must match a gene ID in ref_gff. |
family |
Gene family, for example LRR-RLK, LRR-RLP or NBS-LRR. |
class |
Reference gene class: Canonical for an intact coding model or Non-canonical for a model affected by coding-disrupting mutations. |
For example:
Gene001 LRR-RLK Canonical
Gene002 NBS-LRR Non-canonical
The reference family and class are reported in the output GFF as Origin-Fam and Origin-Class. The transferred gene is independently classified as Gene-Class, so its class may differ from that of the reference gene.
TBLASTN and BLASTN results can optionally be computed beforehand and supplied through tblastn_results and blastn_results.
When omitted, the corresponding search is performed automatically by LRRtransfer.
Precomputed results must be headerless tab-separated BLAST outfmt 6 files using the column order:
qseqid sseqid qlen length qstart qend sstart send nident pident gapopen evalue bitscore positive
For external BLAST results:
qseqidmust correspond to a gene ID inref_gffsseqidmust correspond to a sequence ID intarget_genome
All workflow outputs are written under out_dirname, which defaults to results/.
The main annotation produced by LRRtransfer is:
results/annot_best.gff
LRRtransfer also retains the annotations generated by the individual prediction strategies:
results/annot_<method>.gff
where <method> includes strategies such as mapping, locusAlignment, cdna2genome, cds2genome and prot2genome.
results/stats/GFFstats.txt
contains:
- total number of transferred genes
- number of genes per prediction method
- number of genes per target sequence
When they are not supplied as inputs, the workflow also generates and retains:
results/LRRome/
results/tblastn_refProt.tsv
results/blastn_refProt.tsv
Each run records its execution information under the output directory:
results/run_info.yaml
results/run_history/
results/LRRtransfer.log
run_info.yaml contains the provenance information for the latest run, including the resolved workflow configuration, Git commit, Snakemake and Python versions, command line, hostname and run status.
run_history/ keeps the same provenance information for previous runs, with one YAML file per run.
LRRtransfer.log provides a simple chronological record of workflow starts and completion status.
tools/update_blast_results.py can be used to update an existing BLAST result file without rerunning the complete search.
python3 tools/update_blast_results.py \
-i old_tblastn.tsv \
--queries-to-rm queries_to_remove.txt \
--blast-update updated_queries.tsv \
-o updated_tblastn.tsv--queries-to-rm is a text file containing one query ID per line. All HSPs corresponding to these queries are removed from the initial BLAST file.
--blast-update is a BLAST result file with the same column structure as the initial file. For each query present in this file, all previous HSPs for that query are removed and replaced by the new HSPs. Queries that were not present in the initial file are added.
At least one of --queries-to-rm or --blast-update must be provided. A query cannot be present in both.
If a recalculated query no longer produces any HSP, it cannot appear in --blast-update. In this case, its query ID must be explicitly added to --queries-to-rm to remove its previous results.
If --output is omitted, the updated BLAST table is written to standard output.
LRRprofiler (ranwez/LRRprofiler)
Gottin, C., Dievart, A., Summo, M., Droc, G., Périn, C., Ranwez, V., and Chantret, N. (2021). A new comprehensive annotation of leucine-rich repeat-containing receptors in rice. The Plant Journal, 108(2), 492–508. doi:10.1111/tpj.15456 https://doi.org/10.1111/tpj.15456
Slater, G. S. C., and Birney, E. (2005). Automated generation of heuristics for biological sequence comparison. BMC Bioinformatics, 6, 31. doi:10.1186/1471-2105-6-31 https://doi.org/10.1186/1471-2105-6-31
Camacho, C., Coulouris, G., Avagyan, V., Ma, N., Papadopoulos, J., Bealer, K., and Madden, T. L. (2009). BLAST+: architecture and applications. BMC Bioinformatics, 10, 421. doi:10.1186/1471-2105-10-421 https://doi.org/10.1186/1471-2105-10-421
Katoh, K., and Standley, D. M. (2013). MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Molecular Biology and Evolution, 30(4), 772–780. doi:10.1093/molbev/mst010 https://doi.org/10.1093/molbev/mst010
Quinlan, A. R., and Hall, I. M. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics, 26(6), 841–842. doi:10.1093/bioinformatics/btq033 https://doi.org/10.1093/bioinformatics/btq033
Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., and Durbin, R. (2009). The Sequence Alignment/Map format and SAMtools. Bioinformatics, 25(16), 2078–2079. doi:10.1093/bioinformatics/btp352 https://doi.org/10.1093/bioinformatics/btp352
Mölder, F., Jablonski, K. P., Letcher, B., Hall, M. B., van Dyken, P. C., Tomkins-Tinch, C. H., Sochat, V., Forster, J., Vieira, F. G., Meesters, C., Lee, S., Twardziok, S. O., Kanitz, A., VanCampen, J., Malladi, V., Wilm, A., Holtgrewe, M., Rahmann, S., Nahnsen, S., and Köster, J. (2025). Sustainable data analysis with Snakemake. F1000Research, 10, 33. doi:10.12688/f1000research.29032.3 https://doi.org/10.12688/f1000research.29032.3