Skip to content

About

LRRtransfer is a complex gene family annotation transfer pipeline designed for plant LRR (Leucine-Rich Repeat) containing receptors. The pipeline use annotation of LRR gene families from a reference genome to annotate LRR of a closely related genome.

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Latest commit

 

History

473 Commits

Folders and files

Repository files navigation

LRRtransfer

LRRtransfer is a gene family annotation transfer pipeline designed for plant LRR (Leucine-Rich Repeat) containing receptors.

The pipeline first identifies candidate loci in a closely related target genome using BLAST-based similarity searches, then transfers reference gene models using complementary nucleotide- and protein-based alignment strategies.

LRRtransfer is mainly composed of Bash and Python scripts orchestrated with Snakemake and is designed for local and HPC execution.

Installation and dependencies

Clone the repository:

git clone https://github.com/ranwez/GeneModelTransfer.git
cd GeneModelTransfer

LRRtransfer requires:

  • Snakemake 7 (tested with 7.32.4)
  • Apptainer or Singularity

All workflow runtime dependencies are provided through a versioned container image. No separate installation of BLAST, Exonerate, samtools or the other bioinformatics dependencies is required.

The default dependency image is:

docker://ghcr.io/johgi/genemodeltransfer-deps:0.4.0

The image definition and locked dependency environment are available in container/.

The default image can be overridden through the singularity_image configuration parameter, for example to use a local SIF file on a cluster without access to GHCR.

Running LRRtransfer

A typical HPC execution is:

snakemake \
    --snakefile SNAKEMAKE/LRR_transfer.smk \
    --configfile path/to/config.yaml \
    --profile SNAKEMAKE/PROFILES/<profile> \
    --cluster-config path/to/cluster_config.yaml

For a local execution:

snakemake \
    --snakefile SNAKEMAKE/LRR_transfer.smk \
    --configfile path/to/config.yaml \
    --cores 4 \
    --use-singularity

Configuration

Workflow parameters are defined in a YAML configuration file and validated against SNAKEMAKE/schemas/config.schema.yaml.

Example configuration files are available under:

example/*/CONFIG/config.yaml

Parameters

Parameter Required Default Description
target_genome yes - Target genome FASTA to annotate.
ref_gff yes - Reference GFF containing the gene models to transfer.
ref_genome conditional null Reference genome FASTA. Required when an existing lrrome is not provided.
lrrome conditional null Prebuilt LRRome directory. Required when ref_genome is not provided.
ref_locus_info no null TSV containing reference gene identifier, family and class information.
tblastn_results no null Precomputed TBLASTN results. Generated by the workflow when omitted.
blastn_results no null Precomputed BLASTN results. Generated by the workflow when omitted.
CL_min_sim no 0.25 Minimum similarity used when selecting candidate loci. Must be between 0 and 1.
ignore_exonerate_errors no false Continue after unexpected Exonerate errors instead of stopping the workflow.
out_dirname no results Output directory.
out_feature_id_prefix no LRRt Prefix assigned to transferred feature identifiers.
singularity_image no GHCR dependency image Override the default container with another GHCR image or a local SIF file.

Exactly one of ref_genome and lrrome must be provided.

tblastn_results and blastn_results are independent. Either, both or neither can be supplied. Missing search results are computed automatically by the workflow.

A minimal configuration using a reference genome is therefore:

target_genome: "/path/to/target.fasta"
ref_gff: "/path/to/reference.gff"

ref_genome: "/path/to/reference.fasta"

To reuse an existing LRRome instead:

target_genome: "/path/to/target.fasta"
ref_gff: "/path/to/reference.gff"

lrrome: "/path/to/LRRome"

Execution profiles

Snakemake profiles define execution settings specific to the computing environment, such as the cluster submission command, maximum number of concurrent jobs, latency settings and container execution.

Example profiles are provided in:

SNAKEMAKE/PROFILES/
├── sge_migale/
├── slurm_io/
└── slurm_muse/

Copy the closest profile and adapt its config.yaml to your HPC scheduler.

Cluster-specific resource values such as partition, account, walltime, memory and log paths are defined separately through --cluster-config.

Example cluster configuration files are available under:

example/*/CONFIG/slurm_config.yaml
example/*/CONFIG/sge_config.yaml

Input files and consistency requirements

Input files are expected to be mutually consistent. In particular, identifiers used in the reference GFF must match the corresponding reference data and, when an LRRome is provided, its reference models. Precomputed BLAST results must refer to valid reference and target sequence identifiers.

Reference genome and LRRome

LRRtransfer requires a reference GFF and one source of reference sequences:

  1. ref_genome: LRRtransfer builds the LRRome from the reference genome and GFF.
  2. lrrome: an existing LRRome is validated and reused, and the reference genome is not required.

Reference GFF

The reference GFF must be a valid GFF3 file.

In addition, LRRtransfer requires that:

  • each gene contains exactly one mRNA;
  • each mRNA contains at least one CDS.

Prebuilt LRRome

A prebuilt LRRome must have the following structure:

LRRome/
├── REF_proteins.fasta
├── REF_loci.fasta
├── REF_PEP/
├── REF_EXONS/
├── REF_cDNA/
├── REF_LOCI/
└── REF_LOCI_GFF/

It contains the reference sequences and gene-model information required by LRRtransfer.

  • REF_proteins.fasta: protein sequences for all reference models.
  • REF_loci.fasta: genomic locus sequences for all reference models.
  • REF_PEP/: one protein FASTA file per reference model.
  • REF_EXONS/: individual exon FASTA files for each reference model.
  • REF_cDNA/: one cDNA FASTA file per reference model.
  • REF_LOCI/: one genomic locus FASTA file per reference model.
  • REF_LOCI_GFF/: one GFF file per reference model describing its gene structure.

Reference locus information

ref_locus_info is an optional headerless tab-separated file containing exactly three columns:

Column Description
gene_id Reference gene identifier. It must match a gene ID in ref_gff.
family Gene family, for example LRR-RLK, LRR-RLP or NBS-LRR.
class Reference gene class: Canonical for an intact coding model or Non-canonical for a model affected by coding-disrupting mutations.

For example:

Gene001    LRR-RLK    Canonical
Gene002    NBS-LRR    Non-canonical

The reference family and class are reported in the output GFF as Origin-Fam and Origin-Class. The transferred gene is independently classified as Gene-Class, so its class may differ from that of the reference gene.

Precomputed TBLASTN and BLASTN results

TBLASTN and BLASTN results can optionally be computed beforehand and supplied through tblastn_results and blastn_results.

When omitted, the corresponding search is performed automatically by LRRtransfer.

Precomputed results must be headerless tab-separated BLAST outfmt 6 files using the column order:

qseqid sseqid qlen length qstart qend sstart send nident pident gapopen evalue bitscore positive

For external BLAST results:

  • qseqid must correspond to a gene ID in ref_gff
  • sseqid must correspond to a sequence ID in target_genome

Outputs

All workflow outputs are written under out_dirname, which defaults to results/.

Main annotation

The main annotation produced by LRRtransfer is:

results/annot_best.gff

LRRtransfer also retains the annotations generated by the individual prediction strategies:

results/annot_<method>.gff

where <method> includes strategies such as mapping, locusAlignment, cdna2genome, cds2genome and prot2genome.

Statistics

results/stats/GFFstats.txt

contains:

  • total number of transferred genes
  • number of genes per prediction method
  • number of genes per target sequence

Generated reference and similarity-search files

When they are not supplied as inputs, the workflow also generates and retains:

results/LRRome/
results/tblastn_refProt.tsv
results/blastn_refProt.tsv

Run provenance

Each run records its execution information under the output directory:

results/run_info.yaml
results/run_history/
results/LRRtransfer.log

run_info.yaml contains the provenance information for the latest run, including the resolved workflow configuration, Git commit, Snakemake and Python versions, command line, hostname and run status.

run_history/ keeps the same provenance information for previous runs, with one YAML file per run.

LRRtransfer.log provides a simple chronological record of workflow starts and completion status.

Utility tools

Updating precomputed BLAST results

tools/update_blast_results.py can be used to update an existing BLAST result file without rerunning the complete search.

python3 tools/update_blast_results.py \
    -i old_tblastn.tsv \
    --queries-to-rm queries_to_remove.txt \
    --blast-update updated_queries.tsv \
    -o updated_tblastn.tsv

--queries-to-rm is a text file containing one query ID per line. All HSPs corresponding to these queries are removed from the initial BLAST file.

--blast-update is a BLAST result file with the same column structure as the initial file. For each query present in this file, all previous HSPs for that query are removed and replaced by the new HSPs. Queries that were not present in the initial file are added.

At least one of --queries-to-rm or --blast-update must be provided. A query cannot be present in both.

If a recalculated query no longer produces any HSP, it cannot appear in --blast-update. In this case, its query ID must be explicitly added to --queries-to-rm to remove its previous results.

If --output is omitted, the updated BLAST table is written to standard output.

References

Related software

LRRprofiler (ranwez/LRRprofiler)

Initial work describing the gene model transfer pipeline

Gottin, C., Dievart, A., Summo, M., Droc, G., Périn, C., Ranwez, V., and Chantret, N. (2021). A new comprehensive annotation of leucine-rich repeat-containing receptors in rice. The Plant Journal, 108(2), 492–508. doi:10.1111/tpj.15456 https://doi.org/10.1111/tpj.15456

Exonerate

Slater, G. S. C., and Birney, E. (2005). Automated generation of heuristics for biological sequence comparison. BMC Bioinformatics, 6, 31. doi:10.1186/1471-2105-6-31 https://doi.org/10.1186/1471-2105-6-31

BLAST+

Camacho, C., Coulouris, G., Avagyan, V., Ma, N., Papadopoulos, J., Bealer, K., and Madden, T. L. (2009). BLAST+: architecture and applications. BMC Bioinformatics, 10, 421. doi:10.1186/1471-2105-10-421 https://doi.org/10.1186/1471-2105-10-421

MAFFT

Katoh, K., and Standley, D. M. (2013). MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Molecular Biology and Evolution, 30(4), 772–780. doi:10.1093/molbev/mst010 https://doi.org/10.1093/molbev/mst010

BEDTools

Quinlan, A. R., and Hall, I. M. (2010). BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics, 26(6), 841–842. doi:10.1093/bioinformatics/btq033 https://doi.org/10.1093/bioinformatics/btq033

SAMtools

Li, H., Handsaker, B., Wysoker, A., Fennell, T., Ruan, J., Homer, N., Marth, G., Abecasis, G., and Durbin, R. (2009). The Sequence Alignment/Map format and SAMtools. Bioinformatics, 25(16), 2078–2079. doi:10.1093/bioinformatics/btp352 https://doi.org/10.1093/bioinformatics/btp352

Snakemake

Mölder, F., Jablonski, K. P., Letcher, B., Hall, M. B., van Dyken, P. C., Tomkins-Tinch, C. H., Sochat, V., Forster, J., Vieira, F. G., Meesters, C., Lee, S., Twardziok, S. O., Kanitz, A., VanCampen, J., Malladi, V., Wilm, A., Holtgrewe, M., Rahmann, S., Nahnsen, S., and Köster, J. (2025). Sustainable data analysis with Snakemake. F1000Research, 10, 33. doi:10.12688/f1000research.29032.3 https://doi.org/10.12688/f1000research.29032.3

About

LRRtransfer is a complex gene family annotation transfer pipeline designed for plant LRR (Leucine-Rich Repeat) containing receptors. The pipeline use annotation of LRR gene families from a reference genome to annotate LRR of a closely related genome.

Topics

Resources

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages