funannotate is a pipeline for genome annotation (built specifically for fungi, but will also work with higher eukaryotes). Installation, usage, and more information can be found at http://funannotate.readthedocs.io
You can use docker to run funannotate. Caveats are that GeneMark is not included in the docker image (see licensing below and you can complain to the developers for making it difficult to distribute/use). I've also written a bash script that can run the docker image and auto-detect/include the proper user/volume bindings. This docker image is built off of the latest code in master, so it will be ahead of the tagged releases. The image includes the required databases as well, if you want just funannotate without the databases then that is located on docker hub as well nextgenusfs/funannotate-slim. So this route can be achieved with:
# download/pull the image from docker hub
$ docker pull nextgenusfs/funannotate
# download bash wrapper script (optional)
$ wget -O funannotate-docker https://raw.githubusercontent.com/nextgenusfs/funannotate/master/funannotate-docker
# might need to make this executable on your system
$ chmod +x /path/to/funannotate-docker
# assuming it is in your PATH, now you can run this script as if it were the funannotate executable script
$ funannotate-docker test -t predict --cpus 12
This repo builds its own container image separate from the upstream
nextgenusfs/funannotate Docker Hub image referenced above -- it packages
the rust_EVM_trinity_PASA branch (Rust-optimized PASA/EVidenceModeler/
Trinity forks) via Dockerfile (release) / Dockerfile.dev (iterative dev
builds), installed through conda-pack into /venv. Build it locally:
docker build -t funannotate-live -f Dockerfile .
or convert to a Singularity/Apptainer image for HPC use:
singularity build funannotate-live.sif docker-daemon://funannotate-live:latest
Environment variables baked into the image (all overridable at
container-runtime, either via docker run -e VAR=.../--env VAR=... or by
exporting on the host before singularity exec -- Singularity passes host
env through by default). FUNANNOTATE_DB and EGGNOG_DATA_DIR are both
placeholder subpaths under a common /opt/databases parent -- bind your own
data onto each subpath, or override the variables individually:
| Variable | Image default | Purpose |
|---|---|---|
FUNANNOTATE_DB |
/opt/databases/funannotate_db |
funannotate's core reference DBs. Bind your own onto this path, or override the variable to point elsewhere. |
EGGNOG_DATA_DIR |
/opt/databases/eggnog_db |
EggNOG-mapper DB (~50GB, not baked into the image -- same reason the upstream Docker image omits it, see docs/docker.rst). Bind your own copy onto this path. Without a real DB bound here, funannotate annotate's auto-run of emapper.py crashes rather than skipping cleanly (see docs/containers.rst). |
PASAHOME |
/venv/opt/pasa/src |
PASA install root. |
PERL5LIB |
includes PASA's SAMPLE_HOOKS + PerlLib |
Required for PASA's hook loader to resolve GFF3::GFF3_annot_retriever during funannotate update's annotation-comparison step -- without it, that step crashes with Error, couldn't resolve path for GFF3::GFF3_annot_retriever. |
EVM_HOME |
/venv/opt/evm |
EVidenceModeler install root. |
See docs/containers.rst for a full worked example (build, bind mounts,
Singularity usage on shared HPC storage, and known gotchas).
funannotate 1.9.0: the bioconda package is built separately and may be an older release than this one. For 1.9.0 follow docs/conda.rst (a conda environment for the external tools,
pip install funannotate, and scripts that build the Rust-optimized EVM, PASA and Trinity), or use the container image. The commands below install whatever version bioconda currently has.
The pipeline can be installed with conda (via bioconda):
#add appropriate channels
conda config --add channels defaults
conda config --add channels bioconda
conda config --add channels conda-forge
#then create environment
conda create -n funannotate "python>=3.6,<3.9" funannotate
If conda is taking forever to solve the environment, I would recommend giving mamba a try:
#install mamba into base environment
conda install -n base mamba
#then use mamba as drop in replacmeent
mamba create -n funannotate funannotate
If you want to use GeneMark-ES/ET you will need to install that manually following developers instructions: http://topaz.gatech.edu/GeneMark/license_download.cgi
Note that you will need to change the shebang line for all perl scripts in GeneMark to use /usr/bin/env perl.
You will then also need to add gmes_petap.pl to the $PATH or set the environmental variable $GENEMARK_PATH to the gmes_petap directory.
To install just the python funannotate package, you can do this with pip:
python -m pip install funannotate
To install the most updated code in master you can run:
python -m pip install git+https://github.com/nextgenusfs/funannotate.git
Jonathan M. Palmer, & Jason E. Stajich. (2020). Funannotate v1.8.1: Eukaryotic genome annotation (v1.8). Zenodo. https://doi.org/10.5281/zenodo.1134477
