pyfraglib is a Python library to analyze high-throughput sequencing data of
cell-free DNA (cfDNA). More specifically, it facilitates the investigation of
fragmentomics features of which a list can be found below.
Because fragmentation of cfDNA is non-random, tissue- and disease-specific
patterns emerge in e.g. fragment length or end motif distributions. To date,
no comprehensive tool exists to implement a "fragmentomics workflow".
pyfraglib aims at being such at tool.
It is recommended that pyfraglib is installed into a dedicated conda
environment using pip, the Python package manager:
git clone git@github.com:schwarzlab-ccb/pyfraglib.git
cd pyfraglib
conda env create -f pyfraglib.yml
conda activate pyfraglib
python3 -m pip install .During development, all code in pyfraglib is thoroughly type-checked using
mypy. After an initial installation as described above, do:
./tools/dev_install.sh # typing & linting errors are reportedAlternatively, pyfraglib is available via PyPI, the Python package index:
pip install pyfraglibThe only non-Python dependency samtools/htslib is usually bundled with
pysam and in those cases, no separate installation of it is required.
pyfraglib comes with a command line utility. After successful installation,
it can be used as follows:
pyfrag.py version # show currently installed version
pyfrag.py --help # show available subcommands and flagstools/workflow.sh gives an example of a commonly used sequence of commands.
pyfraglib offers a simple batch mode for most operations, i.e. it can take
a directory of BAM or FRAG files and perform analyses on them. Some operations
(i.e. the extract subcommand) are even parallelized. We explicitly do not
recommend using this functionality for the analysis of large cohorts. Even 10
BAM files cannot be extracted at once without exceeding most workstation's
memory. Thus, we include a convenient Nextflow pipeline in tools/pipeline.nf.
It reads the input data from a TSV file and Nextflow parallelizes pyfraglibs
operations as much as possible. See tools/nextflow.config for all paths and
variables that must be set by the user. Currently, only slurm is supported as
a scheduler (via the ramses profile).
Please refer to the project's documentation for a list of currently supported features (see below for instructions on how to build documentation using Sphinx).
Sphinx is used to create documentation. The easiest way to create repository-wide documentation in HTML or PDF format, run:
./tools/build_docs.sh pdf # or:
./tools/build_docs.sh htmlPlease cite our bioRxiv preprint (DOI XXX) if you use pyfraglib in your work.
pyfraglib is licensed under the GPL-v3 as indicated in the source files.
Please direct requests to daniel.schuette@iccb-cologne.org.
