A set of tools to analyse the output from TraDIS analyses

Quadram Institute Bioscience

Last update: Feb 16, 2022

Related tags

Data Analysis QuaTradis

Overview

QuaTradis (Quadram TraDis)

A set of tools to analyse the output from TraDIS analyses

Introduction
Installation
Usage
License
Feedback/Issues
Citation

Introduction

The QuaTradis pipeline provides software utilities for the processing, mapping, and analysis of transposon insertion sequencing data. The pipeline was designed with the data from the TraDIS sequencing protocol in mind, but should work with a variety of transposon insertion sequencing protocols as long as they produce data in the expected format.

For more information on the TraDIS method, see http://bioinformatics.oxfordjournals.org/content/32/7/1109 and http://genome.cshlp.org/content/19/12/2308.

Installation

QuaTradis has the following dependencies:

Required dependencies

bwa
smalt
samtools
tabix

There are a number of ways to install QuaTradis and details are provided below. If you encounter an issue when installing QuaTradis please contact your local system administrator.

Bioconda

Install conda and enable the bioconda channel.

conda install -c bioconda quatradis=xxx

Docker

QuaTradis can be run in a Docker container. First install Docker, then pull the QuaTradis image from dockerhub:

docker pull quadraminstitute/quatradis

To use QuaTradis use a command like this (substituting in your directories), where your files are assumed to be stored in /home/ubuntu/data:

docker run --rm -it -v /home/ubuntu/data:/data quadraminstitute/quatradis bacteria_tradis -h

Running the tests

The test can be run with pytest from the tests directory. Alternatively you can use the make target from the top-level directory:

make test

Usage

QuaTradis provides functionality to:

detect TraDIS tags in a BAM file
add the tags to the reads
filter reads in a FastQ file containing a user defined tag
remove tags
map to a reference genome
create an insertion site plot file

The functions are available as standalone scripts or as perl modules.

Scripts

Executable scripts to carry out most of the listed functions are available in the bin:

check_tradis_tags - Prints 1 if tags are present in alignment file, prints 0 if not.
add_tradis_tags - Generates a BAM file with tags added to read strings.
filter_tradis_tags - Create a fastq file containing reads that match the supplied tag
remove_tradis_tags - Creates a fastq file containing reads with the supplied tag removed from the sequences
tradis_plot - Creates an gzipped insertion site plot
bacteria_tradis - Runs complete analysis, starting with a fastq file and produces mapped BAM files and plot files for each file in the given file list and a statistical summary of all files. Note that the -f option expects a text file containing a list of fastq files, one per line. This script can be run with or without supplying tags.

Note that default parameters are for comparative experiments, and will need to be modified for gene essentiality studies.

A help menu for each script can be accessed by running the script by adding with "--help".

Analysis Scripts

Three scripts are provided to perform basic analysis of TraDIS results in bin:

tradis_gene_insert_sites - Takes genome annotation in embl format along with plot files produced by bacteria_tradis and generates tab-delimited files containing gene-wise annotations of insert sites and read counts.
tradis_essentiality.R - Takes a single tab-delimited file from tradis_gene_insert_sites to produce calls of gene essentiality. Also produces a number of diagnostic plots.
tradis_comparison.R - Takes tab files to compare two growth conditions using edgeR. This analysis requires experimental replicates.

License

QuaTradis is free software, licensed under GPLv3.

Feedback/Issues

Please report any issues to the issues page or email [email protected]

Citation

If you use this software please cite:

"The TraDIS toolkit: sequencing and analysis for dense transposon mutant libraries", Barquist L, Mayho M, Cummins C, Cain AK, Boinett CJ, Page AJ, Langridge G, Quail MA, Keane JA, Parkhill J. Bioinformatics. 2016 Apr 1;32(7):1109-11. doi: 10.1093/bioinformatics/btw022. Epub 2016 Jan 21.

Comments

fix channel order in readme

Channel order is important for bioconda to work correctly -- the conda-forge has to come first (which means higher priority when specified on the command line with -c). That might be why some users are getting pysam issues requiring a workaround.

FYI might also want to consider suggesting --strict-channel-priority, see the new bioconda docs.

opened by daler 1
Fixes for albatradis compatibility

Fixing name of analysis output files for consumption by albatradis.

Fixing mistake when creating gene names during insertion site analysis.. Shouldn't have ignored underscores in the name.

opened by maplesond 0
requirements.txt should not list bgzip

A followup to the discussion on the Bioconda PR: The requirements.txt file that you are using should not list bgzip. Names in requirements.txt refer to packages on PyPI, so if you list bgzip, you actually pull in a Python package named bgzip (that is meant to be used via import bgzip from within Python). It will not give you the bgzip binary that your project actually seems to want.

You cannot list non-Python dependencies in requirements.txt so you can only list that dependency in the Conda recipe.

opened by marcelm 0
Fixing problems running the job in docker.

The issue was that the mapping stage outputs files to the current working directory which may not have user permissions. The fix is to make sure mapping logs are output to the same place as all other output files.

opened by maplesond 0
Nextflow pipeline to replace bacteria_tradis, and implementation of tradis_gene_insert_sites

Adding nextflow to handle processing of multiple fastq files (similar to bacteria_tradis).

Add the tradis_gene_insert_sites script, and associated functions under isp_analyse. Although there are still some very small diffs between this and old biotradis script in terms of ins_index and ins_count, which I still need to investigate.

Renamed and refactored a few things.

Added a few scripts to get closer to feature parity with old BioTradis.

Tidied up README.

opened by maplesond 0
problem with running tradis pipeline multiple

Hello,

When I try to run following command using quatradis:

tradis pipeline multiple -v -n 12 -o quatradis_out fastqs_filtered_sizecut_all.txt genome.fa

this error appears: Traceback (most recent call last): File "/home/jang/anaconda3/envs/mamba/envs/albatradis/bin/tradis", line 293, in main() File "/home/jang/anaconda3/envs/mamba/envs/albatradis/bin/tradis", line 285, in main args.func(args) File "/home/jang/anaconda3/envs/mamba/envs/albatradis/bin/tradis", line 202, in run_multiple_pipeline tradis.run_multi_tradis(args.fastqs, args.reference, File "/home/jang/anaconda3/envs/mamba/envs/albatradis/lib/python3.9/site-packages/quatradis/tradis.py", line 142, in run_multi_tradis pipeline = find_pipeline_file() File "/home/jang/anaconda3/envs/mamba/envs/albatradis/lib/python3.9/site-packages/quatradis/tradis.py", line 101, in find_pipeline_file if os.path.exists(exe_path): File "/home/jang/anaconda3/envs/mamba/envs/albatradis/lib/python3.9/genericpath.py", line 19, in exists os.stat(path) TypeError: stat: path should be string, bytes, os.PathLike or integer, not NoneType

What I'm doing wrong?

The same input files work smoothly in bacteria_tradis.

Bests, Jan

opened by gaworj 1

Releases(0.8.3)

0.8.3(Jun 21, 2022)

Source code(tar.gz)
Source code(zip)
0.8.2(May 10, 2022)

Source code(tar.gz)
Source code(zip)
0.8.1(May 10, 2022)

Source code(tar.gz)
Source code(zip)
0.8.0(May 10, 2022)

Source code(tar.gz)
Source code(zip)
0.7.0(Apr 2, 2022)

Source code(tar.gz)
Source code(zip)
0.6.2(Mar 6, 2022)

Source code(tar.gz)
Source code(zip)
0.6.1(Mar 5, 2022)

Source code(tar.gz)
Source code(zip)
0.6.0(Mar 5, 2022)

Source code(tar.gz)
Source code(zip)
0.5.4(Mar 4, 2022)

Source code(tar.gz)
Source code(zip)
0.5.3(Mar 4, 2022)

Source code(tar.gz)
Source code(zip)
0.5.2(Mar 4, 2022)

Source code(tar.gz)
Source code(zip)
0.5.1(Mar 4, 2022)

Source code(tar.gz)
Source code(zip)
0.5.0(Mar 4, 2022)

Source code(tar.gz)
Source code(zip)
0.4.10(Mar 4, 2022)

Source code(tar.gz)
Source code(zip)
0.4.9(Mar 2, 2022)

Source code(tar.gz)
Source code(zip)
0.4.8(Mar 2, 2022)

Source code(tar.gz)
Source code(zip)
0.4.7(Mar 2, 2022)

Source code(tar.gz)
Source code(zip)
0.4.6(Mar 2, 2022)

Source code(tar.gz)
Source code(zip)
0.4.5(Feb 16, 2022)

Source code(tar.gz)
Source code(zip)
0.4.4(Feb 16, 2022)

Source code(tar.gz)
Source code(zip)
0.4.3(Feb 16, 2022)

Source code(tar.gz)
Source code(zip)
0.4.2(Feb 14, 2022)

Source code(tar.gz)
Source code(zip)
0.4.1(Feb 14, 2022)

Source code(tar.gz)
Source code(zip)
0.4.0(Feb 14, 2022)

Source code(tar.gz)
Source code(zip)
0.3.4(Feb 10, 2022)

Source code(tar.gz)
Source code(zip)
0.3.3(Feb 10, 2022)

null
Source code(tar.gz)
Source code(zip)

Owner

Quadram Institute Bioscience

GitHub

PyIOmica (pyiomica) is a Python package for omics analyses.

PyIOmica (pyiomica) This repository contains PyIOmica, a Python package that provides bioinformatics utilities for analyzing (dynamic) omics datasets.

13 Jun 29, 2022

Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data.

PremiershipPlayerAnalysis Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data. No

5 Sep 6, 2021

Analyse the limit order book in seconds. Zoom to tick level or get yourself an overview of the trading day.

Analyse the limit order book in seconds. Zoom to tick level or get yourself an overview of the trading day. Correlate the market activity with the Apple Keynote presentations.

2 Jan 4, 2022

A set of functions and analysis classes for solvation structure analysis

SolvationAnalysis The macroscopic behavior of a liquid is determined by its microscopic structure. For ionic systems, like batteries and many enzymes,

19 Nov 24, 2022

Evaluation of a Monocular Eye Tracking Set-Up

Evaluation of a Monocular Eye Tracking Set-Up As part of my master thesis, I implemented a new state-of-the-art model that is based on the work of Che

19 Dec 17, 2022

WaveFake: A Data Set to Facilitate Audio DeepFake Detection

WaveFake: A Data Set to Facilitate Audio DeepFake Detection This is the code repository for our NeurIPS 2021 (Track on Datasets and Benchmarks) paper

27 Dec 22, 2022

CINECA molecular dynamics tutorial set

High Performance Molecular Dynamics Logging into CINECA's computer systems To logon to the M100 system use the following command from an SSH client ss

0 Mar 13, 2022

A set of procedures that can realize covid19 virus detection based on blood.

3 Mar 7, 2022

A project consists in a set of assignements corresponding to a BI process: data integration, construction of an OLAP cube, qurying of a OPLAP cube and reporting.

TennisBusinessIntelligenceProject - A project consists in a set of assignements corresponding to a BI process: data integration, construction of an OLAP cube, qurying of a OPLAP cube and reporting.

1 Jan 2, 2022

MeSH2Matrix - A set of Python codes for the generation of biomedical ontologies from the MeSH keywords of the PubMed scholarly publications

A set of Python codes for the generation of biomedical ontologies from the MeSH keywords of the PubMed scholarly publications

6 Nov 30, 2022

An ETL Pipeline of a large data set from a fictitious music streaming service named Sparkify.

An ETL Pipeline of a large data set from a fictitious music streaming service named Sparkify. The ETL process flows from AWS's S3 into staging tables in AWS Redshift.

1 Feb 11, 2022

A Pythonic introduction to methods for scaling your data science and machine learning work to larger datasets and larger models, using the tools and APIs you know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

This tutorial's purpose is to introduce Pythonistas to methods for scaling their data science and machine learning work to larger datasets and larger models, using the tools and APIs they know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

102 Nov 10, 2022

🧪 Panel-Chemistry - exploratory data analysis and build powerful data and viz tools within the domain of Chemistry using Python and HoloViz Panel.

???? ??. The purpose of the panel-chemistry project is to make it really easy for you to do DATA ANALYSIS and build powerful DATA AND VIZ APPLICATIONS within the domain of Chemistry using using Python and HoloViz Panel.

97 Dec 8, 2022

Visions provides an extensible suite of tools to support common data analysis operations

Visions And these visions of data types, they kept us up past the dawn. Visions provides an extensible suite of tools to support common data analysis

168 Dec 28, 2022

Helper tools to construct probability distributions built from expert elicited data for use in monte carlo simulations.

Elicited Helper tools to construct probability distributions built from expert elicited data for use in monte carlo simulations. Credit to Brett Hoove

3 Nov 4, 2022

Universal data analysis tools for atmospheric sciences

U_analysis Universal data analysis tools for atmospheric sciences Script written in python 3. This file defines multiple functions that can be used fo

1 Oct 10, 2021

GWpy is a collaboration-driven Python package providing tools for studying data from ground-based gravitational-wave detectors

GWpy is a collaboration-driven Python package providing tools for studying data from ground-based gravitational-wave detectors. GWpy provides a user-f

342 Jan 7, 2023

Additional tools for particle accelerator data analysis and machine information

PyLHC Tools This package is a collection of useful scripts and tools for the Optics Measurements and Corrections group (OMC) at CERN. Documentation Au

3 Apr 13, 2022

Tools for working with MARC data in Catalogue Bridge.

catbridge_tools Tools for working with MARC data in Catalogue Bridge. Borrows heavily from PyMarc

1 Nov 11, 2021

A set of tools to analyse the output from TraDIS analyses

Related tags

Overview

QuaTradis (Quadram TraDis)

Contents

Introduction

Installation

Required dependencies

Bioconda

Docker

Running the tests

Usage

Scripts

Analysis Scripts

License

Feedback/Issues

Citation

Comments

fix channel order in readme

Fixes for albatradis compatibility

requirements.txt should not list bgzip

Fixing problems running the job in docker.

Nextflow pipeline to replace bacteria_tradis, and implementation of tradis_gene_insert_sites

problem with running tradis pipeline multiple

Releases(0.8.3)

0.8.3(Jun 21, 2022)

0.8.2(May 10, 2022)

0.8.1(May 10, 2022)

0.8.0(May 10, 2022)

0.7.0(Apr 2, 2022)

0.6.2(Mar 6, 2022)

0.6.1(Mar 5, 2022)

0.6.0(Mar 5, 2022)

0.5.4(Mar 4, 2022)

0.5.3(Mar 4, 2022)

0.5.2(Mar 4, 2022)

0.5.1(Mar 4, 2022)

0.5.0(Mar 4, 2022)

0.4.10(Mar 4, 2022)

0.4.9(Mar 2, 2022)

0.4.8(Mar 2, 2022)

0.4.7(Mar 2, 2022)

0.4.6(Mar 2, 2022)

0.4.5(Feb 16, 2022)

0.4.4(Feb 16, 2022)

0.4.3(Feb 16, 2022)

0.4.2(Feb 14, 2022)

0.4.1(Feb 14, 2022)

0.4.0(Feb 14, 2022)

0.3.4(Feb 10, 2022)

0.3.3(Feb 10, 2022)

Owner

Quadram Institute Bioscience

PyIOmica (pyiomica) is a Python package for omics analyses.

Using Python to scrape some basic player information from www.premierleague.com and then use Pandas to analyse said data.

Analyse the limit order book in seconds. Zoom to tick level or get yourself an overview of the trading day.

A set of functions and analysis classes for solvation structure analysis

Evaluation of a Monocular Eye Tracking Set-Up

WaveFake: A Data Set to Facilitate Audio DeepFake Detection

CINECA molecular dynamics tutorial set

A set of procedures that can realize covid19 virus detection based on blood.

A project consists in a set of assignements corresponding to a BI process: data integration, construction of an OLAP cube, qurying of a OPLAP cube and reporting.

MeSH2Matrix - A set of Python codes for the generation of biomedical ontologies from the MeSH keywords of the PubMed scholarly publications

An ETL Pipeline of a large data set from a fictitious music streaming service named Sparkify.

A Pythonic introduction to methods for scaling your data science and machine learning work to larger datasets and larger models, using the tools and APIs you know and love from the PyData stack (such as numpy, pandas, and scikit-learn).

🧪 Panel-Chemistry - exploratory data analysis and build powerful data and viz tools within the domain of Chemistry using Python and HoloViz Panel.

Visions provides an extensible suite of tools to support common data analysis operations

Helper tools to construct probability distributions built from expert elicited data for use in monte carlo simulations.

Universal data analysis tools for atmospheric sciences

GWpy is a collaboration-driven Python package providing tools for studying data from ground-based gravitational-wave detectors

Additional tools for particle accelerator data analysis and machine information

Tools for working with MARC data in Catalogue Bridge.