get_UMI_genotype_paired
def get_UMI_genotype_paired(
fastq_path_fwd:str, # path to the input fastq file reading the ref_seq in the forward orientation
fastq_path_rev:str, # path to the input fastq file reading the ref_seq in the reverse orientation
ref_seq:str, # sequence of the reference amplicon
fwd_span:tuple, # span of the ref_seq that is read in the forward orientation format: (start, end)
rev_span:tuple, # span of the ref_seq that is read in the reverse orientation format: (start, end)
fwd_ref_read_size:int=None, # number of nucleotides in the fwd read expected to align to the ref_seq. If None the whole read will be used.
rev_ref_read_size:int=None, # number of nucleotides in the rev read expected to align to the ref_seq. If None the whole read will be used.
require_perfect_pair_agreement:bool=True, # if True only pairs of reads that perfectly agree on the sequence within the common span will be used. If False the fwd sequence will be used. Will be set to False by default if there is no overlap.
umi_size_fwd:int=10, # number of nucleotides at the beginning of the fwd read that will be used as the UMI
umi_size_rev:int=0, # number of nucleotides at the beginning of the rev read that will be used as the UMI (if both are provided the umi will be the concatenation of both)
quality_threshold:int=30, # threshold value used to filter out reads of poor average quality. Both reads have to pass the threshold.
ignore_pos:list=[], # list of positions that are ignored in the genotype
max_mutations:int=15, # read pairs with more mutations than this are discarded as bad data. Raise it when genotypes are expected to differ strongly from the reference, e.g. a VR replaced by a TR-derived copy.
base_quality_threshold:int=0, # if >0, bases below this Phred score are masked and never called as mutations. 0 keeps the previous behaviour of filtering on mean read quality only.
N:NoneType=None, # number of reads to consider (useful to get a quick view of the data without going through the whole fastq files). If None the whole data will be used.
**kwargs
)->dict: # alignment parameters can be passed here (match, mismatch, gap_open, gap_extend)Processes paired-end FASTQ files to extract UMI-grouped genotypes. Aligns forward and reverse reads to the reference, optionally requiring perfect agreement in overlapping regions.