get_UMI_genotype
def get_UMI_genotype(
fastq_path:str, # path to the input fastq file
ref_seq:str, # sequence of the reference amplicon
umi_size:int=10, # number of nucleotides at the beginning of the read that will be used as the UMI
ref_read_size:int=None, # number of nucleotides in the read expected to align to the ref_seq. If None the whole read will be used.
quality_threshold:int=30, # threshold value used to filter out reads of poor average quality
ignore_pos:list=[], # list of positions that are ignored in the genotype
max_mutations:int=15, # reads with more mutations than this are discarded as bad data. Raise it when genotypes are expected to differ strongly from the reference, e.g. a VR replaced by a TR-derived copy.
base_quality_threshold:int=0, # if >0, bases below this Phred score are masked and never called as mutations. 0 keeps the previous behaviour of filtering on mean read quality only.
**kwargs
)->dict: # alignment parameters can be passed here (match, mismatch, gap_open, gap_extend)Takes as input a fastq_file of single read amplicon sequencing, and a reference amplicon sequence. Returns a dictionary containing as keys UMIs and as values a Counter of all genotype strings read for that UMI.