estimate_library_size
def estimate_library_size(
observed_counts:list, G:int, # total number of possible unique genotypes
N:int, # number of cells
verbose:bool=False
)->dict:Estimates the expected number of unique genotypes sampled (K_N).
This function uses a hybrid model: 1. The “head” (high-frequency genotypes) uses empirical probabilities. 2. The “tail” (low-frequency and unobserved genotypes) is modeled with a power-law (Zipf) distribution.
Args: observed_counts (list[int]): A list of UMI counts for each observed unique genotype. G (int): The total number of possible unique genotypes. N (int): The number of cells in the culture to sample from.
Returns: dict: A dictionary containing the results: ‘E_KN’ (float): The final estimated E[K_N]. ‘B’ (int): The estimated breakpoint. ‘alpha’ (float): The estimated power-law exponent. ‘E_head’ (float): The contribution to E[K_N] from the head. ‘E_tail’ (float): The contribution to E[K_N] from the tail.
