encoding

Functions to encode TR (Template Repeat) sequences as feature vectors for machine learning models. Uses ViennaRNA to compute RNA secondary structure folding energies. The key feature is ΔE = E(TR+Spacer) - E(TR), which captures whether the spacer interaction disrupts the TR fold — this affects bRT accessibility and predicts TR activity.

source

encode_tr_list

def encode_tr_list(
    list_TRs:list, # A list of DNA sequences (strings) to encode.
    feat:int=1, # 1 -> single-feature model, 2 -> two-feature model.
    n_jobs:NoneType=None, # Worker processes; default all-but-one core. n_jobs=1 forces serial.
    parallel_min:int=200, # Only parallelize batches at least this large (amortizes pool startup).
):

Encodes a list of TR sequences. If feat is 1, uses only the single feature model, if 2 uses the 2 features model.

Output is bit-identical to the original serial implementation. The 3 ViennaRNA partition-function folds per TR are independent across TRs, so they are spread over worker processes (not threads: ViennaRNA’s SWIG bindings hold the GIL). A persistent process pool is reused across calls to pay worker startup only once, and any pool failure falls back to the serial path via try/except.

NOTE: because this uses multiprocessing, callers launching it from a script must guard their entry point with if __name__ == "__main__": (spawn requirement on Windows/macOS).

On a personal computer, cap cores with the DGREC_NJOBS environment variable (e.g. DGREC_NJOBS=4) to keep the machine responsive; workers also run at below-normal priority.

TR_list=['AACTTGCTAGTAAACACAGTGCAGCTAAAGTGCATTCTAATGGGGTACTAAATGTGTTGTTCCATCTGGG','CCAGAGGTTAGGATTATGAGTGTAGTGCGTATCACGTGGGTTGCAGATTAAACAGGTCGTGCCTCATTTA']
encode_tr_list(TR_list,feat=2)
array([[ -8.17911339, -15.36869621],
       [-11.47970581, -13.57634544]])