Instructions to use Taykhoom/UTR-LM-MLMSS with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Taykhoom/UTR-LM-MLMSS with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
UTR-LM-MLMSS
Minimal HuggingFace port of the MLM + secondary structure variant of UTR-LM -- an ESM2-style 5' UTR RNA language model pretrained on endogenous sequences from five species and a large synthetic library.
Architecture
| Parameter | Value |
|---|---|
| Layers | 6 |
| Attention heads | 16 |
| Embedding dimension | 128 |
| FFN hidden dimension | 512 (GELU) |
| Vocabulary size | 10 |
| Positional encoding | RoPE (base=10000) |
| Normalization | LayerNorm |
| Architecture | ESM2-style pre-LN Transformer with GELU FFN |
| Max sequence length | 1024 tokens (1022 nucleotides + <cls> / <eos>) |
Vocabulary: <pad> (0), <eos> (1), <unk> (2), A (3), G (4), C (5), T (6), <cls> (7), <mask> (8), <sep> (9)
Pretraining
- Objective: Masked language modeling + per-token secondary structure prediction (3-class: unpaired, stem, loop)
- Data: Endogenous 5' UTRs from five species (human, mouse, zebrafish, Drosophila, yeast) combined with the Cao et al. random 5' UTR synthetic library
- Source checkpoint:
ESM2SS_FS4.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_lr1e-05_structureweight1.0_MLMLossMin_epoch200.pkl
Only one ESM2SS (secondary structure only, no MFE regression) checkpoint was available; no selection decision was required.
Parity Verification
All 7 representation levels (embedding + 6 transformer blocks) were verified
to be bit-exact (max absolute difference = 0.00) against the original
ESM2SS_FS4.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_lr1e-05_structureweight1.0_MLMLossMin_epoch200.pkl
weights. Verified on GPU with PyTorch 2.7.1 / CUDA 12.9.
Related Models
See the full UTR-LM collection.
| Model | Pretraining Objective | Notes |
|---|---|---|
| UTR-LM-MLM | MLM | Base model |
| UTR-LM-MLMSI | MLM + MFE regression | Recommended for TE / EL tasks |
| UTR-LM-MLMSS | MLM + secondary structure | This model |
| UTR-LM-MLMSISS | MLM + MFE + secondary structure | Recommended for MRL tasks |
Usage
Embedding generation
import torch
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True)
model = AutoModel.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True)
model.eval()
sequences = ["ATGCATGCATGC", "GCTAGCTAGCTAGCTA"]
enc = tokenizer(sequences, return_tensors="pt", padding=True)
with torch.no_grad():
out = model(**enc)
# CLS token embedding (position 0) - recommended for sequence-level tasks
cls_emb = out.last_hidden_state[:, 0, :] # (batch, 128)
# All-token embeddings
token_emb = out.last_hidden_state # (batch, seq_len, 128)
# Intermediate layer representations
out_all = model(**enc, output_hidden_states=True)
layer3_emb = out_all.hidden_states[3] # after layer 3, shape (batch, seq_len, 128)
MLM logits
import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM
tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True)
model = AutoModelForMaskedLM.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True)
model.eval()
enc = tokenizer(["ATGC<mask>ATGC"], return_tensors="pt")
with torch.no_grad():
logits = model(**enc).logits # (1, seq_len, 10)
Faster attention backends
# SDPA (PyTorch 2.0+)
model = AutoModel.from_pretrained(
"Taykhoom/UTR-LM-MLMSS",
trust_remote_code=True,
attn_implementation="sdpa",
)
# Flash Attention 2 (requires flash-attn)
model = AutoModel.from_pretrained(
"Taykhoom/UTR-LM-MLMSS",
trust_remote_code=True,
attn_implementation="flash_attention_2",
dtype=torch.bfloat16,
)
Fine-tuning
The model follows standard HF conventions and can be fine-tuned with any Trainer-compatible setup. For sequence regression tasks, use the CLS token embedding as input to a prediction head (as done in the original UTR-LM paper).
Implementation Notes
The source tokenizer uses the DNA-style A/G/C/T alphabet. Convert U to
T before tokenization when supplying RNA-spelled sequences; a literal U
otherwise maps to <unk>. Secondary structure was an auxiliary prediction
target, not an input channel; this minimal port preserves the backbone and
MLM head but omits the auxiliary structure head.
The original UTR-LM implementation uses eager scaled dot-product attention.
This port additionally supports attn_implementation="sdpa" and
attn_implementation="flash_attention_2".
Citation
@article{chu2024_utrlm,
title = {A 5' {UTR} Language Model for Decoding Untranslated Regions of {mRNA} and Function Predictions},
author = {Chu, Yanyi and Yu, Dan and Li, Yupeng and Huang, Kaixuan and Shen, Yue and Cong, Le and Zhang, Jason and Wang, Mengdi},
journal = {Nature Machine Intelligence},
volume = {6},
number = {4},
pages = {449--460},
year = {2024},
doi = {10.1038/s42256-024-00823-9}
}
Credits
Original model and code by Yanyi Chu et al. Source: UTR-LM GitHub repository. Hugging Face port maintained by Taykhoom Dalal.
License
GPL-3.0, following the original UTR-LM repository.
- Downloads last month
- 27