How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("fill-mask", model="nlpie/modernalbert-large-v1.0", trust_remote_code=True)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModelForMaskedLM
model = AutoModelForMaskedLM.from_pretrained("nlpie/modernalbert-large-v1.0", trust_remote_code=True, device_map="auto")
Quick Links

ModernALBERT-Large

ModernALBERT-Large is a compact, recursive transformer for natural language understanding. It combines ALBERT-style cross-layer parameter sharing with Mixture of LoRAs (MoL) β€” a lightweight, token-conditional routing mechanism that restores the expressivity normally lost when transformer layers share weights β€” plus a set of modern architectural upgrades (RoPE, GeGLU, FlashAttention, Pre-Norm).

It is well suited to text classification, natural language inference, paraphrase/semantic-similarity detection, extractive QA, and dense retrieval.

Model repo: nlpie/modernalbert-large-v1.0

Other sizes in this family: tiny Β· medium Β· base Β· large


Table of Contents

  1. Overview
  2. Model Architecture
  3. How to Use
  4. Training & Dataset
  5. GLUE Benchmark Results
  6. SQuAD-v2 Results
  7. BEIR Retrieval Results
  8. Inference Efficiency
  9. Key Features and Design Choices
  10. Limitations
  11. Citation

Overview

ModernALBERT builds on ALBERT's cross-layer parameter sharing, which reduces model size but can cap representational capacity when layers are fully tied. ModernALBERT addresses this with:

  • Mixture of LoRAs (MoL): low-rank LoRA "experts" injected directly into the weights of the shared feed-forward network, with sparse router-driven activation, simulating a Mixture-of-Experts layer at a fraction of the parameter cost.
  • Modern architecture: Pre-Norm, GeGLU activations, rotary position embeddings (RoPE), and FlashAttention (with an automatic PyTorch SDPA fallback).
  • Distillation-based initialisation: weights are seeded from a fully-parameterised ModernBERT teacher via layer-mapped initialisation, and training uses knowledge distillation from that same teacher β€” critical for reaching strong performance on a comparatively small pretraining budget.

ModernALBERT-Large is the flagship variant of the family: 24 layers organised into 6 shared groups, with MoL layers placed at the end of the last four groups.


Model Architecture

Variant Layers Groups MoL Groups Hidden Dim FFN Intermediate Dim Expert (LoRA) Dim Experts Top-K
Large 24 6 3, 4, 5, 6 1024 2624 4096 8 2
  • Parameter sharing: layers are grouped (group depth = 4); all layers within a group share attention and FFN weights, so the model behaves like a 24-layer network while storing far fewer unique parameters.
  • Mixture of LoRAs: the last MoL group in each recursion replaces the shared FFN with a router over 8 low-rank LoRA experts (top-2 routing), letting different tokens activate different experts.
  • Attention: rotary embeddings for position information, FlashAttention (unpadded/varlen) when available, otherwise scaled-dot-product attention.
  • Embeddings: ALBERT-style factorised token embeddings (small embedding dimension projected up to the hidden size), reducing the size of the embedding matrix.
  • ~120M parameters total (paper-reported figure for this configuration).

How to Use

ModernALBERT ships with custom transformers-compatible modeling code (ModernALBERTConfig, ModernALBERTModel, ModernALBERTForMaskedLM, ModernALBERTForSequenceClassification, ModernALBERTForQuestionAnswering). Load it with trust_remote_code=True. It runs on both CPU and GPU: FlashAttention is used automatically when installed, and the model falls back to PyTorch's built-in SDPA attention otherwise β€” no extra configuration needed either way.

pip install transformers torch
# Optional, for the fastest attention path on supported GPUs (auto-detected; falls back to SDPA if absent):
pip install flash-attn --no-build-isolation

Masked language modeling

import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM

model_id = "nlpie/modernalbert-large-v1.0"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(model_id, trust_remote_code=True)
model.eval()

text = f"Paris is the capital of {tokenizer.mask_token}."
inputs = tokenizer(text, return_tensors="pt")

with torch.no_grad():
    outputs = model(**inputs)

mask_index = (inputs.input_ids == tokenizer.mask_token_id)[0].nonzero(as_tuple=True)[0]
predicted_id = outputs.logits[0, mask_index].argmax(dim=-1)
print(tokenizer.decode(predicted_id))

Sentence / token embeddings

from transformers import AutoTokenizer, AutoModel
import torch

model_id = "nlpie/modernalbert-large-v1.0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True)
model.eval()

inputs = tokenizer("Example sentence for embeddings.", return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

attention_mask = inputs["attention_mask"]
last_hidden = outputs.last_hidden_state
# Mean pooling over valid tokens
embedding = (last_hidden * attention_mask.unsqueeze(-1)).sum(1) / attention_mask.sum(1, keepdim=True)

Fine-tuning for sequence classification

from transformers import AutoTokenizer, AutoModelForSequenceClassification, TrainingArguments, Trainer

model_id = "nlpie/modernalbert-large-v1.0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
    model_id, trust_remote_code=True, num_labels=2
)

# tokenized_train_dataset, tokenized_eval_dataset = ...  # your tokenized datasets

training_args = TrainingArguments(
    output_dir="./modernalbert-large-finetuned",
    per_device_train_batch_size=16,
    num_train_epochs=3,
    learning_rate=2e-5,
    eval_strategy="epoch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=tokenized_train_dataset,
    eval_dataset=tokenized_eval_dataset,
)
trainer.train()

Extractive question answering

from transformers import AutoTokenizer, AutoModelForQuestionAnswering
import torch

model_id = "nlpie/modernalbert-large-v1.0"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForQuestionAnswering.from_pretrained(model_id, trust_remote_code=True)

question, context = "What does MoL stand for?", "ModernALBERT introduces Mixture of LoRAs (MoL), a routing mechanism over low-rank experts."
inputs = tokenizer(question, context, return_tensors="pt")

with torch.no_grad():
    outputs = model(**inputs)

start = outputs.start_logits.argmax()
end = outputs.end_logits.argmax() + 1
print(tokenizer.decode(inputs["input_ids"][0][start:end]))

If Auto-class mapping isn't yet wired up on the Hub

If trust_remote_code=True doesn't resolve the classes automatically, either import the classes directly from the repo's Python files, or add an auto_map block to config.json:

{
  "auto_map": {
    "AutoConfig": "configuration_modernalbert.ModernALBERTConfig",
    "AutoModel": "modeling_modernalbert.ModernALBERTModel",
    "AutoModelForMaskedLM": "modeling_modernalbert.ModernALBERTForMaskedLM",
    "AutoModelForSequenceClassification": "modeling_modernalbert.ModernALBERTForSequenceClassification",
    "AutoModelForQuestionAnswering": "modeling_modernalbert.ModernALBERTForQuestionAnswering"
  }
}

Compatibility

This code has been verified end-to-end β€” building the model, running a forward pass, and a save_pretrained β†’ from_pretrained round trip with bit-identical weights β€” on CPU (PyTorch's SDPA attention path), and structurally validated against the FlashAttention code path. It targets a recent transformers release (tested against 5.x); on much older transformers versions you may need to upgrade, since the weight-tying and rotary-embedding buffer conventions it relies on changed across versions.

Efficient (merged) inference

The routing_strategy field in ModernALBERTConfig controls how the MoL layer behaves:

  • "standard" (default): sparse top-2 token routing, as used during pretraining.
  • "uniform": experts are averaged with equal weight into a single static LoRA adapter (no per-token routing), matching the "Vanilla" merge strategy in the paper.
  • "ema": experts are merged using an exponential moving average of the router's historical activations, matching the paper's dynamic EMA-merging strategy β€” this recovers accuracy close to the unmerged model while removing routing overhead at inference time (see Inference Efficiency).

Training & Dataset

  • Corpus: two-stage curriculum β€” warm-up on RedPajama-1T (20k–30k steps), then continued training on RefinedWeb (70k–80k further steps).
  • Budget: ~30B tokens total, versus 1.7T tokens for the ModernBERT teacher.
  • Initialisation: step-wise, layer-mapped initialisation from a fully-parameterised ModernBERT teacher.
  • Distillation: ModernBERT's predictions are used as soft targets alongside the MLM objective.
  • Optimisation: AdamW, global batch size 384, max sequence length 1024, linear warmup to a peak learning rate of 5Γ—10⁻⁴ or 5Γ—10⁻⁡, followed by linear decay.

GLUE Benchmark Results

Task Category Task Score
Single Sentence CoLA 66.4
Single Sentence SST-2 95.5
Paraphrase / Similarity MRPC 92.7
Paraphrase / Similarity STS-B 92.1
Paraphrase / Similarity QQP 92.0
Natural Language Inference MNLI 88.9
Natural Language Inference QNLI 93.7
Natural Language Inference RTE 88.44
Average 88.72

For context, this surpasses the fully-parameterised ModernBERT-base (149M params, 88.45 avg) and outperforms prior compact baselines such as MiniLM and MosaicBERT, with particularly strong results on RTE, STS-B, and MRPC.


SQuAD-v2 Results

Metric Score
F1 92.9
Exact Match 85.9

BEIR Retrieval Results

Selected BEIR datasets, reported for the ModernALBERT model in the paper:

Dataset Score
NFCorpus 24.30
SciFact 56.90
TREC-COVID 72.85
FiQA 30.43
ArguAna 48.82
Average (subset) 46.66

This is a strong result on domain shift β€” e.g., ArguAna (argument retrieval) improves substantially over ModernBERT (48.82 vs. 35.7).


Inference Efficiency

Latency and throughput measured with and without the expert-merging procedure described in the paper (batch inference, single GPU):

Model Latency (ms) ↓ Throughput (tok/s) ↑ Memory (GB) ↓
ModernALBERT-large, no merging 31.45 32,270 0.459
ModernALBERT-large, merged experts 18.87 54,483 0.459
ModernBERT-large (reference, dense) 21.30 48,117 1.5

Merging collapses the dynamic MoL router into a single static LoRA adapter at deployment time (see Efficient (merged) inference above), cutting latency substantially while keeping the same memory footprint and most of the accuracy gains from routing.


Key Features and Design Choices

  • Compact but flexible: parameter sharing keeps the model small; MoL restores per-token expressivity with low-rank experts.
  • Conditional computation: only the top-2 experts activate per token in MoL layers.
  • Modern training stack: Pre-Norm, GeGLU, RoPE, and FlashAttention (with SDPA fallback) for training stability and speed.
  • Distillation-based warm start: initialised and distilled from ModernBERT, enabling strong results on a fraction of ModernBERT's pretraining token budget.
  • Deployment-friendly: an optional expert-merging step (routing_strategy="ema" or "uniform") compresses MoL into a single dense adapter for lower-latency inference.

Limitations

  • MoE-style routing still carries more computational overhead than a fully dense/shared model, even with the merging optimisation; further work on load balancing and expert selection could close this gap.
  • The model uses global attention only (no local/sliding-window attention), so it may underperform on tasks requiring very long context or fine-grained long-range reasoning compared to architectures with hybrid attention patterns.
  • Benchmark numbers above are as reported in the accompanying paper; results can vary with fine-tuning setup, hardware, and library versions.

Citation

This model accompanies the paper "Improving Recursive Transformers with Mixture of LoRAs" (currently under anonymous review). Formal citation details will be added once the paper is published β€” check back on this model card or the paper's repository for an updated BibTeX entry.

Downloads last month
61
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including nlpie/modernalbert-large-v1.0