MOSS voice-profile LoRAs β€” ten pilot voices, and the hub for the whole LoRA ecosystem

Ten LoRA adapters for laion/moss-tts-local-transformer-4.55b-voice-acting-v2. One adapter per voice, each holding a single speaker identity across the whole expressive range: 40 emotions, 57 VoiceNet dimensions, acting edge cases, character clusters, vocal bursts, English and German.

They are also the entry point to a family of MOSS adapters β€” emotions, character clusters, VoiceNet dimensions, vocal bursts, domain styles β€” that all attach to the same base model and are designed to be stacked on top of a voice. Β§ The ecosystem is the map; Β§ Stacking is the part nobody gets right on the first try.

If you have never seen this project before, read Β§ Quickstart then Β§ Read this first.


Read this first: what is wrong with this release

Not buried at the bottom, because these change how you should use it. Full detail in Β§ Limitations.

  1. No human has listened to any of this in a controlled study. Every number on this page β€” speaker similarity, reward, genuineness, blend, emotion strength β€” is the output of a learned scorer. They have been observed disagreeing with listening judgements. Treat them as relative signals between arms, never as absolute quality.
  2. 43.4 % of the training corpus falls below the 0.40 speaker-similarity floor β€” 57.1 % in the intense-emotion block. Identity is ranked, not gated, so the adapters were trained on takes that partly drift off the reference voice. The adapters improve this substantially; they do not remove it.
  3. Identity is bought with expressiveness. Across the pilot the adapters raise speaker similarity far above the base model and lower genuineness. If your application values spontaneity over identity, the base model may be the better starting point.
  4. The reference dataset is superseded-in-waiting. It encodes run PPILOT2, which has two measured text defects: 99.96 % of its 135,630 burst tags are Title-Case (MOSS spells (Growl) out letter by letter instead of performing it), and burst density is 33.7 % against 50 % intended. Both are fixed in the in-flight 500-voice build, whose indices 490–499 will replace these same ten voices.

The ecosystem: what attaches to what

Everything below is a PEFT LoRA adapter for one base model. They differ only in what they were trained to move.

                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚  laion/moss-tts-local-transformer-4.55b-         β”‚   the frozen base:
                    β”‚  voice-acting-v2                                 β”‚   4.55 B, reference-
                    β”‚      +  OpenMOSS-Team/MOSS-Audio-Tokenizer-v2    β”‚   conditioned TTS
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                            β”‚  every adapter below merges into THIS
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β”‚               β”‚               β”‚               β”‚               β”‚               β”‚
  β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”
  β”‚  WHO      β”‚  β”‚  FEELING    β”‚  β”‚  TIMBRE   β”‚  β”‚  NON-SPEECH β”‚  β”‚  CHARACTERβ”‚  β”‚  DOMAIN     β”‚
  β”‚  speaks   β”‚  β”‚             β”‚  β”‚  knobs    β”‚  β”‚  events     β”‚  β”‚  archetypeβ”‚  β”‚  style      β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ THIS REPO β”‚  β”‚ emotion-    β”‚  β”‚ voicenet- β”‚  β”‚ vocal-burst β”‚  β”‚ character β”‚  β”‚ mediathek   β”‚
  β”‚ 10 voices β”‚  β”‚ loras-v3    β”‚  β”‚ dimension β”‚  β”‚ -lora-      β”‚  β”‚ -loras-*  β”‚  β”‚ sports      β”‚
  β”‚           β”‚  β”‚ 40 emotions β”‚  β”‚ -loras    β”‚  β”‚ adapters    β”‚  β”‚ 120 clust.β”‚  β”‚ explicitnessβ”‚
  β”‚ velvet-   β”‚  β”‚             β”‚  β”‚ 114       β”‚  β”‚ 64 classes  β”‚  β”‚           β”‚  β”‚             β”‚
  β”‚ sage      β”‚  β”‚             β”‚  β”‚ (57 Γ— Β±)  β”‚  β”‚             β”‚  β”‚           β”‚  β”‚             β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
        β”‚
        β”‚ trained on, and evaluated against
        β–Ό
  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚  TTS-AGI/moss-voice-profile-references  (dataset)      β”‚  402,560 takes Β· 853.5 h
  β”‚  pilot/<voice>/reference.wav  ← the conditioning clip  β”‚  50 WebDataset shards
  β”‚  pilot/<voice>/metadata.parquet  ← 83 annotation cols  β”‚  every score component kept
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
repo what it is adapters subfolder key typical merge weight Ξ»
laion/moss-tts-local-transformer-4.55b-voice-acting-v2 the base model β€” everything here attaches to it β€” β€” β€”
OpenMOSS-Team/MOSS-Audio-Tokenizer-v2 the audio tokenizer/codec; pass as codec_path= to the processor β€” β€” β€”
TTS-AGI/moss-voice-profile-loras (this repo) who speaks β€” one speaker identity across the full expressive range 10 shipped (+60 ablation) <voice> e.g. k325_age3_bg1 1.0
TTS-AGI/moss-voice-lora-velvet-sage-baritone the single-voice predecessor of this repo, ranks 16 and 32 2 shipped rank16, rank32 1.0
TTS-AGI/moss-emotion-loras-v3 what it feels like β€” one adapter per emotion 40 Anger, Sadness, … 0.5 moderate, 0.5–1.9 intense (per emotion)
laion/moss-voicenet-dimension-loras timbre/prosody knobs β€” 57 VoiceNet dimensions, each in both directions 114 vn_<CODE>__high / __low per dimension & direction, 0.25–1.25
laion/vocal-burst-lora-adapters non-speech events β€” sighs, sobs, gasps, laughs 64 sobs, sharp_inhale, … 0.5
TTS-AGI/moss-character-loras-refined-public character archetypes, refined; three epoch checkpoints each. human/ is the genuineness scaffold 120 Γ— ep1–3 <slug>/ep3 1.0
TTS-AGI/moss-character-loras-genuine-public same 120 archetypes, genuine (unrefined) variant β€” one checkpoint each 120 <slug> 1.0
laion/moss-mediathek-emotion-lora German broadcast expressiveness; rank/epoch grid 9 r64_e2 (used) 0.25
laion/moss-sports-commentator-lora live sports commentary register 6 real_r64_e8 (used) 1.0
TTS-AGI/moss-explicitness-loras adult/unguarded register; gated by an age check upstream 6 aesthetic_mix_r32 (used) 0.8
TTS-AGI/moss-voice-profile-references the dataset these adapters were trained and evaluated on, plus every reference clip β€” pilot/<voice>/ β€”

The Ξ» column is not decoration. See Β§ Stacking β€” activating an adapter in peft applies it at its trained strength, which for most of these is the wrong dose.


Quickstart

Environment

Python 3.11, one CUDA GPU. The bf16 base is ~9.1 GB of weights and the audio tokenizer and KV cache sit on top of it; 24 GB is a safe floor for single-sentence generation. Tested on one GH200 (96 GB). Exactly the versions it was tested with:

pip install "torch==2.8.0" "transformers==5.14.1" "peft==0.20.0" \
            "accelerate==1.14.0" "huggingface_hub==1.25.1" \
            "safetensors==0.8.0" "soundfile==0.14.0" "numpy==1.26.4"

trust_remote_code=True is required β€” the MOSS TTS modelling code ships in the base-model repo, not in transformers.

The code

Also in this repo as quickstart.py, which is the file that was actually executed.

import numpy as np, soundfile as sf, torch
from huggingface_hub import hf_hub_download, snapshot_download
from peft import PeftModel
from transformers import AutoModel, AutoProcessor

BASE  = "laion/moss-tts-local-transformer-4.55b-voice-acting-v2"
CODEC = "OpenMOSS-Team/MOSS-Audio-Tokenizer-v2"
LORAS = "TTS-AGI/moss-voice-profile-loras"
REFS  = "TTS-AGI/moss-voice-profile-references"
VOICE = "k325_age3_bg1"          # Velvet Sage Baritone
GEN_SR = 48000                   # what proc.decode() returns

# 1. processor + codec
proc = AutoProcessor.from_pretrained(BASE, trust_remote_code=True, codec_path=CODEC)
proc.audio_tokenizer = proc.audio_tokenizer.to("cuda").eval()

# 2. frozen base (the "weights not initialised: audio_lm_heads/text_lm_head" warnings are benign)
model = AutoModel.from_pretrained(BASE, trust_remote_code=True, dtype=torch.bfloat16,
                                  attn_implementation="sdpa").to("cuda").eval()

# 3. the voice adapter. The top level of a voice folder is the SHIPPED adapter.
#    Online, `PeftModel.from_pretrained(model, LORAS, subfolder=VOICE)` is enough. Resolving
#    to a local directory first also works on an air-gapped machine -- see the note below.
root = snapshot_download(LORAS, allow_patterns=[f"{VOICE}/adapter_*"])
model = PeftModel.from_pretrained(model, f"{root}/{VOICE}").eval()

# 4. the reference clip. The adapter carries the identity, but the base model is STILL
#    reference-conditioned: generate without one and you get a random speaker wearing the
#    adapter. Use the same reference the adapter was trained against.
ref = hf_hub_download(REFS, f"pilot/{VOICE}/reference.wav", repo_type="dataset")

text = ("I have read the file. There is nothing in it that surprises me, "
        "and that is exactly what worries me.")
instruction = ("A warm, aged baritone, unhurried and contemplative, "
               "speaking just above a murmur.")

conv = [[proc.build_user_message(text=text, instruction=instruction, language="English",
                                 reference=[ref], tokens=max(8, len(text.split())))]]
batch = proc(conv, mode="generation")

torch.manual_seed(0)
out = model.generate(input_ids=batch["input_ids"].cuda(),
                     attention_mask=batch["attention_mask"].cuda(),
                     max_new_frames=400, do_sample=True,
                     text_temperature=0.7, text_top_k=50, text_top_p=1.0,
                     audio_temperature=1.0, audio_top_p=0.95, audio_top_k=30,
                     audio_repetition_penalty=1.1)

# 5. ALWAYS check audio_codes_list. An empty decode is a normal silent failure of this
#    model, not an exception -- on one run every candidate came back empty.
msg = proc.decode(out)[0]
assert msg.audio_codes_list, "empty decode; retry with another seed"
w = msg.audio_codes_list[0].cpu().float().numpy()
w = np.ascontiguousarray(w.mean(0) if w.ndim > 1 else w)
sf.write("quickstart.wav", w, GEN_SR)
print(f"{len(w)/GEN_SR:.2f}s @ {GEN_SR} Hz mono")

peft subfolder= is broken offline. In peft 0.20.0, PeftModel.from_pretrained(..., subfolder="X") works normally when the hub is reachable, but under HF_HUB_OFFLINE=1 it takes a different code path (load_peft_weights) that puts the subfolder into the filename and passes it again as a hf_hub_download kwarg. It then looks for X/X/adapter_model.safetensors, does not find it, and raises LocalEntryNotFoundError: Cannot find the requested files in the disk cache β€” which reads like a missing download rather than a doubled path. The config loads fine (PeftConfig.from_pretrained handles the subfolder correctly), so only the weights fail. Resolving the repo to a local directory first, as above, avoids the branch entirely and behaves identically online and offline.

Expected output

This was run. quickstart.py was executed end to end on one GH200 (Slurm job pvdoctest, seed 0, the default text above) with the versions pinned above, and it produced:

[66s] base loaded
[67s] adapter k325_age3_bg1 attached (r=4)
[81s] wrote quickstart.wav: 5.04s @ 48000 Hz mono, peak 0.824
what to check value
sample rate 48 000 Hz, mono β€” this is what proc.decode() returns. The scorers in this project resample to 16 kHz; the audio itself is 48 k.
duration 5.04 s for the 20-word default line (~4.0 words/s)
file size 473 KB (483,884 bytes) as 16-bit PCM WAV β€” β‰ˆ 96 KB per second of audio
peak / RMS peak 0.824, RMS 0.107, 8.8 % of samples near-silent
wall clock 66 s to load the base from a cold shared filesystem, ~1 s to attach the adapter, 14 s to generate and decode

If your file is a few hundred bytes, or the duration is ~0.1 s, the decode came back empty β€” regenerate with a different seed. If it is 5 s of noise, check that you passed reference=[...].

The decoder returns float and can exceed Β±1.0 (a base-model take in the stacking run below peaked at 1.047), which soundfile silently clips when it writes 16-bit PCM. If that matters, write float (subtype="FLOAT") or normalise before writing.

The exact duration and peak will not reproduce bit-for-bit on different hardware or library versions (sampling is stochastic even at a fixed seed once the kernel schedule changes), but the sample rate, the rough length and the byte-per-second ratio will.

The three parameters that matter

parameter what it does notes
text the words, and the direction Inline tags in round brackets are directions, not spoken words: (sobs), (sharp inhale). Square brackets are pauses. Lower-case them β€” MOSS spells a capitalised token out letter by letter, so (Growl) is delivered "gee-are-oh-doubleyou-el". This is the single most common mistake with this model and it is what damaged the reference dataset.
instruction the caption: who is speaking and how Free text. The adapters were trained with captions resampled every epoch from measured attributes, so they are robust to phrasing β€” but the caption still steers.
reference list of paths to the conditioning clip(s) Required. Use pilot/<voice>/reference.wav from the dataset repo.

max_new_frames=400 is roughly a 32 s ceiling; the sampling defaults above (audio_temperature=1.0, top_p=0.95, top_k=30, repetition_penalty=1.1) are the corpus defaults and are a reasonable starting point for all ten voices.


The ten voices

Numbers are the shipped adapter measured on held-out groups it never saw (192 paired clips each, same prompts and seeds as every other arm). reference is a path inside TTS-AGI/moss-voice-profile-references; each voice folder there also holds voice.json (the identity card), metadata.parquet (83 annotation columns) and five WebDataset shards of every take generated for it.

voice (subfolder) name ships params spk-sim base reward genuineness blend WER reference
anime_088 Breathless Exile's Whisper r4 8.6 M 0.4566 0.3286 7.06 0.89 5.90 0.086 pilot/anime_088/reference.wav
emolia_c0542 Measured Slavic Historian r8 17.2 M 0.6167 0.3922 2.53 0.58 1.09 0.086 pilot/emolia_c0542/reference.wav
emolia_c1682 Cynical Streetwise Youth r4 8.6 M 0.6083 0.3872 3.31 1.23 1.38 0.105 pilot/emolia_c1682/reference.wav
emolia_c1699 Poised Intellectual Professional r4 8.6 M 0.5480 0.3659 4.48 0.97 2.63 0.075 pilot/emolia_c1699/reference.wav
emolia_c2570 Scholarly Matriarch Historian r4 8.6 M 0.6123 0.4447 3.11 1.32 1.22 0.117 pilot/emolia_c2570/reference.wav
k10_age3_bg1 The Serene Storyteller r4 8.6 M 0.6374 0.4066 3.11 0.37 1.94 0.076 pilot/k10_age3_bg1/reference.wav
k325_age3_bg1 Velvet Sage Baritone r4 8.6 M 0.6369 0.4134 4.02 0.43 2.84 0.080 pilot/k325_age3_bg1/reference.wav
k395_age3_bg1 Fragile Matriarch r4 8.6 M 0.5915 0.3105 6.76 1.10 5.45 0.105 pilot/k395_age3_bg1/reference.wav
k91_age5_bg0 Fading Elder Prophetess r4 8.6 M 0.6574 0.3753 5.39 0.47 4.37 0.068 pilot/k91_age5_bg0/reference.wav
mediathek_0184 Whispering Teutonic Chronicler r4 8.6 M 0.6880 0.5139 5.14 0.46 4.15 0.089 pilot/mediathek_0184/reference.wav

Identity cards, in the same order:

voice gender age language / accent source pool
anime_088 Male Late 30s–40s English, Japanese-accented anime
emolia_c0542 Male Late 40s–50s English, Slavic-accented emolia
emolia_c1682 Male Early–mid 20s English, AAVE emolia
emolia_c1699 Female Late 20s–mid 30s English, Standard American emolia
emolia_c2570 Female Late 60s–70s English, Standard American with AAVE undercurrents emolia
k10_age3_bg1 Androgynous Adult (30s–50s) β€” character cluster
k325_age3_bg1 Male Late 40s–60s β€” character cluster
k395_age3_bg1 Female Late 40s–60s β€” character cluster
k91_age5_bg0 Female Late 70s–80s β€” character cluster
mediathek_0184 Male Late 60s–70s German, Standard German German broadcast

Reading the numbers. spk-sim is ECAPA cosine to the reference clip (higher = more the same person; 0.40 is this project's floor). base is the same measurement with no adapter β€” the column that shows what the adapter bought. reward is the corpus's own composite ranking score and is not comparable across voices, only across arms of the same voice. genuineness (0–6) and blend (0–10) are learned heads. WER is Whisper-large-v3-turbo on the generated audio with inline tags stripped.

anime_088 is the hard case: a breathy, gravelly, heavily-accented voice whose base similarity is the lowest of the ten (0.3286) and whose adapted similarity (0.4566) is barely above the 0.40 floor. It is also the voice with 80.4 % of its corpus takes below the floor. Expect identity drift.

Listen: profile pages for all ten voices.


Stacking adapters

The non-obvious part, and the reason this page exists.

Why you can't just call set_adapter

A voice LoRA gives you who. An emotion LoRA gives you how it feels. A vocal-burst LoRA gives you the sob in the middle of the line. You want all three at once, at different strengths β€” and peft has no dose parameter.

PeftModel.set_adapter(name) activates one adapter at its trained scaling (alpha / r). PeftModel.base_model.set_adapter([a, b, c]) activates several, still each at its trained scaling. There is no Ξ» argument anywhere. The dose lives one level down, in LoraLayer.scaling[name], and a stack is therefore always the same two moves:

  1. activate the set through pm.base_model.set_adapter([...])
  2. rewrite m.scaling[name] on every LoraLayer, relative to the trained value

The peft trap that killed four runs

PeftModel.active_adapter is a plain attribute, not a property. base_model.set_adapter() does not update it.

peft sets PeftModel.active_adapter once in __init__, to the name of the first adapter ever loaded, and thereafter only updates it inside PeftModel.set_adapter(). Step 1 above goes through pm.base_model.set_adapter() β€” which updates the LoraModel and leaves the PeftModel attribute pinned to the first adapter, forever.

That is harmless right up until the first adapter is deleted or evicted. From that moment every generate() evaluates peft_config[self.active_adapter] and raises:

KeyError: 'A_emo_Fear'

permanently, on an adapter name that no longer exists. It killed four sweep arms, four separate times, each stopping at exactly the group where the resident adapter set first exceeded the cache limit. Two attempted fixes missed it because peft's delete_adapter() does try to repair active_adapter β€” but only when exactly one adapter is left active, which is never true for a stack.

The fix is one line, after activating the set:

pm.base_model.set_adapter(names)        # names = the list you want active
pm.active_adapter = names[0]            # the attribute peft actually indexes

The second trap: a in m.scaling is not enough

Rewriting m.scaling[a] is destructive, so you cannot read the trained value back out of it later β€” you have to snapshot m.scaling first and always compute the dose as snapshot Γ— Ξ». The trap is that the snapshot goes stale: it is taken over the modules and adapters that exist at the moment you take it, and any adapter loaded afterwards β€” or reloaded after an LRU eviction β€” is present in m.scaling but absent from your snapshot. The naive if a in m.scaling: m.scaling[a] = base_scaling[nm][a] * s then raises KeyError on an adapter that is loaded and is active, which reads as impossible.

So: check both dicts, and when only m.scaling has the key, treat its current value as the trained one and record it. Re-snapshot after every load_adapter as well (the load() above does).

The same failure would also be produced by adapters that target genuinely different module sets. Worth knowing that in this ecosystem they do not: the voice, emotion, vocal-burst, VoiceNet-dimension, character and Mediathek adapters all declare the same 23 target-module patterns (q,k,v,o,gate,up,down_proj, c_attn, c_proj, fc_in, fc_out, audio_lm_heads.0–11) and differ only in rank β€” 4, 32 and 64 respectively. Checked against the published adapter_config.json of each. If you bring in an adapter from outside this family, check its target_modules before assuming.

The working pattern

Complete, runnable version in this repo as stack_adapters.py β€” it also generates base / voice / voice+emotion / voice+emotion+burst and then deletes an adapter and generates again, which is where trap 1 fires.

from huggingface_hub import snapshot_download
from peft import PeftModel
from peft.tuners.lora import LoraLayer

class AdapterStack:
    """Several LoRAs on one frozen base, each at its own merge weight."""

    def __init__(self, base_model):
        self.pm = None
        self.base = base_model
        self.base_scaling = {}                       # module -> {adapter: trained scaling}

    def load(self, name, model_id, subfolder=None):
        path = model_id
        if subfolder:                                # see the offline note in Β§ Quickstart
            path = f"{snapshot_download(model_id, allow_patterns=[subfolder + '/adapter_*'])}/{subfolder}"
        if self.pm is None:
            self.pm = PeftModel.from_pretrained(self.base, path, adapter_name=name).eval()
        else:
            self.pm.load_adapter(path, adapter_name=name)
        # Re-snapshot: a module first reached by adapter B has no entry for B in a
        # snapshot taken while only A was loaded.
        self.base_scaling = {nm: dict(m.scaling) for nm, m in self.pm.named_modules()
                             if isinstance(m, LoraLayer)}
        return self

    def set_active(self, spec):
        """spec: {name: lambda}. Weight 0 or absent = off; {} = pure base model."""
        pm = self.pm
        keys = {n: float(v) for n, v in spec.items() if v}
        if not keys:
            pm.base_model.disable_adapter_layers()
            return
        pm.base_model.enable_adapter_layers()
        pm.base_model.set_adapter(list(keys))

        # TRAP 1 -- see above. One line, four dead runs.
        first = next(iter(keys))
        if first in getattr(pm, "peft_config", {}):
            pm.active_adapter = first

        # TRAP 2 -- check BOTH dicts, and record the trained value on first sight.
        for nm, m in pm.named_modules():
            if not isinstance(m, LoraLayer):
                continue
            for a, s in keys.items():
                if a not in m.scaling:
                    continue
                if a in self.base_scaling.get(nm, {}):
                    m.scaling[a] = self.base_scaling[nm][a] * s
                else:
                    self.base_scaling.setdefault(nm, {})[a] = m.scaling[a]
                    m.scaling[a] = m.scaling[a] * s


stack = (AdapterStack(model)
         .load("voice",   "TTS-AGI/moss-voice-profile-loras", subfolder="k325_age3_bg1")
         .load("emotion", "TTS-AGI/moss-emotion-loras-v3",    subfolder="Anger")
         .load("burst",   "laion/vocal-burst-lora-adapters",  subfolder="sobs"))

stack.set_active({"voice": 1.0, "emotion": 0.25, "burst": 0.5})
# ... then generate() exactly as in the quickstart, on stack.pm

This was run too. stack_adapters.py was executed on one GH200, same seed for every arm, same sentence except where a burst tag is added:

adapters: ['voice', 'emotion', 'burst']
    voice    r=4   alpha=8   targets=23 modules
    emotion  r=32  alpha=64  targets=23 modules
    burst    r=32  alpha=64  targets=23 modules
  00_base                 {}                                            3.84s  peak 1.047
  01_voice                {'voice': 1.0}                                3.84s  peak 0.652
  02_voice_emotion        {'voice': 1.0, 'emotion': 1.9}                1.68s  peak 0.902
  03_voice_emotion_burst  {'voice': 1.0, 'emotion': 0.25, 'burst': 0.5} 3.36s  peak 0.680

Note arm 02 against arm 03: the capped emotion dose (0.25, from the burst rule) keeps the line intact at 3.36 s, while the uncapped intense dose (1.9) stacked on the voice adapter cuts it to 1.68 s. Also note that all three adapters declare the same 23 target-module patterns and differ only in rank β€” 4 for the voice, 32 for the other two.

And the trap was reproduced, on purpose. The script then deletes the first-loaded adapter β€” the exact trigger β€” and prints what peft leaves behind:

  active_adapter before delete: 'voice'
  active_adapter after deleting 'voice': 'voice'  (peft_config now holds ['burst', 'emotion'])
  ^ DANGLING. Without the one-line repair in set_active(), the next generate() raises
    KeyError from inside peft_config[self.active_adapter].
  04_after_deleting_first_adapter {'emotion': 0.25, 'burst': 0.5}  3.60s  peak 0.801
  active_adapter after set_active: 'emotion'

PeftModel.active_adapter is still 'voice' after delete_adapter('voice') β€” a pointer to a name that is no longer in peft_config. Every subsequent generate() would raise KeyError: 'voice'. set_active() repairs it on the next call and generation proceeds normally. That is the whole bug, in four lines of output.

Both runs are bit-reproducible: the same arms on two different nodes produced identical durations and peaks.

Choosing Ξ»

Measured, not guessed. Sources are the recipe pages linked at the bottom.

adapter Ξ» why
voice (this repo) 1.0 it is the identity; the ablation below is about rank, not dose
emotion, moderate 0.5 flat across all 40
emotion, intense 0.5 – 1.9, per emotion measured per emotion. e.g. Awe 0.5, Sadness 1.25, Fear 1.75, Anger 1.9. Using 1.9 for Awe overdrives it.
emotion, on a line that also carries a burst ≀ 0.5 Γ— burst Ξ» see below
vocal burst 0.5 measured optimum for a mid-utterance burst: presence 50.3 %, tail coverage 0.90
VoiceNet dimension per dimension and direction best dose measured per (dim, direction) in 0.25–1.25; Γ—0.40 for a "somewhat/notably" step instead of an extreme. Some are traces: vn_VULN__low wants 0.06–0.18, not a full dose
mediathek (German broadcast) 0.25
sports commentator 1.0
explicitness 0.8
character archetype 1.0 the character is the condition in that block

The burst/emotion interaction is the one measured conflict on record. At burst Ξ» = 0.5:

emotion Ξ» burst presence blend genuineness
0.00 0.505 4.76 1.89
0.25 0.569 4.90 1.91
0.50 0.441 4.45 1.57

Emotion at half the burst dose beats both dropping it and matching it, on all three metrics. So the rule is a cap, not a set: Ξ»_emotion = min(Ξ»_emotion, 0.5 Γ— Ξ»_burst). A condition already below the cap keeps its own smaller dose.

Five burst classes (hiss, kissing_noises, lip_smack, person_whistling_playfully, slurping_noises) produce no located burst at any dose when asked for mid-utterance; generate them as isolated events instead.

What happens when adapters conflict

  • Two adapters pulling the same modules compound, and the bigger one wins. Every family here targets the same 23 module patterns, so a voice LoRA and an emotion LoRA are always fighting over the same weights. The voice adapters are rank 4; the emotion, burst, VoiceNet and character adapters are rank 32 and Mediathek is rank 64. At equal Ξ» the larger adapter dominates and identity drifts. That asymmetry β€” not any target-module difference β€” is why every non-voice Ξ» in the table above is below 1 unless the adapter is the condition.
  • Overdriving truncates. In the stacking run below, the same sentence and seed produced 3.84 s with the voice adapter alone, 3.36 s with voice + emotion capped at 0.25 + burst 0.5, and 1.68 s β€” under half β€” with the emotion adapter at its intense dose of 1.9 on top of the voice adapter. An overdriven stack does not merely sound wrong; it stops early. (One sentence, one seed: an illustration, not a measurement.)
  • Order does not matter, dose does. LoRA deltas are additive; set_adapter([a, b]) is symmetric. Only the scalings differentiate them.
  • Keep the resident set bounded. Loading many adapters costs VRAM. Evict LRU β€” and re-read trap 1, because eviction is exactly what turns the latent bug into a KeyError.
  • Sort your work by adapter. With 114 dimension adapters and 40 emotion adapters, thrashing the loader between every generation is the easiest way to waste a GPU-hour.

The rank ablation, in full

The headline: rank 4 is enough

9 of the ten voices ship a rank-4 adapter β€” 8.6 M trainable parameters, a quarter of rank 16's 34.4 M and an eighth of the rank 32 the single-voice predecessor shipped. emolia_c0542 ships rank 8 (17.2 M) because rank 4 failed the non-inferiority test for that voice.

Each voice was trained at rank 16, 8 and 4 and the three were compared on held-out groups. The shipped rank is the smallest rank that is not significantly worse than the best rank on speaker similarity (paired t, p β‰₯ 0.05, and no more than 0.03 absolute WER worse). Speaker similarity is the primary axis because these are identity adapters; an argmax on a noisy mean would have answered "16 always" by construction.

Pooled: n = 1,920 held-out clips per arm

Every arm generated the same prompts with the same seeds; only the adapter differs.

arm clips spk-sim reward genuineness blend WER emotion strength
base 1920 0.3938 3.990 0.866 2.58 0.1297 1.500
stage1_r16 1920 0.5933 4.235 0.747 2.89 0.0971 1.462
stage1_r8 1920 0.5976 4.213 0.765 2.87 0.0937 1.461
stage1_r4 1920 0.5944 4.257 0.753 2.89 0.0883 1.465
stage2_r16 1920 0.6017 4.496 0.799 3.09 0.0926 1.456
stage2_r8 1920 0.6034 4.453 0.809 3.08 0.0934 1.459
stage2_r4 1920 0.6040 4.499 0.785 3.11 0.0893 1.462

The two effects, side by side

comparison Ξ” spk-sim 95 % CI p
base β†’ stage2_r4 +0.2102 [+0.2030, +0.2175] < 1e-300
stage2_r16 β†’ stage2_r4 +0.0023 [-0.0014, +0.0060] 0.22
stage2_r16 β†’ stage2_r8 +0.0017 [-0.0018, +0.0052] 0.35
stage2_r16 β†’ stage1_r16 -0.0084 [-0.0123, -0.0046] 1.8e-5

The base→r4 effect is ~90× the size of the r16→r4 gap, and the inter-rank gaps are indistinguishable from zero on every axis measured (reward +0.003, p = 0.94; WER -0.003, p = 0.32). Meanwhile stage 2 of the curriculum is worth something: +0.0084 spk-sim (p = 1.8e-5) and +0.26 reward (p = 1.7e-8) over stage 1.

So the rank knob is not where the quality is. Training these at rank 16 would have cost 4Γ— the adapter parameters to buy nothing measurable.

Per-voice decisions

voice ships best rank spk-sim @ ships @ best Ξ” p WER
anime_088 r4 r16 0.4566 0.4623 -0.0058 0.542 0.0860
emolia_c0542 r8 r8 0.6167 0.6167 +0.0000 1.000 0.0864
emolia_c1682 r4 r4 0.6083 0.6083 +0.0000 1.000 0.1052
emolia_c1699 r4 r4 0.5480 0.5480 +0.0000 1.000 0.0745
emolia_c2570 r4 r4 0.6123 0.6123 +0.0000 1.000 0.1172
k10_age3_bg1 r4 r4 0.6374 0.6374 +0.0000 1.000 0.0758
k325_age3_bg1 r4 r16 0.6369 0.6433 -0.0064 0.163 0.0795
k395_age3_bg1 r4 r8 0.5915 0.5998 -0.0084 0.162 0.1049
k91_age5_bg0 r4 r4 0.6574 0.6574 +0.0000 1.000 0.0684
mediathek_0184 r4 r8 0.6880 0.6951 -0.0071 0.159 0.0886

Four voices have a nominally better rank than the one they ship; in all four the gap is within noise (p β‰₯ 0.16), so the smaller adapter wins. emolia_c0542 is the exception: rank 4 was significantly worse and rank 8 ships.

The subtlety worth stating: validation loss would have been wrong

Held-out validation loss does separate the ranks, monotonically, for every one of the ten voices, without exception: r16 < r8 < r4. If you had run this ablation on loss alone, you would have shipped rank 16 ten times out of ten, with clean, consistent, unanimous evidence.

The size of it is the point:

value
base validation loss 3.94 – 4.20 nats
rank 16 (stage 2) 3.73 – 3.95 nats
mean base β†’ r16 improvement 0.217 nats
mean r4 βˆ’ r16 gap +0.0084 nats
gap as a share of the gain 3.9 % (range 3.2 – 5.3 % across the ten voices)

The generation-based evaluation turns that 3.9 % into no measurable difference at all on 1,920 held-out clips per arm. The two measurements agree on the ordering and disagree only on whether the remaining gap is worth paying for.

The lesson: validation loss alone would have recommended rank 16, and it would have been wrong. A loss difference can be perfectly consistent, perfectly monotonic, replicate across ten independent training problems β€” and still be four times too small to matter in the artefact you actually ship. Ablate on the output, not on the objective. (And note the honest limit of that claim: the generation evaluation is itself a learned scorer, so what has really been shown is that the gap is below the resolution of every instrument that was pointed at it. No human was asked.)


Repo layout

<voice>/adapter_model.safetensors      SHIPPED   the adapter to use
<voice>/adapter_config.json            SHIPPED   rank, alpha = 2r, target modules
<voice>/RECOMMENDED.json               SHIPPED   which rank and stage, and why
<voice>/holdout_gids.json                        the groups this adapter never saw
<voice>/ranks/r16/{stage1,stage2}/     audit     every rank, both curriculum stages
<voice>/ranks/r8/{stage1,stage2}/      audit
<voice>/ranks/r4/{stage1,stage2}/      audit
ablation/pooled.csv                              the pooled arm table above
ablation/per_voice_arm.csv                       10 voices Γ— 7 arms
ablation/decisions.json                          the shipped rank per voice, with p-values
ablation/rank_ablation.json                      everything, including all paired tests
quickstart.py                                    the tested quickstart
stack_adapters.py                                the tested stacking pattern

166 adapter and metadata files, 5.19 GB, plus this card and the two scripts. Use the top level of a voice folder. The ranks/ tree is kept so the comparison can be redone or a different rank chosen deliberately β€” it is not a menu of recommendations.


Training

Two-stage curriculum, identical for all ten voices and all three ranks:

stage 1 stage 2
pool the better half of every group's candidates by the group's own reward ranking the winners only (rank == 0 per (gid, subset))
rows ~18,800 per voice ~1,580 per voice
peak LR 2e-4 5e-5 (0.25Γ—)
epochs 1 1
effective batch 16 (micro 4 Γ— accum 4) 16
schedule linear warmup 5 % then linear decay to 0 warmup 10 %

AdamW (weight_decay=0, betas 0.9/0.999), grad-norm clip 1.0, max 380 codec frames (~30 s), seed 42, bf16, frozen base. LoRA alpha = 2r, dropout 0.05, bias=none, targeting the global Qwen3 stack (q,k,v,o,gate,up,down_proj), the local GPT-2 decoder (c_attn, c_proj, fc_in, fc_out) and all twelve audio heads (audio_lm_heads.0..11).

The three ranks are trained in one process against one shared frozen base, stepping on the same micro-batches with the same seed. Rank is then the only difference between the three adapters β€” same data order, same sampled captions, same dropout draws β€” which is what makes the paired comparison meaningful.

Captions are resampled every epoch from the stored measurements, seeded by (uid, epoch). The dataset's own caption_gen column is not used for training: it reproduces a known-broken generator that slices the first nine words off paragraph-long VoiceNet anchors. The training caption is rebuilt from the same measurements with varied skeletons, synonyms and dimension subsets, and with probability 0.35 the group's authored caption is used instead β€” so the adapter sees both the kind of prompt a user writes and the kind a scorer produces.

Held-out groups are removed before encoding, not at training time, so no held-out audio is ever in the training container. The split is by group and stratified by (block, language): the candidates of a group are the same condition, so holding out candidates rather than groups would leak the line.

Compute: 38.8 GPU-hours on GH200s (698 core-hours) for all ten voices β€” encoding, three ranks Γ— two stages each, and the held-out evaluation.


Limitations, in full

No human listening study was run. Every number on this page is an automatic scorer's output: speaker similarity is ECAPA cosine (speechbrain/spkrec-ecapa-voxceleb); genuineness, blend and emotion strength are learned heads. They have been observed disagreeing with listening judgements. Nothing here has been validated against human preference.

43.4 % of the training corpus falls below the 0.40 speaker-similarity floor. By block:

block takes spk-sim % below 0.40
voicenet 218,880 0.434 40.1
emotion 153,600 0.411 46.8
β€” of which intense, free 38,400 0.363 57.1
β€” of which intense, contained 38,400 0.378 52.4
β€” of which moderate 76,800 0.452 38.7–39.0
edge cases 13,440 0.393 49.6
character 11,520 0.387 52.1
burst isolated 3,200 0.451 36.7
explicit 960 0.398 50.5
sports 960 0.295 78.4
all 402,560 0.422 43.4

Identity is ranked at 0.19 weight, not gated β€” a hard 0.68 gate was rejected because 55 % of genuinely same-speaker pairs fall below it. The consequence is that the adapters were trained on takes that partly drift off the reference. They improve on it (pooled 0.3938 β†’ 0.6040 on held-out clips) but the intense-emotion conditions remain the hardest place to hold a voice, and sports is worse still.

Identity is bought with expressiveness. The adapters raise speaker similarity and lower genuineness β€” the same trade the single-voice predecessor measured (genuineness 0.784 β†’ 0.581 at rank 16). If spontaneity matters more than identity to you, start from the base model.

The reference dataset is superseded-in-waiting. It encodes run PPILOT2 and carries two measured text defects:

defect intended measured how measured
burst tags Title-Case β€” MOSS spells (Growl) out letter by letter instead of performing it lower-case (growl) 135,580 of 135,630 tags (99.96 %) are Title-Case; median 13,544 per voice counting \(…\) spans in the released text column
burst density 50 % of lines carry a burst 33.7 % of rows do same

Both are fixed in the in-flight 500-voice build. Its indices 490–499 are these same ten voices and will replace them. Nothing here is wrong about rank β€” that conclusion is independent of the tag casing β€” but the audio itself will be redone.

WER on this page is Whisper-large-v3-turbo, not the Parakeet used for the dataset's own wer column: the eval process already holds the generator, the scorer and ECAPA, and NeMo cannot be imported beside them. These numbers are comparable across arms, which is what a rank comparison needs, but not to the dataset's column.

Two languages (English, German). Synthetic training data throughout β€” these adapters were trained on MOSS output, scored by learned heads, not on recorded human speech.

anime_088 is weak. Adapted spk-sim 0.4566 is barely above the floor, and 80.4 % of its corpus takes are below it. Use it knowing that.


Related work and further reading

The manual β€” MOSS voice-acting manual. Directly relevant recipe pages:

Companion sites β€” emotion/voice conditions (per-emotion Ξ») Β· VoiceNet manual (per-dimension doses) Β· character voice clusters Β· these ten voices, with audio

Repos β€” see Β§ The ecosystem for the full table with links.

Citation

@misc{moss_voice_profile_loras_2026,
  title  = {MOSS voice-profile LoRAs: ten pilot voices and a rank ablation},
  author = {LAION and TTS-AGI},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/TTS-AGI/moss-voice-profile-loras}}
}

Apache-2.0, matching the base model.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for TTS-AGI/moss-voice-profile-loras

Dataset used to train TTS-AGI/moss-voice-profile-loras