CARA-native Stable Audio Open Small

This is the Phase 2 CARA-native Stable Audio Open Small checkpoint used in the CARA cross-architecture attribution research. The fork exposes batch-aligned native DiT features and trains a checkpoint-owned hierarchical 98-pool/9-family head together with the diffusion model.

The release is an open-weight peer-review artifact, not an OSI open-source model. Its use and redistribution remain subject to the Stability AI Community License.

Contents

  • phase2/native_stable_audio_full.ckpt: completed 7,665-step model/head checkpoint.
  • phase2/native_stable_audio_training_report.json: training contract.
  • registry/: exact CARA pool/family ordering used by the head.
  • evidence/: authoritative prompt-visible and fixed-audio score reports.
  • source/: exact source snapshot used to load and evaluate the checkpoint.
  • LICENSE and NOTICE: required upstream license and attribution.
  • cara_model_manifest.json: byte sizes and SHA-256 values for release files.

Use the immutable phase2-v1 tag, or the Hub commit hash it resolves to.

Reload

hf download sammoran-phd/cara-native-stable-audio \
  --revision phase2-v1 \
  --local-dir cara-native-stable-audio-release

mkdir cara-native-stable-audio-source
tar -xzf \
  cara-native-stable-audio-release/source/cara-native-stable-audio-source.tar.gz \
  -C cara-native-stable-audio-source
cd cara-native-stable-audio-source

Load stabilityai/stable-audio-open-small with stable-audio-tools, attach the included fork's CARAAttributionHead, and load phase2/native_stable_audio_full.ckpt. The included benchmark scripts perform the registry-hash, global-step, native-head, and feature-shape checks before evaluation. The complete invocation is in the cara-native-musicmodels Phase 2 job specification.

Evaluation boundary

On the primary 780-waveform balanced fixed-audio core, this checkpoint scored 7.82% exact top-1, 25.64% top-3, 33.21% pool-derived family accuracy, 6.90% ECE, and 100% registry-valid output. On the earlier 320-row prompt-visible benchmark it scored 23.75% exact top-1, 56.56% top-3, and 99.38% family accuracy.

The fixed-audio result is the primary estimate of audio-conditioned attribution. The much higher prompt-visible family score is dominated by visible semantic taxonomy and is not evidence of reliable exact source identification.

Intended use and limitations

This release is intended for academic reproduction, interface auditing, and controlled attribution experiments. It is not a provenance, copyright identification, royalty allocation, or safety system. The primary fixed-audio core covers 39 of 98 pools and represents one source corpus and one training run.

License

The stable-audio-tools code is MIT licensed. The base model and this derivative checkpoint are governed by the Stability AI Community License. Redistribution must include that agreement and the required NOTICE; commercial conditions depend on the user's circumstances and the current upstream license.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sammoran-phd/cara-native-stable-audio

Finetuned
(5)
this model