SheetSage2 for MLX Swift โ€” 4-bit prequantized pack

Music recording โ†’ editable ABC score, on Apple silicon (Mac, iPhone), with yue2-mlx-swift (SheetSage2Core, yue2 transcribe), for the 4bit-fast / 4bit-lean transcription profiles.

Source and changes

A derivative of SheetSage2 (m-a-p/SheetSage2, revision 398b22834dac) and of its encoder parent MERT2 (m-a-p/MERT-v2-FullSong, revision d8ba1c745e73), both by m-a-p, CC BY-NC 4.0. Changes made here:

  • the LoRA adapters merged into MERT2's attention projections (upstream merge_lora, float32);
  • the Conformer encoder's linear layers quantized to 4 bits (MLX affine, group size 64); the decoder in float16; the mel front-end buffers, the ConvNeXt front and the layer-mix logits in float32;
  • stored in MLX layout (weights_format: mlx-quantized), SHA-256 in model.safetensors.sha256. No retraining. The weights are not endorsed by m-a-p. Quantization changes some scores against the float16 model (rubato orchestral material especially); see the repository's docs/References.md.

Use

yue2 download --model sheetsage2-q4
yue2 transcribe --audio song.m4a --profile 4bit-lean --out run/score

The pack is bit-identical to what the 4bit-* profiles build by quantizing the float16 model at load, without the cost: it loads directly at its own size.

License

CC BY-NC 4.0 (see LICENSE, copied from the upstream release): non-commercial use only, with attribution to MERT2 and SheetSage2 (m-a-p) and their repositories above.

Downloads last month
25
Safetensors
Model size
0.7B params
Tensor type
F16
ยท
F32
ยท
U32
ยท
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for VincentGOURBIN/sheetsage2-mlx-q4

Finetuned
(3)
this model