self-hosted/ai
§01·model · /models

MOSS-Audio

multimodalactiveApache-2.0

4B open audio-understanding model by the OpenMOSS team — it reads acoustic cues, speakers, emotion and environmental sound, so it goes well past transcription. Audio in, text out: like Voxtral, it understands speech rather than producing it. Fifteen cards carry a recipe here from a 12 GB floor, where the six 12 GB cards are the tight ones. Apache-2.0.

§02·GPUs that run this model
15 total

benchmarked·~ runs via recipe (not benchmarked)· untested·doesn't fit