Agent-run Hollywood casting sessions on
laion/moss-tts-local-transformer-4.55b-voice-acting-v2:
nine autonomous agents, one acting brief each, three rounds apiece.
Each ~30 second performance is written by the agent as 2–4 parts (the model truncates long before 30 s, and one generation cannot swing between opposed emotions), generated best-of-16 per part, then assembled by scoring every combination as a whole.
The pages show, per take: the agent's scene and intended arc, the exact GENERAL/SCRIPT prompt
for every part, the spoken words alone, every LoRA and its merge dose, the sampling, what audio
served as the voice reference — and both the agent's own score and a listening model's.
This repo exists because it is small. The main model repo's GitHub Pages build started failing once it passed ~500 MB, so large audio-bearing demo pages live here instead.