AI PLAY Studio

Songs, art and video.
On your machine.

Write a style description and some lyrics. Get a song. No account, no upload, no credits, no per-song cost — nothing leaves the building. No suitable GPU? There is a hosted mode, off by default.

Start with the output

Everything below was made locally on one 16 GB card. Judge it before you read a word of how it works.

A full song, generated locally MiniMax Music 3 from a caption and lyrics. The timed-lyrics pass measured 98% of its words rather than interpolating them.
LTX 2.5 — the default engine. Five seconds at 1280×704, about two minutes of GPU.

MiniMax H3 is supported as a second engine and there is no sample of it here on purpose. The one published previously was rendered at 1280×720 and 56 frames while its caption claimed H3's native 1344×768 — and rendering H3 below native size and below its trained frame count is precisely what makes it look vague, so that clip demonstrated the defect rather than the engine. Measured on the same prompt and seed: H3 at 8 steps 308 s, H3 at 20 steps 660 s, LTX 121 s and visibly better. LTX is the default for that reason.

What it does

One window, six things, and a rule that keeps them out of each other's way.

Music

Full songs from a caption and lyrics, or instrumentals from a structure. Re-roll the mix in about 15 seconds.

Extend & merge

Continue a take, branch it, then flatten the whole chain into one song.

Start from real audio

Feed it a recording and vary it. The interesting one — see below.

Cover art

Drawn automatically while the card is otherwise idle, then embedded in the file.

Stems & timed lyrics

Four-way separation, and word-level .lrc for karaoke and visualisers.

Video & a small editor

Clips under a track, then a timeline to cut them together.

Post-processing never competes with music. Covers, stems and lyrics run only when the queue is empty, and a new song preempts them. That is a rule, not a setting.
The Create view: title, lyrics, style and a Create button
Write a caption and some lyrics. That is the whole interface.
The library grouped into sessions, each track with generated cover art
The library, grouped into sessions. Covers appear on their own while the card is idle.

It ships no models. That is deliberate.

This matters more than the feature list, so it goes first.

AIPLAY Studio contains no model weights and redistributes none. Not one byte. Every capability declares what it needs — in server/models.js, which you can read — and downloads it from the publisher, at your request, on your machine.

Before anything downloads, the Models screen shows you the real byte count, the licence, and whether your card can run it. Then there is one button.

CapabilityDownloadLicence
Music — MiniMax Music 311.9 GBMiniMax Music3 Community
Cover art — FLUX.2 klein 4B12.5 GBApache-2.0
Stem separation — HTDemucs336 MBMIT
Timed lyrics — Whisper large-v33.1 GBMIT
Video — MiniMax H334.5 GBH3 Community ⚠ excludes the EU, UK and South Korea
Video — LTX 2.539.7 GBLTX-2.x Community ⚠ paid above $10M revenue; gated repo
The video licences are genuinely restrictive. Studio treats them as blocking rather than as a footnote — but "we showed you the licence" is not the same as "you are fine". Read them.
The Models screen listing each capability with its size and licence
Nothing downloads until you ask. The size and the licence are on screen first.
The Thanks page listing every model with its licence, each name linking to the publisher
The other half of the same claim: every model credited, every licence named, and each name linking to the publisher it came from. The H3 entry carries its territory restriction in the open, because that one is not ours to soften.

It is a face on ComfyUI, and it says so

Everything Studio does, it does by submitting a graph to ComfyUI over its HTTP API. Not linked, not bundled, not modified. ComfyUI is GPL-3.0 and a separate program, running as its own process, exactly as its authors shipped it.

You install ComfyUI yourself. That is a real cost and I am not going to pretend otherwise.

The upside is that the graphs are the product. The pipelines Studio actually submits are in workflows/ as ordinary ComfyUI JSON, exported straight from the code that builds them. Drag one onto the canvas and every value is right there — and those values are not ComfyUI's defaults. They were arrived at by measurement, and the measurement scripts ship in scripts/.

A claim nobody can re-run is just an assertion.

Pictures that change on the beat

The Reactive page analyses a song, finds its peaks, and cross-fades your reference images against each other in time with them — so a new picture arrives on the hit rather than on a timer. Images to video, video to video, or text to video.

It is optional, and it is not part of a Studio install. The pack that does the work — ComfyUI_Yvann-Nodes, by Yvann Barbot and Lilia — is GPL-3.0, and GPL code cannot ship inside an Apache-2.0 application. So it runs on a second ComfyUI you set up yourself, which Studio talks to over HTTP: the same boundary described above, one step further out. Studio never downloads, installs or redistributes any of it.

With no engine running the page shows the setup, naming each pack and its licence, rather than controls that fail when you click them.

Why it cannot simply live in Studio’s own engine: Studio pins a reference picture at a chosen frame, and measured across four strengths and several densities, that either replaces the frame or does nothing. The reactive effect needs two pictures conditioning every frame at once, with the mix travelling in time with the audio. That is a difference in kind, not in tuning.

The Reactive page: three modes, a song picker, and a grid of reference images to blend between
Pick a song and two or more reference images. The peaks decide when a new picture arrives; the pictures decide what it looks like. If the second engine is not running, this page shows the setup instead — naming each pack and its licence rather than offering buttons that fail.

Setup, licences and the traps: REACTIVE.md.

Installing it

A zip and a launcher. Three steps, and the third one is a double-click.

  1. Node.js — the LTS installer, about 30 MB.
  2. ComfyUI — the Windows portable NVIDIA build is the easy route. Run it once on its own first: that launch is what proves your card, your driver and PyTorch agree with each other, and it is far easier to read that failure in ComfyUI's own window.
  3. Double-click Start AIPLAY Studio.cmd.

On first run it fetches exactly one npm package (ws, MIT — that is the entire dependency list), finds your ComfyUI, and opens your browser at 127.0.0.1:4173.

It is a short batch file rather than an .exe, on purpose. Open it in Notepad and read exactly what it is about to do before you let it do it.

You do not need to set up PyTorch or CUDA yourself — ComfyUI's portable build brings its own. One caveat worth knowing: install the cu130 build. A cu128 build runs, and runs about 4.9× slower, with nothing on screen to explain why. Studio checks at startup and says so rather than letting you wonder.

Full walkthrough, including the handful of things that actually go wrong on a first run, is in INSTALL.md.

What your machine needs

Music and video are completely different questions, so here are two answers. Everything was measured on an RTX 4070 Ti SUPER (16 GB), 32 GB RAM, Windows 11.

For music

GPUNVIDIA, 6 GB min · 12 GB recommended
RAM16 GB min · 32 GB recommended
Disk12 GB

There is no CPU fallback. The first stage requires CUDA. AMD, Intel and Apple graphics will not run this, and macOS is out for that reason rather than a packaging one.

For video

GPU16 GB VRAM minimum — either engine
RAM32 GB minimum
Disk34.5 GB (H3) or 39.7 GB (LTX)

Video and cover art are never resident alongside the music engine on a 16 GB card. That is exactly why they wait for idle time instead of fighting for it.

What "fast" means, measured

TaskOn the 16 GB card
Engine cold start~15 s
A fresh song39 s of audio in 65 s — 1.66× realtime
A full-length render135 s of audio in ~207 s — ~1.5× realtime
Re-rolling the mix50 s → 15 s
A cover~3 s
Stems, 30 s track~12 s
A 5 s clip — LTX 2.5, 1280×704121 s
The same clip — H3, 1344×768308 s at 8 steps · 660 s at 20
The small-VRAM tiers are simulated. 6 GB and 8 GB were tested on a 16 GB card with a reserved-VRAM budget, not on real hardware. Nothing ran out of memory, but the proxy is imperfect. I would rather say "unproven" than publish a minimum I have not seen hold.

No GPU? There is a second way in.

Same model, someone else's hardware, your API key. Off by default.

The numbers above rule out most laptops. But the library, cover art, timed lyrics, the studio timeline and overnight batching have nothing to do with a GPU — and that is the majority of the app. API mode runs the music on a hosted MiniMax Music 3 so all of it works on a machine that could never load the model, and switching to a local GPU later is a setting rather than a migration. Same library, same files.

Studio calls the provider directly from your machine. Nothing is proxied through AI PLAY, so the account and the transaction are yours and there is no middle for keys to pool in.

Two things change, and Studio says both on screen. It costs money per song — about $0.36 for three minutes — so there is a hard monthly cap, default $20, checked immediately before every call rather than when a batch is queued. And audio reference stops working: it encodes a real recording into the model's own latent, and hosted endpoints take text and return audio with no latent to hand them. The control is disabled with the reason, not left to fail at submit.

Your key is stored with Windows DPAPI, tied to your Windows account and that machine — a copied settings file is inert anywhere else. It is write-only across Studio's own interface: the browser is told a key exists, how it is protected and its last four characters, never the key. On platforms without DPAPI it falls back to a 0600 file and says so, because file permissions are not encryption.

Or bring your own graph

Studio's pipelines are tuned, and that tuning is most of what it knows — but they are opinionated: one image model, two video engines, a fixed sampler. If you already live in ComfyUI, drop an API-format graph into workflows/custom/ and pick it in Settings. Studio fills in a few placeholders and submits it otherwise untouched. It does not try to understand your graph.

If you hand it ComfyUI's ordinary Save instead of Save (API Format), Studio says so and tells you which one to use. Both are valid JSON, both load without complaint, and only one can execute — it is the most common way this goes wrong.
Settings: engine paths, graphics-memory tier, API mode and custom workflows
One page for the things that decide how everything else behaves — where ComfyUI lives, how much graphics memory to assume, and whether the music runs locally or on someone else's hardware. Your own graphs are picked here too.

Video, loops, and a small editor

Two engines, a starting frame, a closing frame, and somewhere to put the results.

Loop videos are the thing people actually make, and the way you make one is boring: end on the picture you started from. The clip has nowhere to go but back to its own first frame, so it cuts to the start with no seam.

There is a real worry here — a lot of models freeze when both ends are the same, satisfying the constraint by not moving. This one does not, and the reason is a number rather than a virtue: the end frames are pinned at strength 0.7, not 1.0. At 1.0 the ends dominate and the middle stalls. Measured on the shipped setting: motion 1.90, mid-point divergence 6.00, loop closure 1.62.

The Video view with a starting frame and a closing frame, both previewed
Starting frame and closing frame, each previewed — with a warning underneath, because covers are square and every render size is 16:9 or 9:16.
The studio: a multi-track timeline with two clips overlapping into a crossfade
The studio is shaped like Vegas on purpose: stacked tracks, mute and solo, and a crossfade that is an overlap — slide one clip over the end of another and the dissolve is there. Karaoke comes from the .lrc Studio already generates.
The clip library beside the Video form, grouped by day, each clip showing its render time
Every clip keeps what made it: engine, size, and how long it took. Grouped by day or by track, because a render you cannot find again is a render you will do twice.

Feeding it a real song

The part worth reading. ComfyUI cannot encode audio for this model. The model can.

MiniMax Music 3 has two stages. The composition stage speaks in eight discrete token streams whose tokenizer was never published — so "cover this exact song" is genuinely closed, and the best community substitute measures 41% semantic and 7.2% acoustic top-1 accuracy, with several heads near chance. That is a dead end with numbers attached.

The render stage is different. It uses a continuous 64-dimensional latent, and its encoder had been sitting in the HuggingFace cache the whole time — the weights were already on disk. Encode real audio into that latent, hand it back as a starting point, and you get variation from a real recording with no training and no tokenizer. Round trip: +26.26 dB, r = 0.999.

The dial is denoise. Same reference, three settings:

denoise 0.60 — still a copy Waveform correlation 0.94 against the near-copy end of the dial. It is not remixing; it is repeating.
denoise 0.85 — the shipped default Correlation has collapsed to 0.004. A new waveform that still follows the reference.
denoise 0.95 — past the useful range The reference has stopped steering it.
Where the cliff is. Waveform correlation against the near-copy end: 0.4 → 0.987, 0.6 → 0.942, 0.8 → 0.708, then 0.85 → 0.004. That collapse is the dial ceasing to copy and starting to compose. Correlation near zero does not mean the reference is ignored — it means the output is a new waveform, which is what a remix is.
A measurement that lied. A packed-stereo interleaving bug in the audio loader produced two confident, wrong conclusions before anyone caught it — and a round-trip SI-SDR cannot catch that class of bug, because it happily scored 14 dB on garbage. Only decoding the latent and comparing it to the real file exposed it. Measure the thing you are about to call proof.

The extras, in plain language

Stem separation

Splits a finished track into drums, bass, vocals and other — four files for a DAW. Off by default: most takes get discarded, and the ones worth pulling apart are the ones you already starred.

Timed lyrics

A .lrc at line level and at word level, for karaoke and visualisers. It reports what share of words were timed by measurement rather than interpolated, so you know how far to trust it.

Cover art

A picture per track, drawn while the card is idle and embedded into the audio file itself, so it travels with the song.

Overnight runs

A list of ideas, N takes each, and a full library by morning — with the post-processing stages you choose.

Planning an overnight run
Plan a run, pick which stages it should do, leave it.
The Images page: a prompt on the left, a grid of generated pictures on the right
Cover art is one use of it. The Images page is the general one — and the pictures it makes are what the Reactive page blends between.