Skip to content
AI Primer
release

Pika Music launches 4-input diffusion model via API Club

Pika Music accepts text, lyrics, voice, and music references separately or together in one diffusion decoder. It is available through Pika API Club, and Pika claims up to 10 times better cost efficiency than other music models.

3 min read
Pika Music launches 4-input diffusion model via API Club
Pika Music launches 4-input diffusion model via API Club

TL;DR

Pika says a vocal reference can shape a performance's character and delivery, while a reference track can seed a new song in its audio-model launch. The same release puts Music beside video-to-soundtrack, sound-effects, and speech models in one API catalog.

Four inputs, one decoder

The launch post assigns four distinct jobs to the conditioning inputs:

  • Text prompt: musical direction.
  • Lyrics: the words to perform.
  • Vocal reference: vocal character and delivery.
  • Music reference: a starting point for a fresh track.

Pika says those signals meet in one shared latent-diffusion decoder. Its technical description lists mixed local and full attention, flow matching, and specialized conditioning for text, lyrics, voice, and source-song context.

Six-minute tracks, six-second renders

Pika says Music can output tracks up to six minutes long. For 90-second output, the official launch reports a 6.21-second local-test average, while Pika's follow-up gives 6.25 seconds, with both figures described as roughly 14.5 times faster than playback.

The company attributes that latency to its optimized diffusion path. Pika's technical post names the attention design, flow matching, and multimodal conditioning as the relevant components.

Pika's up-to-10-times cost-efficiency claim compares provider list prices current on August 14, according to the launch post. It is a price comparison, separate from the quality evaluation below.

API Club access

Pika Music is initially available only through API Club. Pika's API Club announcement says membership starts at $10 a month and provides one API for more than 100 video, image, audio, and LLM models.

The Club's broader claim is prices up to 88% below other API aggregators; Music's up-to-10-times figure is its own model-specific comparison. Pika's sample post also demonstrates the intended workflow of layering multiple inputs into a track.

Fixed-lyrics benchmark

Pika evaluated fixed-lyrics output across 15 creative directions spanning genres, languages, BPMs, keys, and time signatures. Its benchmark post called the result top-tier on Audiobox Aesthetics, while reporting Suno ahead on both published measures:

  • Content Enjoyment: Pika Music, 7.403; Suno, 7.412.
  • Production Quality: Pika Music, 8.064; Suno, 8.151.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR2 posts
Four inputs, one decoder1 post
Six-minute tracks, six-second renders1 post
API Club access1 post
Fixed-lyrics benchmark1 post
Share on X