Latent Music World Model

Submission 14423

THE IDEA

Abstract

Latent dynamicsHierarchical planningControllable musicLong-horizon generation

Long-horizon controllable symbolic music generation requires persistent musical state, explicit actions, and structural organization across recursive rollout. We formulate this problem as plan-conditioned latent world modeling and develop a hierarchical, action-conditioned Music World Model (MWM) with fast segment states and slow phrase-level plans. Recursive MWM deployment creates a predicted-plan structural drift pathway: later planner inputs contain representations produced by earlier model rollouts. Under matched conditions, we compare a predicted-plan (PP) path with a teacher-plan (TT) path. PP recursively uses its predicted executor-facing slow trajectory, whereas TT replaces only that future trajectory with a target-derived reference. This PP–TT diagnostic isolates the additional divergence associated with recursive structural self-conditioning while holding the executor, actions, codec, decoder, and rollout horizon fixed. Across chained long-form rollouts, the discrepancy concentrates in late phrase transitions and is inherited by subsequent planner contexts. We then introduce Boundary-Aware History Correction (BHC), a training-free, parameter-free inference procedure that combines a local context-regularized anchoring correction with a persistent forward-decaying boundary displacement. The intervention leaves the generator frozen and uses no future teacher information at inference time. Across 16–64-bar validation rollouts, song-level decoded and internal trajectory analyses show that BHC suppresses part of the PP–TT divergence, with the clearest gains after structural mismatch has accumulated. Controlled mismatch experiments further show that the recoverable gain increases with the injected gap. These results support BHC as a targeted suppressor of predicted-plan structural drift, rather than a guarantee of global stability or drift elimination.

MODEL OVERVIEW

Model architecture

From musical context to future states
Comparison of diffusion-based music generation, transformer-based music generation, and our latent music world model. The world model encodes musical inputs, evolves latent states under an action policy, and decodes future outputs. View full size
FIGURE 01 A comparison of music generation paradigms. The latent world model represents musical state and models its evolution under explicit actions.
01 /

Represent

Encode musical context into a compact latent state.

02 /

Plan

Predict phrase-level structure from the evolving history.

03 /

Generate

Roll out segment states under actions, then decode the music.

Recording status. The five selected songs include real ground-truth and generated continuations. Other models' recordings are not yet available. BHC and ablations remain labelled previews.

PARTIAL PREVIEW
01 / MUSIC GENERATION

Autonomous continuation

Five musical contexts, each followed by a 32-bar multi-track continuation. The same contexts are used in the control experiments below.

05 musical contexts

Listen from the beginning, or skip the observed context to hear the continuation.

Context Generated
Example selection and verification

The full continuation validation set contains 233 songs and 1,760 windows. Each window uses the same 16 generation configurations. Scores are averaged equally over 11 evaluation categories, then over all windows of each song, including failed conditions.

These five songs have the highest mean scores under this fixed protocol. Each displayed window is representative of its song's average performance. This is a selected qualitative showcase, not an unbiased evaluation or a human listening-quality ranking.

Scoring protocol · Full ranking · Verification records

Song-level scores under the selection protocol
ContextBPMWindowsMean / 100
02 / MODEL COMPARISON

Ground truth and generation

Compare the actual next 32 bars with our autonomous continuation for each of the same five songs.

GT / OURS

GT uses the source song's actual quantized notes and velocities. Both recordings share the same 8-bar context, playback tempo, SoundFont, and gain. Other models' recordings are not yet available. GT source and verification

03 / LONG-HORIZON GENERATION

Long-horizon continuation

Compare 32- and 64-bar excerpts from the same autonomous rollout for each song.

Listen for recurring motifs, evolving texture, and how the arrangement develops at phrase boundaries.

04 / CONTROL DIMENSIONS

Controlled generation

Compare control settings on the same five musical contexts, each paired with an autonomous reference.

06 control dimensions
05 / ACTION SEQUENCES

Phrase-level action sequences

Apply one action every eight bars and compare the response against the same-context autonomous reference.

08 action schedules
06 / STRUCTURAL DRIFT & CORRECTION

Structural drift and BHC

Compare predicted plans, boundary-aware correction, and a teacher-plan diagnostic reference.

PP / BHC / TT
PP

Predicted planRecursive generation under the model’s own structural plan.

TT

Teacher-plan referenceUse target-derived slow structure for diagnosis while retaining recursive execution.

These slots currently contain preview recordings, not matched PP / BHC / TT results. TT is a diagnostic reference, not a deployable generation setting.

07 / MECHANISM ABLATIONS

BHC ablations

Compare the complete correction with its individual components.

04 correction variants
Listening guide & recording notes

What to listen for. Compare musical continuity, phrase transitions, density, instrument roles, and unintended changes. Use the same context across controls, and switch players to keep the listening position aligned.

Recording status. Generation, horizons, and controls use five songs selected by their mean scores across all validation windows. The same settings are used across songs. The comparison pairs each generated continuation with the actual 32-bar ground truth. Other models have no recordings yet; BHC and ablations remain labelled previews. Selection scores do not measure human listening quality.

Control meanings. Role restoration uses an explicit MIDI guide from the corresponding autonomous reference. Mute and density use output constraints; constrained music feeds subsequent generation. Guidance strength measures target following, not volume. Recordings use one SoundFont and common gain, with no per-clip normalization, fixed prefix tempo, and BHC disabled.

Playback. Only one recording plays at a time. The floating player follows the active recording. MP3 and MIDI links download the files associated with each slot. Verification records · Earlier three-song examples