Latent Music World Model

Submission 14423

THE IDEA

Abstract

Latent dynamicsHierarchical planningControllable musicLong-horizon generation

Long-horizon controllable symbolic music generation requires persistent musical state, explicit actions, and structural organization across recursive rollout. We formulate this problem as plan-conditioned latent world modeling and develop a hierarchical, action-conditioned Music World Model (MWM) with fast segment states and slow phrase-level plans. Recursive MWM deployment creates a predicted-plan structural drift pathway: later planner inputs contain representations produced by earlier model rollouts. Under matched conditions, we compare a predicted-plan (PP) path with a teacher-plan (TT) path. PP recursively uses its predicted executor-facing slow trajectory, whereas TT replaces only that future trajectory with a target-derived reference. This PP–TT diagnostic isolates the additional divergence associated with recursive structural self-conditioning while holding the executor, actions, codec, decoder, and rollout horizon fixed. Across chained long-form rollouts, the discrepancy concentrates in late phrase transitions and is inherited by subsequent planner contexts. We then introduce Boundary-Aware History Correction (BHC), a training-free, parameter-free inference procedure that combines a local context-regularized anchoring correction with a persistent forward-decaying boundary displacement. The intervention leaves the generator frozen and uses no future teacher information at inference time. Across 16–64-bar validation rollouts, song-level decoded and internal trajectory analyses show that BHC suppresses part of the PP–TT divergence, with the clearest gains after structural mismatch has accumulated. Controlled mismatch experiments further show that the recoverable gain increases with the injected gap. These results support BHC as a targeted suppressor of predicted-plan structural drift, rather than a guarantee of global stability or drift elimination.

MODEL OVERVIEW

Model architecture

From musical context to future states
Comparison of diffusion-based music generation, transformer-based music generation, and our latent music world model. The world model encodes musical inputs, evolves latent states under an action policy, and decodes future outputs. View full size
FIGURE 01 A comparison of music generation paradigms. The latent world model represents musical state and models its evolution under explicit actions.
01 /

Represent

Encode musical context into a compact latent state.

02 /

Plan

Predict phrase-level structure from the evolving history.

03 /

Generate

Roll out segment states under actions, then decode the music.

Recording status. Five selected songs include real ground truth, our continuations, and six baseline models. BHC and ablations remain labelled previews.

PARTIAL PREVIEW
01 / MUSIC GENERATION

Autonomous continuation

Five musical contexts, each followed by a 32-bar multi-track continuation. The same contexts are used in the control experiments below.

05 musical contexts

Listen from the beginning, or skip the observed context to hear the continuation.

Context Generated
Example selection and verification

We screened the existing 1,760 validation windows from 233 songs for lower repetition while retaining clear control responses. Every selected window passes all 11 control groups and all 16 individual checks, has an active melody and nonempty restoration targets, and keeps its continuation active through 64 bars. Exported MIDI notes also meet precision, recall, restoration and background-preservation thresholds.

Among eligible windows, we prioritized lower overlap between adjacent eight-bar phrases in both the arrangement and melody, including the later continuation, and selected one window per song. The recordings are unchanged model outputs with the same control settings as before. This is a retrospective qualitative selection, not an unbiased evaluation or a human listening-quality ranking. Recurring motifs remain.

Selection and before/after audit · Original song ranking · Verification records

Adjacent eight-bar note overlap (lower = less literal repetition); control score for the displayed window
ContextBPMRepeat · 32 barsRepeat · 64 barsControl / 100Checks passed
02 / MODEL COMPARISON

Ground truth and generation

Listen to the same five musical contexts: real ground truth, our continuation, and six baseline models prompted with GT music.

GT / OURS / 6 BASELINES

All players share the same 8-bar GT prelude, playback tempo, SoundFont, and gain. Baselines receive the observed music only; the actual next 32 bars are withheld. Short outputs retain a silent remainder, with their last sounding bar noted on each card. Generation and listening protocol · Model and recording provenance · GT sources

The requested continuation is 32 bars. For Context 2, BEAT and Text2midi use the most recent 4 observed bars to fit their native context limits. Ours retains its original history of up to 48 observed bars, while baseline inputs contain 8 bars (or the noted 4-bar fallback). MIDI and ABC adapters differ in quantization and expressive detail. These fixed qualitative examples use model-specific inference settings; they are not a controlled model ranking.

03 / LONG-HORIZON GENERATION

Long-horizon continuation

Compare 32- and 64-bar excerpts from the same autonomous rollout for each song.

Listen for recurring motifs, evolving texture, and how the arrangement develops at phrase boundaries.

04 / CONTROL DIMENSIONS

Controlled generation

Compare control settings on the same five musical contexts, each paired with an autonomous reference.

06 control dimensions
05 / ACTION SEQUENCES

Phrase-level action sequences

Apply one action every eight bars and compare the response against the same-context autonomous reference.

08 action schedules
06 / STRUCTURAL DRIFT & CORRECTION

Structural drift and BHC

Compare predicted plans, boundary-aware correction, and a teacher-plan diagnostic reference.

PP / BHC / TT
PP

Predicted planRecursive generation under the model’s own structural plan.

TT

Teacher-plan referenceUse target-derived slow structure for diagnosis while retaining recursive execution.

These slots currently contain preview recordings, not matched PP / BHC / TT results. TT is a diagnostic reference, not a deployable generation setting.

07 / MECHANISM ABLATIONS

BHC ablations

Compare the complete correction with its individual components.

04 correction variants
Listening guide & recording notes

What to listen for. Compare musical continuity, phrase transitions, density, instrument roles, and unintended changes. Use the same context across controls, and switch players to keep the listening position aligned.

Recording status. Generation, horizons, and controls use five songs selected for lower phrase repetition among windows that passed every control check. The same settings are used across songs. The comparison pairs each generated continuation with the actual 32-bar ground truth. Six baseline models have real GT-conditioned recordings; BHC and ablations remain labelled previews. Selection scores do not measure human listening quality.

Control meanings. Role restoration uses an explicit MIDI guide from the corresponding autonomous reference. Mute and density use output constraints; constrained music feeds subsequent generation. Guidance strength measures target following, not volume. Recordings use one SoundFont and common gain, with no per-clip normalization, fixed prefix tempo, and BHC disabled.

Playback. Only one recording plays at a time. The floating player follows the active recording. MP3 and MIDI links download the files associated with each slot. Verification records