Abstract
Long-horizon controllable symbolic music generation requires persistent musical state, explicit actions, and structural organization across recursive rollout. We formulate this problem as plan-conditioned latent world modeling and develop a hierarchical, action-conditioned Music World Model (MWM) with fast segment states and slow phrase-level plans. Recursive MWM deployment creates a predicted-plan structural drift pathway: later planner inputs contain representations produced by earlier model rollouts. Under matched conditions, we compare a predicted-plan (PP) path with a teacher-plan (TT) path. PP recursively uses its predicted executor-facing slow trajectory, whereas TT replaces only that future trajectory with a target-derived reference. This PP–TT diagnostic isolates the additional divergence associated with recursive structural self-conditioning while holding the executor, actions, codec, decoder, and rollout horizon fixed. Across chained long-form rollouts, the discrepancy concentrates in late phrase transitions and is inherited by subsequent planner contexts. We then introduce Boundary-Aware History Correction (BHC), a training-free, parameter-free inference procedure that combines a local context-regularized anchoring correction with a persistent forward-decaying boundary displacement. The intervention leaves the generator frozen and uses no future teacher information at inference time. Across 16–64-bar validation rollouts, song-level decoded and internal trajectory analyses show that BHC suppresses part of the PP–TT divergence, with the clearest gains after structural mismatch has accumulated. Controlled mismatch experiments further show that the recoverable gain increases with the injected gap. These results support BHC as a targeted suppressor of predicted-plan structural drift, rather than a guarantee of global stability or drift elimination.
Model architecture
View full size
Represent
Encode musical context into a compact latent state.
Plan
Predict phrase-level structure from the evolving history.
Generate
Roll out segment states under actions, then decode the music.
Autonomous continuation
Five musical contexts, each followed by a 32-bar multi-track continuation. The same contexts are used in the control experiments below.
Listen from the beginning, or skip the observed context to hear the continuation.
Context GeneratedExample selection and verification
The full continuation validation set contains 233 songs and 1,760 windows. Each window uses the same 16 generation configurations. Scores are averaged equally over 11 evaluation categories, then over all windows of each song, including failed conditions.
These five songs have the highest mean scores under this fixed protocol. Each displayed window is representative of its song's average performance. This is a selected qualitative showcase, not an unbiased evaluation or a human listening-quality ranking.
| Context | BPM | Windows | Mean / 100 |
|---|
Ground truth and generation
Compare the actual next 32 bars with our autonomous continuation for each of the same five songs.
GT uses the source song's actual quantized notes and velocities. Both recordings share the same 8-bar context, playback tempo, SoundFont, and gain. Other models' recordings are not yet available. GT source and verification
Long-horizon continuation
Compare 32- and 64-bar excerpts from the same autonomous rollout for each song.
Listen for recurring motifs, evolving texture, and how the arrangement develops at phrase boundaries.
Controlled generation
Compare control settings on the same five musical contexts, each paired with an autonomous reference.
Phrase-level action sequences
Apply one action every eight bars and compare the response against the same-context autonomous reference.
Shape the musical density
Click a phrase to listen. Switch players to compare the same moment.
Structural drift and BHC
Compare predicted plans, boundary-aware correction, and a teacher-plan diagnostic reference.
Predicted planRecursive generation under the model’s own structural plan.
Boundary-aware correctionAnchor the slow trajectory and carry a decaying boundary correction forward.
Teacher-plan referenceUse target-derived slow structure for diagnosis while retaining recursive execution.
These slots currently contain preview recordings, not matched PP / BHC / TT results. TT is a diagnostic reference, not a deployable generation setting.
BHC ablations
Compare the complete correction with its individual components.
Listening guide & recording notes
What to listen for. Compare musical continuity, phrase transitions, density, instrument roles, and unintended changes. Use the same context across controls, and switch players to keep the listening position aligned.
Recording status. Generation, horizons, and controls use five songs selected by their mean scores across all validation windows. The same settings are used across songs. The comparison pairs each generated continuation with the actual 32-bar ground truth. Other models have no recordings yet; BHC and ablations remain labelled previews. Selection scores do not measure human listening quality.
Control meanings. Role restoration uses an explicit MIDI guide from the corresponding autonomous reference. Mute and density use output constraints; constrained music feeds subsequent generation. Guidance strength measures target following, not volume. Recordings use one SoundFont and common gain, with no per-clip normalization, fixed prefix tempo, and BHC disabled.
Playback. Only one recording plays at a time. The floating player follows the active recording. MP3 and MIDI links download the files associated with each slot. Verification records · Earlier three-song examples