Plan, Generate, ReplanHierarchical Symbolic Music Generation
with Boundary-Aware History Correction

Submission 14432

THE METHOD

Method Overview

Our architecture treats music generation as a latent world model with two interacting time scales. A phrase-level planner updates at a lower frequency to determine the next section’s structural direction, while a segment-level executor runs at a higher frequency to turn that plan into local musical states and concrete multi-track events. The two modules exchange recursively updated latent states instead of repeatedly passing complete discrete note sequences, allowing the system to preserve long-range structure without directly modeling an unwieldy full-length sequence. User actions are first-class inputs to the same generation interface: density changes, MIDI guidance, and the addition, removal, or restoration of melody, chords, bass, and drums are handled through one action-conditioned executor rather than separate task-specific models. After each generated segment, the system writes the local state and phrase summary back into a structured history, which the planner uses for the next rollout step. This planning–execution–feedback loop gives one model a unified interface for long-form continuation, phrase-level control, and multi-track editing.

GENERATION METHODS COMPARISON

Generation Methods Comparison

From musical context to future states
Comparison of diffusion-based music generation, transformer-based music generation, and our latent music world model. The world model encodes musical inputs, predicts latent-state transitions under explicit actions, and decodes future outputs.
FIGURE 01 A comparison of music generation paradigms. The latent world model represents musical state and models its evolution under explicit actions.
01 /

Represent

Encode musical context into a compact latent state.

02 /

Plan

Predict phrase-level structure from the evolving history.

03 /

Generate

Roll out segment states under actions, then decode the music.

01 / MODEL COMPARISON

Model Comparison

Listen across five fixed musical contexts: the ground-truth continuation, Music World Model, and six baseline systems. The reference row shows the actual future, while each model generates a continuation from the observed music. Switch players within a column to hear how different generation methods extend the same starting point.

REFERENCE / OURS / 6 BASELINES

Every card begins with the same eight-bar observed prelude. Reference rows show the held-out continuation; model rows show autonomous generations from the observed context. Playback uses each context's fixed tempo, shared SoundFont, and common gain, with no per-clip normalization.

02 / LONG-HORIZON COMPARISON

32-Bar Comparison

Compare 32-bar continuations from the same five musical contexts. Each recording begins with eight bars of observed music, followed by either the reference continuation or music generated by Music World Model, Amadeus-S, NotaGen-RL3, or Music2Music-PT. Listen for how each model develops motifs, maintains the arrangement, and connects successive phrases as generation moves beyond the starting context.

REFERENCE / MWM / 3 BASELINES

The reference row contains the actual next 32 bars. Model continuations are generated without access to that future, and each baseline’s later steps use its own generated history. Recordings share the playback tempo, SoundFont, and gain within each context. Switch players within a column to compare the same passage across systems.

Generation settings and recording details

Amadeus-S and NotaGen-RL3 generate four successive eight-bar segments, conditioning each step on the latest eight bars. Music2Music-PT generates sixteen two-bar segments, using the latest two bars to fit its token limit. MWM retains its original history of up to 48 observed bars.

The five source windows and inference settings are fixed. Decoded timing is preserved within the fixed listening window.

03 / EXTENDED MWM GENERATION

64-Bar MWM Showcase

Follow Music World Model through 64 generated bars from each of the same five starting contexts. After the eight-bar observed prelude, the model continues for eight consecutive phrases, carrying its generated musical history forward. Listen to the second half for recurring motifs, changes in texture, and phrase transitions as the rollout extends further from the original context.

The first 32 generated bars are the same MWM continuation heard in the comparison above. Use “Bars 33–64” to jump directly to the extended portion; these bar numbers refer to generated music and exclude the prelude. This gallery presents the complete MWM rollout for listening beyond the 32-bar comparison.

Action-conditioned Music Editing

Compare five fixed song contexts under one four-stage drum edit. The opening bars stay identical across systems; the action removes the drum role, restores it with the same guide, and then repeats the change. Listen for timing, role recovery, and stability in the surrounding arrangement across Music World Model and five baselines.

Each column keeps one song context fixed. Rows are separate completed runs using the same four-stage edit schedule.

04 / EXTENDED ACTION CONTROL

Extended Action Control

Explore Music World Model’s response to density targets, MIDI guidance, and edits to melody, bass, chords, or drums. Select a context and compare each controlled continuation with its autonomous reference. Density controls the number of note onsets, MIDI guidance varies target following, and role edits alternate between muting and restoration with a supplied MIDI guide. Listen for both the requested change and changes in the accompanying parts.

05 / ACTION SEQUENCES

Action Sequences

Follow a changing instruction across four eight-bar phrases, with each generated segment feeding the next step. Choose alternating or gradually increasing density, changing MIDI guidance, octave targets, or instrument removal and restoration. Compare the result with the autonomous continuation from the same starting context, listening for responses at action boundaries and continuity between phrases. Select a phrase in the schedule to jump to that point.

06 / STRUCTURAL DRIFT & CORRECTION

BHC Analysis

Compare predicted-plan rollout (PP), boundary-aware correction (BHC), and the teacher-plan diagnostic (TT) across the 32-bar and 64-bar horizons reported in the paper. BHC anchors the slow trajectory and carries a decaying correction across later steps, while TT provides the diagnostic reference for measuring the planning gap.

PP / BHC / TT
PP

Predicted planRecursive generation under the model’s own structural plan.

TT

Teacher-plan referenceUse target-derived slow structure for diagnosis while retaining recursive execution.

Paper Figure 3: decoded F1, diagnostic F1 gaps, and latent and logit RMS distances across rollout horizons.
Long-horizon results. BHC recovers 5.5% and 29.1% of the PP–TT terminal-F1 gap at 32 and 64 bars, respectively. At 64 bars, the absolute F1 gain over PP is 0.0080. Distances are averaged across rollout steps and equally across songs.
Paper Figure 2: PP–TT gaps by segment position and song, followed by recovery under persistent plan mismatch.
Plan-mismatch robustness. Paired persistent Gaussian perturbations use autoregressive coefficient 0.9. As innovation strength increases from 0 to 1, BHC recovery rises from 0.0080 to 0.0331, recovering 29.1%–42.6% of the diagnostic gap.
07 / MECHANISM ABLATIONS

Component Ablations

Full BHC and five controls compare anchoring, propagation, smoothing, and correction direction.

06 correction settings

All settings share the frozen checkpoint, prefixes, action schedules, and decoder. Six deterministic inference configurations are averaged within song. Confidence intervals use 10,000 paired song-bootstrap replicates.

64-bar continuation · 44 held-out songs
Correction settingF1 gain ↑95% CILatent reduction ↑Logit reduction ↑
Anchoring only+0.0068[+0.0005,+0.0132]0.1190.062
Propagation only+0.0045[+0.0007,+0.0093]0.0720.038
EMA slow-plan+0.0023[-0.0012,+0.0059]0.0510.024
Reverse propagation+0.0012[-0.0026,+0.0050]0.0350.019
Norm-matched random-0.0002[-0.0013,+0.0011]0.0120.005

F1 gain is measured at the terminal step relative to uncorrected MWM. Latent and logit reductions are decreases in rollout-averaged RMS distance to TT. At 64 bars, full BHC has the largest mean F1 gain and distance reductions.

Listening Guide & Recording Setup

What to listen for. Compare continuity at phrase boundaries, requested changes, role recovery, and unintended changes in the surrounding parts.

How the recordings are organized. The generation and control galleries reuse five fixed contexts. The 32-bar section compares completed continuations, the 64-bar section extends the same MWM contexts, and the BHC section separates structural-path listening from the reported measurements.

How actions work. Density changes the onset budget, MIDI guidance changes target following, and role actions mute or restore melody, bass, chords, or drums. Action Sequences apply one instruction every eight bars, with each generated segment feeding the next.

Playback. Only one recording plays at a time. Use the native controls to switch contexts or compare rows; long-form players include a shortcut to the second half.