Method Overview
Our architecture treats music generation as a latent world model with two interacting time scales. A phrase-level planner updates at a lower frequency to determine the next section’s structural direction, while a segment-level executor runs at a higher frequency to turn that plan into local musical states and concrete multi-track events. The two modules exchange recursively updated latent states instead of repeatedly passing complete discrete note sequences, allowing the system to preserve long-range structure without directly modeling an unwieldy full-length sequence. User actions are first-class inputs to the same generation interface: density changes, MIDI guidance, and the addition, removal, or restoration of melody, chords, bass, and drums are handled through one action-conditioned executor rather than separate task-specific models. After each generated segment, the system writes the local state and phrase summary back into a structured history, which the planner uses for the next rollout step. This planning–execution–feedback loop gives one model a unified interface for long-form continuation, phrase-level control, and multi-track editing.
Generation Methods Comparison
Represent
Encode musical context into a compact latent state.
Plan
Predict phrase-level structure from the evolving history.
Generate
Roll out segment states under actions, then decode the music.
Model Comparison
Listen across five fixed musical contexts: the ground-truth continuation, Music World Model, and six baseline systems. The reference row shows the actual future, while each model generates a continuation from the observed music. Switch players within a column to hear how different generation methods extend the same starting point.
Every card begins with the same eight-bar observed prelude. Reference rows show the held-out continuation; model rows show autonomous generations from the observed context. Playback uses each context's fixed tempo, shared SoundFont, and common gain, with no per-clip normalization.
32-Bar Comparison
Compare 32-bar continuations from the same five musical contexts. Each recording begins with eight bars of observed music, followed by either the reference continuation or music generated by Music World Model, Amadeus-S, NotaGen-RL3, or Music2Music-PT. Listen for how each model develops motifs, maintains the arrangement, and connects successive phrases as generation moves beyond the starting context.
The reference row contains the actual next 32 bars. Model continuations are generated without access to that future, and each baseline’s later steps use its own generated history. Recordings share the playback tempo, SoundFont, and gain within each context. Switch players within a column to compare the same passage across systems.
Generation settings and recording details
Amadeus-S and NotaGen-RL3 generate four successive eight-bar segments, conditioning each step on the latest eight bars. Music2Music-PT generates sixteen two-bar segments, using the latest two bars to fit its token limit. MWM retains its original history of up to 48 observed bars.
The five source windows and inference settings are fixed. Decoded timing is preserved within the fixed listening window.
64-Bar MWM Showcase
Follow Music World Model through 64 generated bars from each of the same five starting contexts. After the eight-bar observed prelude, the model continues for eight consecutive phrases, carrying its generated musical history forward. Listen to the second half for recurring motifs, changes in texture, and phrase transitions as the rollout extends further from the original context.
The first 32 generated bars are the same MWM continuation heard in the comparison above. Use “Bars 33–64” to jump directly to the extended portion; these bar numbers refer to generated music and exclude the prelude. This gallery presents the complete MWM rollout for listening beyond the 32-bar comparison.
Action-conditioned Music Editing
Compare five fixed song contexts under one four-stage drum edit. The opening bars stay identical across systems; the action removes the drum role, restores it with the same guide, and then repeats the change. Listen for timing, role recovery, and stability in the surrounding arrangement across Music World Model and five baselines.
Each column keeps one song context fixed. Rows are separate completed runs using the same four-stage edit schedule.
Extended Action Control
Explore Music World Model’s response to density targets, MIDI guidance, and edits to melody, bass, chords, or drums. Select a context and compare each controlled continuation with its autonomous reference. Density controls the number of note onsets, MIDI guidance varies target following, and role edits alternate between muting and restoration with a supplied MIDI guide. Listen for both the requested change and changes in the accompanying parts.
Action Sequences
Follow a changing instruction across four eight-bar phrases, with each generated segment feeding the next step. Choose alternating or gradually increasing density, changing MIDI guidance, octave targets, or instrument removal and restoration. Compare the result with the autonomous continuation from the same starting context, listening for responses at action boundaries and continuity between phrases. Select a phrase in the schedule to jump to that point.
Shape the musical density
Click a phrase to listen. Switch players to compare the same moment.
BHC Analysis
Compare predicted-plan rollout (PP), boundary-aware correction (BHC), and the teacher-plan diagnostic (TT) across the 32-bar and 64-bar horizons reported in the paper. BHC anchors the slow trajectory and carries a decaying correction across later steps, while TT provides the diagnostic reference for measuring the planning gap.
Predicted planRecursive generation under the model’s own structural plan.
Boundary-aware correctionAnchor the slow trajectory and carry a decaying boundary correction forward.
Teacher-plan referenceUse target-derived slow structure for diagnosis while retaining recursive execution.
Component Ablations
Full BHC and five controls compare anchoring, propagation, smoothing, and correction direction.
All settings share the frozen checkpoint, prefixes, action schedules, and decoder. Six deterministic inference configurations are averaged within song. Confidence intervals use 10,000 paired song-bootstrap replicates.
| Correction setting | F1 gain ↑ | 95% CI | Latent reduction ↑ | Logit reduction ↑ |
|---|---|---|---|---|
| Full BHC | +0.0031 | [-0.0021,+0.0085] | 0.058 | 0.024 |
| Anchoring only | +0.0026 | [-0.0024,+0.0072] | 0.061 | 0.027 |
| Propagation only | +0.0010 | [-0.0030,+0.0046] | 0.032 | 0.013 |
| EMA slow-plan | +0.0012 | [-0.0028,+0.0051] | 0.036 | 0.015 |
| Reverse propagation | +0.0005 | [-0.0031,+0.0034] | 0.025 | 0.011 |
| Norm-matched random | -0.0009 | [-0.0045,+0.0014] | -0.006 | -0.002 |
| Correction setting | F1 gain ↑ | 95% CI | Latent reduction ↑ | Logit reduction ↑ |
|---|---|---|---|---|
| Full BHC | +0.0080 | [+0.0030,+0.0178] | 0.151 | 0.081 |
| Anchoring only | +0.0068 | [+0.0005,+0.0132] | 0.119 | 0.062 |
| Propagation only | +0.0045 | [+0.0007,+0.0093] | 0.072 | 0.038 |
| EMA slow-plan | +0.0023 | [-0.0012,+0.0059] | 0.051 | 0.024 |
| Reverse propagation | +0.0012 | [-0.0026,+0.0050] | 0.035 | 0.019 |
| Norm-matched random | -0.0002 | [-0.0013,+0.0011] | 0.012 | 0.005 |
F1 gain is measured at the terminal step relative to uncorrected MWM. Latent and logit reductions are decreases in rollout-averaged RMS distance to TT. At 64 bars, full BHC has the largest mean F1 gain and distance reductions.