ICML 2026
MIMOMamba: From Scalar Duality to Matrix-Valued Attention
TL;DR. Mamba's state space duality (SSD) only works for single-input single-output systems, so it cannot mix information across state dimensions without a separate feed-forward layer. MIMOMamba restores this mixing by parameterizing the state matrices as polynomials of a shared base matrix, which guarantees the commutativity SSD needs while allowing full matrix-valued cross-channel dynamics, and the resulting matrix-valued dual attention form matches or exceeds Transformers on a sequence-modeling benchmark with about one-third the parameters.
Abstract
The state space duality (SSD) framework, central to modern state-space models (SSMs) such as Mamba, has established an efficient attention-like mechanism by leveraging the commutative property of linear recurrences. However, existing formulations are limited to single-input single-output (SISO) systems that enforce commutativity with a restrictive scalar-identity constraint, which prevents cross-dimensional interactions within the state dynamics. In this work, we generalize SSD to the MIMO setting by introducing a matrix polynomial parameterization. This approach not only provides a principled way to ensure commutativity for generalized duality but also induces a shared algebraic structure across state transitions, thereby significantly reducing parameter redundancy. Building on this foundation, we present MIMOMamba, a multi-head SSM architecture that captures rich cross-dimensional dynamics while retaining linear-time training. Empirical results on a sequence modeling benchmark show that MIMOMamba matches or exceeds the performance of standard Transformers with only approximately one-third the parameters of the baseline.
From Scalar to Matrix-Valued Duality
A selective state space model updates a latent state ht with ht = Atht−1 + Btxt, yt = Ctht. In its most general form this is a MIMO system, with At a full D×D matrix that can rotate and mix all D feature dimensions at once. The 1 framework behind Mamba-2 gives this recurrence a dual, parallel attention form, but only after collapsing At to a scalar-identity matrix atI — the constraint that lets the recurrence commute across time and be reordered into an attention-like product. The price is that each of the D channels then evolves independently, pushing all cross-channel mixing onto a separate feed-forward layer after the SSM block2. Linear-attention variants45678 take a related but different shortcut, leaning on the associative property of matrix multiplication rather than commutativity. MIMOMamba's question is whether the commutativity route can be walked without shrinking the state matrices down to scalars.

Figure 1. Softmax attention vs. linear attention. Left: softmax attention's quadratic complexity arises from computing the large L×L matrix QK⊤. Right: linear attention leverages associativity to first compute the smaller D×D matrix K⊤V. SSD instead leans on commutativity, the property MIMOMamba generalizes to the matrix-valued case.
Matrix Polynomial Parameterization
A direct extension of SSD to MIMO faces a dichotomy: drop commutativity and lose the duality, or keep the scalar-identity constraint and lose all cross-dimensional interaction. MIMOMamba only needs At, Bt, and Ct for a given head to commute with each other, and realizes this by drawing all three from a matrix polynomial of a single, shared, time-invariant base matrix A: each is a linear combination of I, A, A2, …, Ap with input-dependent coefficients. Matrix powers commute automatically (AiAj = AjAi), so this construction supplies commutativity for free while letting every state matrix act as a full D×D transformation.
This restriction costs nothing in practice. The Cayley–Hamilton theorem bounds the polynomial degree at D−1, keeping the parameterization compact, and the centralizer theorem shows the set of matrices commuting with A equals exactly the polynomials of A whenever A is non-derogatory — true almost surely under random initialization. MIMOMamba's polynomial family is therefore, with probability 1, the entire space of commutativity-compatible transformations rather than a proper subset of it.
The Matrix-Valued Dual Attention Form
Unrolling the MIMO recurrence expresses output yi as a sum over past inputs xj, each weighted by CiAiAi−1…Aj+1Bj. Because these matrices all commute, this product reorders into Ai:j+1·(CiBj), separating a state-dynamics term that depends only on the distance i−j from a projection term depending only on the absolute positions. Applied to every entry, this turns the MIMO recurrence into a block-wise Hadamard product Y = (L ○ (QK⊤)) ▹ X, where Q stacks the Ct matrices, K the Bt matrices, and L is a block-lower-triangular causal mask. Setting D = 1 recovers exactly the scalar SSD attention form in Mamba-2, so the SISO case is a special case rather than being replaced.
The key difference from ordinary attention is that every entry of this L×L grid is a D×D matrix block, not a scalar: each off-diagonal block CiAi:j+1Bj rotates and mixes a past input across feature dimensions rather than just decaying it by a scalar weight, and because the ordered matrix products Ai:j+1 enforce sequential order structurally, the mask needs no external positional encoding. The factorization also preserves the efficiency SSMs are built on: the global transformation stays block D-semiseparable, the same property Mamba-2 exploits11, and the polynomial parameterization turns the D×D matrix-product chain needed for parallel training into scalar polynomial multiplications evaluable by FFT-based convolution, compressing what must be communicated across GPUs from O(D2) down to O(p) coefficients per head, in the spirit of FlashAttention12 and Monarch matrices13.
Multi-Head MIMO Architecture
A separate constraint is the state expansion ratio N/D: larger expansion is known to matter for information-dense sequences9, but MIMOMamba's per-head state matches D for parameter efficiency, which would otherwise cap the ratio at N = D. Following multi-head attention310, MIMOMamba instead runs H parallel MIMO SSM heads, each with its own base matrix and polynomial coefficients, giving an effective joint dimension of H·D without inflating per-head computation. The H outputs are concatenated and combined by a linear merge layera, letting heads specialize in different frequency bands, timescales, or channel-interaction patterns.
Experiments
The benchmark generates sequences from a linear dynamical system whose state transition matrix is block-diagonal 2×2 rotation blocks, each rotating a pair of feature dimensions at its own frequency and decaying for stability, a direct test of whether a model can represent rotation-like mixing inside its own state recurrence rather than relegating it to a downstream feed-forward layer. MIMOMamba is compared against Gated DeltaNet14, Mamba-21, Mamba-315, a standard Transformer3, PD-SSM16, and an LSTM17, using RMSE (lower is better) and R² (higher is better).
| Model | RMSE (mean) | RMSE (max) | R² |
|---|---|---|---|
| MIMOMamba | 0.687 | 1.592 | 0.480 |
| Gated DeltaNet | 0.699 | 1.633 | 0.431 |
| Mamba-2 | 0.717 | 1.746 | 0.459 |
| Mamba-3 | 0.715 | 1.765 | 0.475 |
| Transformer | 0.749 | 1.966 | 0.465 |
| PD-SSM | 0.774 | 1.909 | 0.454 |
| LSTM | 0.974 | 2.888 | 0.438 |
Table 1. Sequence-prediction results on the rotational-dynamics benchmark, averaged over evaluation runs. MIMOMamba obtains the lowest mean and worst-case RMSE and the highest R² of all seven models, including the SISO counterpart Mamba-2 and the strongest non-MIMO competitor, Gated DeltaNet.
MIMOMamba improves over Mamba-2 on every metric (RMSE mean 0.687 vs. 0.717; R² 0.480 vs. 0.459), isolating the benefit of the MIMO formulation over the SISO duality it generalizes, and it beats every other baseline on every metric while using roughly a third of the Transformer's parameters at matched settingsb.
Ablation Study
Ablations vary the three knobs the polynomial parameterization introduces: polynomial degree p, number of heads H, and the rank of the low-rank base-matrix factorization.
| Degree | RMSE | Params |
|---|---|---|
| 2 | 0.6935 | 32,329 |
| 3 | 0.6951 | 33,097 |
| 4 | 0.6746 | 33,865 |
| 5 | 0.6843 | 34,633 |
| 6 | 0.6811 | 35,401 |
| 7 | 0.6839 | 36,169 |
| 8 | 0.6805 | 36,937 |
Table 2. Polynomial degree ablation. Accuracy peaks at degree 4 and does not improve monotonically beyond it.
| Heads | RMSE | Params |
|---|---|---|
| 1 | 0.6841 | 18,499 |
| 2 | 0.6890 | 23,621 |
| 4 | 0.6746 | 33,865 |
| 8 | 0.6770 | 54,353 |
| 16 | 0.6735 | 95,329 |
Table 3. Head-count ablation. Accuracy improves from 1 to 4 heads with diminishing returns per added parameter beyond 4, the paper's default.
| Rank | RMSE | Params |
|---|---|---|
| 1 | 0.6862 | 33,481 |
| 2 | 0.6843 | 33,609 |
| 4 | 0.6746 | 33,865 |
| 8 | 0.6918 | 34,377 |
| 16 | 0.6682 | 35,401 |
Table 4. Rank ablation on the low-rank base-matrix factorization. Rank 4 gives strong accuracy at low parameter cost.
Together these support degree 4, 4 heads, and rank 4 as the default: each knob shows a clear, non-monotonic optimum rather than accuracy simply tracking parameter count.
Open Question
A matrix polynomial restores the commutativity that scalar Mamba gets from an identity constraint, so channels can mix inside the state instead of staying decoupled. Deriving which cross-channel mixing patterns that algebra can express, and which still need attention, remains open.
Citation
@article{li2026mimomamba,
title={MIMOMamba: From Scalar Duality to Matrix-Valued Attention},
author={Li, Yanbo and Suwandi, Richard Cornelius and Yin, Feng and Sun, Yiyong and Huang, Wei and Pu, Wenqiang},
journal={43rd International Conference on Machine Learning (ICML)},
year={2026}
}
Footnotes
- An earlier draft of the architecture description also proposed routing this merge stage through a mixture-of-experts layer. The authors have clarified that this was left over from an early draft and was not implemented: all reported results use a plain linear merge layer with no expert routing. [↩]
- The Transformer baseline is run with PyTorch's built-in, highly optimized nn.Transformer implementation, while MIMOMamba is implemented in plain PyTorch without custom CUDA kernels. The authors note that the Transformer's faster wall-clock step time reflects this difference in kernel engineering maturity rather than a difference in algorithmic complexity, where MIMOMamba retains Mamba's linear scaling in sequence length. [↩]
References
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality [PDF]
Dao, T. and Gu, A., 2024. In International Conference on Machine Learning, pp. 10041–10071. PMLR. - Mamba: Linear-Time Sequence Modeling with Selective State Spaces [PDF]
Gu, A. and Dao, T., 2023. arXiv preprint arXiv:2312.00752. - Attention Is All You Need [PDF]
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I., 2017. Advances in Neural Information Processing Systems, 30. - Transformers Are RNNs: Fast Autoregressive Transformers with Linear Attention [PDF]
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F., 2020. In International Conference on Machine Learning, pp. 5156–5165. PMLR. - Rethinking Attention with Performers [PDF]
Choromanski, K.M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J.Q., Mohiuddin, A., Kaiser, L., Belanger, D.B., Colwell, L.J., and Weller, A., 2021. In International Conference on Learning Representations. - Retentive Network: A Successor to Transformer for Large Language Models [PDF]
Sun, Y., Dong, L., Huang, S., Ma, S., Xia, Y., Xue, J., Wang, J., and Wei, F., 2023. arXiv preprint arXiv:2307.08621. - Gated Linear Attention Transformers with Hardware-Efficient Training [PDF]
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y., 2024. In Proceedings of the 41st International Conference on Machine Learning, PMLR 235, pp. 56501–56523. - RWKV: Reinventing RNNs for the Transformer Era [PDF]
Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., et al., 2023. arXiv preprint arXiv:2305.13048. - HiPPO: Recurrent Memory with Optimal Polynomial Projections [PDF]
Gu, A., Dao, T., Ermon, S., Rudra, A., and Ré, C., 2020. Advances in Neural Information Processing Systems, 33, pp. 1474–1487. - Are Sixteen Heads Really Better Than One? [PDF]
Michel, P., Levy, O., and Neubig, G., 2019. Advances in Neural Information Processing Systems, 32. - Time and Space Efficient Generators for Quasiseparable Matrices
Pernet, C. and Storjohann, A., 2018. Journal of Symbolic Computation, 85, pp. 224–246. - FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness [PDF]
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C., 2022. Advances in Neural Information Processing Systems, 35, pp. 16344–16359. - Monarch: Expressive Structured Matrices for Efficient and Accurate Training [PDF]
Dao, T., Chen, B., Sohoni, N.S., Desai, A., Poli, M., Grogan, J., Liu, A., Rao, A., Rudra, A., and Ré, C., 2022. In International Conference on Machine Learning, pp. 4690–4721. PMLR. - Gated Delta Networks: Improving Mamba2 with Delta Rule [PDF]
Yang, S., Kautz, J., and Hatamizadeh, A., 2025. In International Conference on Learning Representations. - Mamba-3: Improved Sequence Modeling Using State Space Principles
Anonymous Authors, 2026. Under review, International Conference on Learning Representations. - Structured Sparse Transition Matrices to Enable State Tracking in State-Space Models [PDF]
Terzić, A., Menet, N., Hersche, M., Hofmann, T., and Rahimi, A., 2025. Advances in Neural Information Processing Systems. - Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J., 1997. Neural Computation, 9(8), pp. 1735–1780.