{"componentChunkName":"component---src-templates-publication-js","path":"/research/chimera-arxiv2026/","result":{"pageContext":{"publication":{"id":"chimera-arxiv2026","title":"Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers","authors":["Chongjian Ge","Hanwen Jiang","Tianyu Wang","Jiuxiang Gu","Yiran Xu","Ziwen Chen","Shaoteng Liu","Jing Shi","Yicong Hong","Zefan Cai","Hailin Jin","Hao Tan"],"highlightAuthor":"Tianyu Wang","venue":"arXiv","year":2026,"tldr":"A hybrid visual diffusion backbone that replaces most full attention with linear attention, paired with Chinchilla-style scaling laws that make its compute-optimal training recipe predictable.","links":{"paper":"https://arxiv.org/abs/2607.28611"},"insight":"Combines O(N) Kimi Delta Attention for long-context state tracking with interleaved Multi-head Latent Attention for global interaction, then introduces HeteroP to transfer hyperparameters across the heterogeneous architecture — yielding fitted compute-optimal scaling laws, 7.3x the pretraining compute efficiency of a full-attention baseline, and zero-shot extrapolation from 5-second clips to 30-second videos.","teaser":"/images/information/chimera.png","bibtex":"@article{ge2026chimera,\n  title={Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers},\n  author={Ge, Chongjian and Jiang, Hanwen and Wang, Tianyu and Gu, Jiuxiang and Xu, Yiran and Chen, Ziwen and Liu, Shaoteng and Shi, Jing and Hong, Yicong and Cai, Zefan and Jin, Hailin and Tan, Hao},\n  journal={arXiv preprint arXiv:2607.28611},\n  year={2026}\n}","role":"Core Contributor","authorMarkers":{"Chongjian Ge":["*","§","†"],"Hanwen Jiang":["*","§"],"Tianyu Wang":["*","§"],"Jiuxiang Gu":["§"],"Yiran Xu":["§"],"Ziwen Chen":["§"],"Hao Tan":["†"]},"markerLegend":{"*":"Equal contribution","§":"Core contribution","†":"Project lead"},"abstract":"Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor's functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.","abstractSource":"https://arxiv.org/abs/2607.28611","version":"arXiv version linked; publication venue as listed above","technicalContext":{"themes":["Kimi Delta Attention (KDA)","Multi-head Latent Attention (MLA)","hybrid attention","multimodal generation","HeteroP hyperparameter transfer","compute-optimal scaling"],"method":"Chimera adapts Kimi Delta Attention (KDA) to image and video diffusion. It combines three KDA layers per Multi-head Latent Attention (MLA) layer, using recurrent state tracking alongside bidirectional global attention. Modality-aware short convolutions supply local spatiotemporal structure, allowing a single raster-ordered text–image–video stream without explicit positional embeddings. The architecture builds on Kimi Linear and pairs hybrid attention with HeteroP hyperparameter transfer and compute-optimal scaling.","relevance":"The work addresses reusable architecture and optimization questions—efficient long-context modeling, multimodal token interfaces, hyperparameter transfer, and compute-optimal allocation—within visual diffusion.","scope":"The evaluated system is hybrid visual diffusion; periodic global-attention layers remain part of the model. Claims about compute efficiency depend on the specific architecture, baseline and evaluation protocol in the paper.","sources":[{"label":"Chimera: Sections 3.3–3.4 (hybrid attention and modality-aware convolutions), Section 4 (scaling recipe)","url":"https://arxiv.org/html/2607.28611v1"},{"label":"Kimi Linear: original Kimi Delta Attention architecture","url":"https://arxiv.org/html/2510.26692v2"}],"description":"Chimera adapts Kimi Delta Attention (KDA) and Multi-head Latent Attention to image–video diffusion, with HeteroP hyperparameter transfer and compute-optimal scaling."}}}},"staticQueryHashes":["63159454"]}