Repository navigation
Expand file tree
/
Copy pathindex.html
More file actions
212 lines (190 loc) · 11.3 KB
/
Copy pathindex.html
File metadata and controls
212 lines (190 loc) · 11.3 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="description" content="UniAR: Unified Multimodal Autoregressive Modeling with Shared Context --- Visual Tokenizer is Key to Unification">
<meta name="keywords" content="UniAR, unified multimodal model, autoregressive modeling, visual tokenizer, image generation, image editing, multimodal understanding">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>UniAR</title>
<link rel="icon" href="data:,">
<!-- Google tag (gtag.js) -->
<script async src="https://www.googletagmanager.com/gtag/js?id=G-562412C65F"></script>
<script>
window.dataLayer = window.dataLayer || [];
function gtag(){dataLayer.push(arguments);}
gtag('js', new Date());
gtag('config', 'G-562412C65F');
</script>
<link rel="stylesheet" href="https://fonts.googleapis.com/css?family=Google+Sans|Noto+Sans|Castoro">
<link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/bulma@0.9.1/css/bulma.min.css">
<link rel="stylesheet" href="https://cdn.jsdelivr.net/gh/jpswalsh/academicons@1/css/academicons.min.css">
<link rel="stylesheet" href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/5.15.1/css/all.min.css">
<link rel="stylesheet" href="./static/css/index.css">
</head>
<body>
<!-- ===== Logo Bar ===== -->
<div class="logo-bar">
<div class="container is-max-desktop">
<img src="./figs/qwen-logo.png" alt="Qwen logo" class="logo-left">
<img src="./figs/TEAI_logo.png" alt="TEAI logo" class="logo-right">
</div>
</div>
<!-- ===== Hero ===== -->
<section class="hero">
<div class="hero-body">
<div class="container is-max-desktop">
<div class="columns is-centered">
<div class="column has-text-centered">
<h1 class="title is-2 publication-title">
UniAR: Unified Multimodal Autoregressive Modeling with Shared Context
</h1>
<p class="publication-subtitle">Visual Tokenizer is Key to Unification</p>
<p class="publication-venue">ICML 2026</p>
<br>
<div class="is-size-5 publication-authors authors-block">
<span class="author-block"><a href="https://wjpoom.github.io/" target="_blank">Wujian Peng</a><sup>1,2,*,#</sup> </span>
<span class="author-block"><a href="https://scholar.google.com/citations?user=pUl9H8gAAAAJ&hl=en" target="_blank">Lingchen Meng</a><sup>3,*,‡</sup> </span>
<span class="author-block">Yuxuan Cai<sup>3</sup> </span>
<span class="author-block">Xianwei Zhuang<sup>3</sup> </span>
<span class="author-block">Yuhuan Yang<sup>3</sup> </span>
<span class="author-block">Rongyao Fang<sup>3</sup> </span>
<span class="author-block">Chenfei Wu<sup>3</sup> </span>
<span class="author-block">Junyang Lin<sup>3</sup> </span>
<span class="author-block"><a href="https://zxwu.azurewebsites.net/" target="_blank">Zuxuan Wu</a><sup>1,2,†</sup> </span>
<span class="author-block"><a href="https://scholar.google.com/citations?user=ylhI1JsAAAAJ&hl=zh-CN" target="_blank">Shuai Bai</a><sup>3,†</sup></span>
</div>
<div class="is-size-6 publication-authors corresponding-note">
<span><sup>*</sup> Equal contributions <sup>#</sup> Work done during internship at Qwen <sup>†</sup> Corresponding authors <sup>‡</sup> Project lead</span>
</div>
<div class="is-size-6 publication-authors publication-affiliations">
<span class="author-block"><sup>1</sup> Institute of Trustworthy Embodied AI, Fudan University</span>
<span class="author-block"><sup>2</sup> Shanghai Innovation Institute</span>
</div>
<div class="is-size-6 publication-authors publication-affiliations">
<span class="author-block"><sup>3</sup> Qwen Team, Alibaba Group</span>
</div>
<div class="publication-links">
<span class="link-block">
<a href="https://arxiv.org/pdf/2606.18249" target="_blank">
<span class="icon"><i class="ai ai-arxiv"></i></span>
<span>Arxiv</span>
</a>
</span>
<span class="link-block">
<a href="https://github.com/ShareLab-SII/UniAR" target="_blank">
<span class="icon"><i class="fab fa-github"></i></span>
<span>GitHub</span>
</a>
</span>
<span class="link-block">
<a href="https://huggingface.co/collections/ShareLab-SII/uniar" target="_blank">
<span class="icon"><i class="fas fa-cube"></i></span>
<span>Checkpoints</span>
</a>
</span>
</div>
<br>
</div>
</div>
</div>
</div>
<!-- Teaser + Abstract -->
<div class="container is-max-desktop abstract-block">
<div class="has-text-centered">
<img id="teaser" width="88%" src="./figs/teaser.jpg" alt="UniAR teaser">
<div class="content has-text-justified abstract-text">
<p>
<strong>UniAR</strong> is a unified autoregressive multimodal model that handles image understanding, image generation,
and image editing in a single Transformer. Unlike prior unified models that rely on two separate visual tokenizers
(splitting the representation space), UniAR uses a <strong>single discrete visual tokenizer</strong> as the key bridge
between understanding and generation, enabling a shared context in which the model can directly interpret its own
generated visual tokens without additional re-encoding.
</p>
<p><strong>Key design choices:</strong></p>
<ul>
<li><strong>Multi-level BSQ tokenizer</strong> — fuses shallow (low-level detail) and deep (high-level semantic) visual features via lookup-free Binary Spherical Quantization, scaling the effective vocabulary to 2<sup>64</sup> codes with minimal overhead.</li>
<li><strong>Parallel bitwise prediction</strong> — jointly predicts spatially grouped, multi-level visual codes per AR step, achieving a 32x visual compression ratio (a 1024×1024 image needs only 256 AR tokens).</li>
<li><strong>DiT-based visual decoder</strong> — an SD3-medium transformer with semantic visual feature injection that reconstructs high-fidelity images from discrete visual tokens, with resolution upsampling support.</li>
</ul>
</div>
</div>
</div>
</section>
<!-- ===== Highlights ===== -->
<section class="section section-compact" id="highlights">
<div class="container is-max-desktop">
<h2 class="title is-3 has-text-centered">Highlights</h2>
<span class="section-divider"></span>
<div class="highlight-box content">
<ul>
<li><strong>True unification via shared context.</strong> A single visual tokenizer is used for both understanding and generation, allowing UniAR to interpret its own generated visual tokens directly.</li>
<li><strong>Bitwise visual tokenization at scale.</strong> Lookup-free BSQ quantization expands the effective visual vocabulary with low overhead while preserving semantic alignment.</li>
<li><strong>Multi-level visual features.</strong> Hierarchical feature fusion retains both high-level semantics and fine-grained details, which is especially important for text rendering and editing.</li>
<li><strong>Fast autoregressive generation.</strong> Parallel bitwise prediction and a diffusion-based visual decoder with resolution upsampling reduce sequence length and accelerate image synthesis.</li>
<li><strong>Strong multimodal performance.</strong> UniAR achieves state-of-the-art or highly competitive results on image generation, image editing, OCR-heavy understanding, and long-text rendering.</li>
</ul>
</div>
</div>
</section>
<!-- ===== Method Overview ===== -->
<section class="section section-compact" id="method">
<div class="container is-max-desktop has-text-centered">
<h2 class="title is-3">Method Overview</h2>
<span class="section-divider"></span>
<img class="figure-image" src="./figs/arch.png" alt="UniAR architecture">
<div class="figure-caption">
UniAR consists of three core components: a unified visual tokenizer that discretizes semantic visual features into shared bitwise tokens,
a unified autoregressive backbone that jointly models text and visual tokens, and a diffusion-based visual decoder that reconstructs
high-fidelity images from predicted visual tokens.
</div>
</div>
</section>
<!-- ===== Interleaved Generation & Understanding ===== -->
<section class="section section-compact" id="shared-context">
<div class="container is-max-desktop has-text-centered">
<h2 class="title is-3">Interleaved Generation and Understanding</h2>
<span class="section-divider"></span>
<img class="figure-image" src="./figs/decoder.png" alt="UniAR interleaved generation-understanding example">
<div class="figure-caption">
UniAR unifies generation and understanding in the same discrete visual space. This enables an emergent interleaved capability:
after generating an image, the model can answer follow-up questions about its own output without re-encoding the image.
</div>
</div>
</section>
<!-- ===== BibTeX ===== -->
<section class="section bibtex-section" id="BibTeX">
<div class="container is-max-desktop content">
<h2 class="title has-text-centered">BibTeX</h2>
<span class="section-divider"></span>
<div class="bibtex-block">
<button class="bibtex-copy-btn" onclick="copyBibtex()" id="copyBtn">Copy</button>
<pre><code id="bibtexCode">@article{peng2026unified,
title={Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification},
author={Peng, Wujian and Meng, Lingchen and Cai, Yuxuan and Zhuang, Xianwei and Yang, Yuhuan and Fang, Rongyao and Wu, Chenfei and Lin, Junyang and Wu, Zuxuan and Bai, Shuai},
journal={arXiv preprint arXiv:2606.18249},
year={2026}
}</code></pre>
</div>
</div>
</section>
<!-- ===== Footer ===== -->
<footer class="site-footer">
<div class="container has-text-centered">
<p>
This website is adapted from <a href="https://nerfies.github.io/">Nerfies</a> and <a href="https://mathvista.github.io/">MathVista</a>,
licensed under a <a rel="license" href="http://creativecommons.org/licenses/by-sa/4.0/">Creative Commons Attribution-ShareAlike 4.0 International License</a>.
</p>
</div>
</footer>
<script>
function copyBibtex() {
var code = document.getElementById('bibtexCode').textContent;
navigator.clipboard.writeText(code).then(function() {
var btn = document.getElementById('copyBtn');
btn.textContent = 'Copied!';
setTimeout(function() { btn.textContent = 'Copy'; }, 2000);
});
}
</script>
</body>
</html>