-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathmtp_sweep.log
More file actions
394 lines (377 loc) · 20.8 KB
/
Copy pathmtp_sweep.log
File metadata and controls
394 lines (377 loc) · 20.8 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
=== phase_e mtp_scale=0.0 seed=42 ===
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
======================================================================
Phase E: mtp_scale=0.0 (w2=0.0, w3=0.0) seed=42
500M tokens | 30,517 steps | batch=16x1024
======================================================================
Loading data...
Downloading FineWeb-Edu and tokenizing (500M tokens)...
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
step 1/30517 next-tok CE 10.1844 (0.0 min)
step 500/30517 next-tok CE 5.6844 (0.7 min)
step 1000/30517 next-tok CE 5.1734 (1.5 min)
[step-1000 snapshot] /root/perrow-gradient-interference/results/checkpoints/model_scale0.0_seed42_step1000.pt
step 1500/30517 next-tok CE 4.8531 (2.2 min)
step 2000/30517 next-tok CE 4.6828 (2.9 min)
step 2500/30517 next-tok CE 4.5453 (3.6 min)
step 3000/30517 next-tok CE 4.4938 (4.3 min)
step 3500/30517 next-tok CE 4.3875 (5.1 min)
step 4000/30517 next-tok CE 4.3922 (5.8 min)
step 4500/30517 next-tok CE 4.3141 (6.5 min)
step 5000/30517 next-tok CE 4.3078 (7.2 min)
step 5500/30517 next-tok CE 4.2422 (7.9 min)
step 6000/30517 next-tok CE 4.2109 (8.7 min)
step 6500/30517 next-tok CE 4.1992 (9.4 min)
step 7000/30517 next-tok CE 4.2172 (10.1 min)
step 7500/30517 next-tok CE 4.1273 (10.8 min)
step 8000/30517 next-tok CE 4.1242 (11.5 min)
step 8500/30517 next-tok CE 4.1141 (12.3 min)
step 9000/30517 next-tok CE 4.1266 (13.0 min)
step 9500/30517 next-tok CE 4.1070 (13.7 min)
step 10000/30517 next-tok CE 4.1367 (14.4 min)
step 10500/30517 next-tok CE 4.1016 (15.1 min)
step 11000/30517 next-tok CE 4.0977 (15.9 min)
step 11500/30517 next-tok CE 4.0875 (16.6 min)
step 12000/30517 next-tok CE 4.0727 (17.3 min)
step 12500/30517 next-tok CE 4.0781 (18.0 min)
step 13000/30517 next-tok CE 4.0719 (18.7 min)
step 13500/30517 next-tok CE 4.0664 (19.5 min)
step 14000/30517 next-tok CE 4.0867 (20.2 min)
step 14500/30517 next-tok CE 4.0445 (20.9 min)
step 15000/30517 next-tok CE 4.0336 (21.6 min)
step 15500/30517 next-tok CE 4.0430 (22.3 min)
step 16000/30517 next-tok CE 4.0891 (23.0 min)
step 16500/30517 next-tok CE 4.0117 (23.8 min)
step 17000/30517 next-tok CE 4.0281 (24.5 min)
step 17500/30517 next-tok CE 4.0469 (25.2 min)
step 18000/30517 next-tok CE 4.0469 (25.9 min)
step 18500/30517 next-tok CE 3.9883 (26.6 min)
step 19000/30517 next-tok CE 4.0172 (27.4 min)
step 19500/30517 next-tok CE 4.0242 (28.1 min)
step 20000/30517 next-tok CE 4.0187 (28.8 min)
step 20500/30517 next-tok CE 3.9852 (29.5 min)
step 21000/30517 next-tok CE 4.0320 (30.2 min)
step 21500/30517 next-tok CE 3.9867 (31.0 min)
step 22000/30517 next-tok CE 4.0094 (31.7 min)
step 22500/30517 next-tok CE 4.0031 (32.4 min)
step 23000/30517 next-tok CE 3.9773 (33.1 min)
step 23500/30517 next-tok CE 4.0227 (33.8 min)
step 24000/30517 next-tok CE 4.0039 (34.6 min)
step 24500/30517 next-tok CE 3.9984 (35.3 min)
step 25000/30517 next-tok CE 4.0000 (36.0 min)
step 25500/30517 next-tok CE 4.0094 (36.7 min)
step 26000/30517 next-tok CE 4.0289 (37.4 min)
step 26500/30517 next-tok CE 3.9852 (38.2 min)
step 27000/30517 next-tok CE 3.9844 (38.9 min)
step 27500/30517 next-tok CE 3.9555 (39.6 min)
step 28000/30517 next-tok CE 4.0016 (40.3 min)
step 28500/30517 next-tok CE 4.0242 (41.0 min)
step 29000/30517 next-tok CE 3.9672 (41.7 min)
step 29500/30517 next-tok CE 4.0055 (42.5 min)
step 30000/30517 next-tok CE 3.9773 (43.2 min)
step 30500/30517 next-tok CE 3.9969 (43.9 min)
step 30517/30517 next-tok CE 3.9680 (43.9 min)
=== FINAL ENRICHED EVAL ===
next-token (t+1) CE : 3.9941
t+2 CE : 8.9750
t+3 CE : 9.5516
mean output entropy : 4.3713 nats
Saved checkpoint /root/perrow-gradient-interference/results/checkpoints/model_scale0.0_seed42.pt
Saved /root/perrow-gradient-interference/results/phase_e/sweep_scale0.0_seed42_30517steps.json
=== phase_e mtp_scale=1.0 seed=42 ===
======================================================================
Phase E: mtp_scale=1.0 (w2=0.5, w3=0.25) seed=42
500M tokens | 30,517 steps | batch=16x1024
======================================================================
Loading data...
Loading cached tokens from fineweb_train_500M.pt
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
step 1/30517 next-tok CE 10.1937 (0.0 min)
step 500/30517 next-tok CE 5.9453 (1.0 min)
step 1000/30517 next-tok CE 5.5234 (2.0 min)
[step-1000 snapshot] /root/perrow-gradient-interference/results/checkpoints/model_scale1.0_seed42_step1000.pt
step 1500/30517 next-tok CE 5.2094 (3.1 min)
step 2000/30517 next-tok CE 5.0359 (4.1 min)
step 2500/30517 next-tok CE 4.9047 (5.1 min)
step 3000/30517 next-tok CE 4.8594 (6.1 min)
step 3500/30517 next-tok CE 4.7531 (7.1 min)
step 4000/30517 next-tok CE 4.7500 (8.1 min)
step 4500/30517 next-tok CE 4.6859 (9.1 min)
step 5000/30517 next-tok CE 4.6875 (10.1 min)
step 5500/30517 next-tok CE 4.6125 (11.2 min)
step 6000/30517 next-tok CE 4.5797 (12.2 min)
step 6500/30517 next-tok CE 4.5750 (13.2 min)
step 7000/30517 next-tok CE 4.5891 (14.2 min)
step 7500/30517 next-tok CE 4.5172 (15.2 min)
step 8000/30517 next-tok CE 4.5000 (16.2 min)
step 8500/30517 next-tok CE 4.4938 (17.2 min)
step 9000/30517 next-tok CE 4.5109 (18.2 min)
step 9500/30517 next-tok CE 4.4922 (19.3 min)
step 10000/30517 next-tok CE 4.5172 (20.3 min)
step 10500/30517 next-tok CE 4.4891 (21.3 min)
step 11000/30517 next-tok CE 4.4844 (22.3 min)
step 11500/30517 next-tok CE 4.4672 (23.3 min)
step 12000/30517 next-tok CE 4.4578 (24.3 min)
step 12500/30517 next-tok CE 4.4688 (25.3 min)
step 13000/30517 next-tok CE 4.4562 (26.3 min)
step 13500/30517 next-tok CE 4.4500 (27.3 min)
step 14000/30517 next-tok CE 4.4531 (28.4 min)
step 14500/30517 next-tok CE 4.4172 (29.4 min)
step 15000/30517 next-tok CE 4.4156 (30.4 min)
step 15500/30517 next-tok CE 4.4156 (31.4 min)
step 16000/30517 next-tok CE 4.4266 (32.4 min)
step 16500/30517 next-tok CE 4.3844 (33.4 min)
step 17000/30517 next-tok CE 4.3906 (34.4 min)
step 17500/30517 next-tok CE 4.4234 (35.4 min)
step 18000/30517 next-tok CE 4.4281 (36.5 min)
step 18500/30517 next-tok CE 4.3703 (37.5 min)
step 19000/30517 next-tok CE 4.4000 (38.5 min)
step 19500/30517 next-tok CE 4.3969 (39.5 min)
step 20000/30517 next-tok CE 4.3969 (40.5 min)
step 20500/30517 next-tok CE 4.3703 (41.5 min)
step 21000/30517 next-tok CE 4.4172 (42.5 min)
step 21500/30517 next-tok CE 4.3641 (43.5 min)
step 22000/30517 next-tok CE 4.3797 (44.5 min)
step 22500/30517 next-tok CE 4.3719 (45.6 min)
step 23000/30517 next-tok CE 4.3453 (46.6 min)
step 23500/30517 next-tok CE 4.4016 (47.6 min)
step 24000/30517 next-tok CE 4.3844 (48.6 min)
step 24500/30517 next-tok CE 4.3719 (49.6 min)
step 25000/30517 next-tok CE 4.3781 (50.6 min)
step 25500/30517 next-tok CE 4.3875 (51.6 min)
step 26000/30517 next-tok CE 4.4125 (52.6 min)
step 26500/30517 next-tok CE 4.3688 (53.7 min)
step 27000/30517 next-tok CE 4.3656 (54.7 min)
step 27500/30517 next-tok CE 4.3391 (55.7 min)
step 28000/30517 next-tok CE 4.3781 (56.7 min)
step 28500/30517 next-tok CE 4.3984 (57.7 min)
step 29000/30517 next-tok CE 4.3547 (58.7 min)
step 29500/30517 next-tok CE 4.3812 (59.7 min)
step 30000/30517 next-tok CE 4.3578 (60.7 min)
step 30500/30517 next-tok CE 4.3703 (61.7 min)
step 30517/30517 next-tok CE 4.3438 (61.8 min)
=== FINAL ENRICHED EVAL ===
next-token (t+1) CE : 4.3656
t+2 CE : 6.0516
t+3 CE : 6.9680
mean output entropy : 5.6609 nats
Saved checkpoint /root/perrow-gradient-interference/results/checkpoints/model_scale1.0_seed42.pt
Saved /root/perrow-gradient-interference/results/phase_e/sweep_scale1.0_seed42_30517steps.json
=== phase_e mtp_scale=0.25 seed=42 ===
======================================================================
Phase E: mtp_scale=0.25 (w2=0.125, w3=0.0625) seed=42
500M tokens | 30,517 steps | batch=16x1024
======================================================================
Loading data...
Loading cached tokens from fineweb_train_500M.pt
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
step 1/30517 next-tok CE 10.1844 (0.0 min)
step 500/30517 next-tok CE 5.7188 (1.0 min)
step 1000/30517 next-tok CE 5.2406 (2.0 min)
[step-1000 snapshot] /root/perrow-gradient-interference/results/checkpoints/model_scale0.25_seed42_step1000.pt
step 1500/30517 next-tok CE 4.9156 (3.1 min)
step 2000/30517 next-tok CE 4.7406 (4.1 min)
step 2500/30517 next-tok CE 4.6141 (5.1 min)
step 3000/30517 next-tok CE 4.5641 (6.1 min)
step 3500/30517 next-tok CE 4.4578 (7.1 min)
step 4000/30517 next-tok CE 4.4578 (8.1 min)
step 4500/30517 next-tok CE 4.3891 (9.1 min)
step 5000/30517 next-tok CE 4.3969 (10.1 min)
step 5500/30517 next-tok CE 4.3234 (11.1 min)
step 6000/30517 next-tok CE 4.2875 (12.2 min)
step 6500/30517 next-tok CE 4.2781 (13.2 min)
step 7000/30517 next-tok CE 4.3016 (14.2 min)
step 7500/30517 next-tok CE 4.2203 (15.2 min)
step 8000/30517 next-tok CE 4.2094 (16.2 min)
step 8500/30517 next-tok CE 4.2047 (17.2 min)
step 9000/30517 next-tok CE 4.2156 (18.2 min)
step 9500/30517 next-tok CE 4.2016 (19.2 min)
step 10000/30517 next-tok CE 4.2391 (20.2 min)
step 10500/30517 next-tok CE 4.2172 (21.2 min)
step 11000/30517 next-tok CE 4.2109 (22.3 min)
step 11500/30517 next-tok CE 4.2000 (23.3 min)
step 12000/30517 next-tok CE 4.1953 (24.3 min)
step 12500/30517 next-tok CE 4.2000 (25.3 min)
step 13000/30517 next-tok CE 4.1977 (26.3 min)
step 13500/30517 next-tok CE 4.1937 (27.3 min)
step 14000/30517 next-tok CE 4.1883 (28.3 min)
step 14500/30517 next-tok CE 4.1461 (29.3 min)
step 15000/30517 next-tok CE 4.1398 (30.3 min)
step 15500/30517 next-tok CE 4.1344 (31.3 min)
step 16000/30517 next-tok CE 4.1547 (32.4 min)
step 16500/30517 next-tok CE 4.0969 (33.4 min)
step 17000/30517 next-tok CE 4.1023 (34.4 min)
step 17500/30517 next-tok CE 4.1266 (35.4 min)
step 18000/30517 next-tok CE 4.1437 (36.4 min)
step 18500/30517 next-tok CE 4.0891 (37.4 min)
step 19000/30517 next-tok CE 4.1141 (38.4 min)
step 19500/30517 next-tok CE 4.1133 (39.4 min)
step 20000/30517 next-tok CE 4.1203 (40.4 min)
step 20500/30517 next-tok CE 4.0820 (41.5 min)
step 21000/30517 next-tok CE 4.1305 (42.5 min)
step 21500/30517 next-tok CE 4.0820 (43.5 min)
step 22000/30517 next-tok CE 4.1000 (44.5 min)
step 22500/30517 next-tok CE 4.0992 (45.5 min)
step 23000/30517 next-tok CE 4.0602 (46.5 min)
step 23500/30517 next-tok CE 4.1188 (47.5 min)
step 24000/30517 next-tok CE 4.0953 (48.5 min)
step 24500/30517 next-tok CE 4.0914 (49.5 min)
step 25000/30517 next-tok CE 4.0859 (50.5 min)
step 25500/30517 next-tok CE 4.1023 (51.6 min)
step 26000/30517 next-tok CE 4.1297 (52.6 min)
step 26500/30517 next-tok CE 4.0797 (53.6 min)
step 27000/30517 next-tok CE 4.0836 (54.6 min)
step 27500/30517 next-tok CE 4.0477 (55.6 min)
step 28000/30517 next-tok CE 4.0938 (56.6 min)
step 28500/30517 next-tok CE 4.1211 (57.6 min)
step 29000/30517 next-tok CE 4.0594 (58.6 min)
step 29500/30517 next-tok CE 4.0938 (59.6 min)
step 30000/30517 next-tok CE 4.0656 (60.6 min)
step 30500/30517 next-tok CE 4.0852 (61.7 min)
step 30517/30517 next-tok CE 4.0594 (61.7 min)
=== FINAL ENRICHED EVAL ===
next-token (t+1) CE : 4.0855
t+2 CE : 6.8680
t+3 CE : 7.7609
mean output entropy : 4.8811 nats
Saved checkpoint /root/perrow-gradient-interference/results/checkpoints/model_scale0.25_seed42.pt
Saved /root/perrow-gradient-interference/results/phase_e/sweep_scale0.25_seed42_30517steps.json
=== phase_e mtp_scale=0.5 seed=42 ===
======================================================================
Phase E: mtp_scale=0.5 (w2=0.25, w3=0.125) seed=42
500M tokens | 30,517 steps | batch=16x1024
======================================================================
Loading data...
Loading cached tokens from fineweb_train_500M.pt
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
step 1/30517 next-tok CE 10.1875 (0.0 min)
step 500/30517 next-tok CE 5.7938 (1.0 min)
step 1000/30517 next-tok CE 5.3344 (2.0 min)
[step-1000 snapshot] /root/perrow-gradient-interference/results/checkpoints/model_scale0.5_seed42_step1000.pt
step 1500/30517 next-tok CE 5.0016 (3.0 min)
step 2000/30517 next-tok CE 4.8344 (4.1 min)
step 2500/30517 next-tok CE 4.7031 (5.1 min)
step 3000/30517 next-tok CE 4.6562 (6.1 min)
step 3500/30517 next-tok CE 4.5469 (7.1 min)
step 4000/30517 next-tok CE 4.5500 (8.1 min)
step 4500/30517 next-tok CE 4.4813 (9.1 min)
step 5000/30517 next-tok CE 4.4844 (10.1 min)
step 5500/30517 next-tok CE 4.4141 (11.1 min)
step 6000/30517 next-tok CE 4.3844 (12.1 min)
step 6500/30517 next-tok CE 4.3750 (13.1 min)
step 7000/30517 next-tok CE 4.3922 (14.1 min)
step 7500/30517 next-tok CE 4.3156 (15.2 min)
step 8000/30517 next-tok CE 4.3078 (16.2 min)
step 8500/30517 next-tok CE 4.2969 (17.2 min)
step 9000/30517 next-tok CE 4.3016 (18.2 min)
step 9500/30517 next-tok CE 4.2984 (19.2 min)
step 10000/30517 next-tok CE 4.3297 (20.2 min)
step 10500/30517 next-tok CE 4.2984 (21.2 min)
step 11000/30517 next-tok CE 4.2906 (22.2 min)
step 11500/30517 next-tok CE 4.2781 (23.2 min)
step 12000/30517 next-tok CE 4.2641 (24.2 min)
step 12500/30517 next-tok CE 4.2797 (25.2 min)
step 13000/30517 next-tok CE 4.2859 (26.2 min)
step 13500/30517 next-tok CE 4.2703 (27.2 min)
step 14000/30517 next-tok CE 4.2734 (28.3 min)
step 14500/30517 next-tok CE 4.2344 (29.3 min)
step 15000/30517 next-tok CE 4.2234 (30.3 min)
step 15500/30517 next-tok CE 4.2266 (31.3 min)
step 16000/30517 next-tok CE 4.2438 (32.3 min)
step 16500/30517 next-tok CE 4.1813 (33.3 min)
step 17000/30517 next-tok CE 4.1922 (34.3 min)
step 17500/30517 next-tok CE 4.2141 (35.3 min)
step 18000/30517 next-tok CE 4.2188 (36.3 min)
step 18500/30517 next-tok CE 4.1719 (37.3 min)
step 19000/30517 next-tok CE 4.1953 (38.4 min)
step 19500/30517 next-tok CE 4.2031 (39.4 min)
step 20000/30517 next-tok CE 4.1953 (40.4 min)
step 20500/30517 next-tok CE 4.1656 (41.4 min)
step 21000/30517 next-tok CE 4.2125 (42.4 min)
step 21500/30517 next-tok CE 4.1703 (43.5 min)
step 22000/30517 next-tok CE 4.1859 (44.5 min)
step 22500/30517 next-tok CE 4.1789 (45.5 min)
step 23000/30517 next-tok CE 4.1492 (46.5 min)
step 23500/30517 next-tok CE 4.2078 (47.6 min)
step 24000/30517 next-tok CE 4.1813 (48.6 min)
step 24500/30517 next-tok CE 4.1766 (49.6 min)
step 25000/30517 next-tok CE 4.1742 (50.6 min)
step 25500/30517 next-tok CE 4.1891 (51.6 min)
step 26000/30517 next-tok CE 4.2156 (52.6 min)
step 26500/30517 next-tok CE 4.1703 (53.6 min)
step 27000/30517 next-tok CE 4.1734 (54.6 min)
step 27500/30517 next-tok CE 4.1398 (55.6 min)
step 28000/30517 next-tok CE 4.1750 (56.6 min)
step 28500/30517 next-tok CE 4.2047 (57.6 min)
step 29000/30517 next-tok CE 4.1469 (58.7 min)
step 29500/30517 next-tok CE 4.1742 (59.7 min)
step 30000/30517 next-tok CE 4.1562 (60.7 min)
step 30500/30517 next-tok CE 4.1797 (61.7 min)
step 30517/30517 next-tok CE 4.1484 (61.7 min)
=== FINAL ENRICHED EVAL ===
next-token (t+1) CE : 4.1773
t+2 CE : 6.4094
t+3 CE : 7.3258
mean output entropy : 5.1882 nats
Saved checkpoint /root/perrow-gradient-interference/results/checkpoints/model_scale0.5_seed42.pt
Saved /root/perrow-gradient-interference/results/phase_e/sweep_scale0.5_seed42_30517steps.json
=== T1: zero-training KL prediction from the CE-only checkpoint (the real hypothesis test) ===
Device: cuda
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
CE-only checkpoint: final_val_loss=3.994140625
k= 64 offset3=p3eqp2 half=0 coverage=0.751 KL(1)=0.3528 KL(0.5)=0.1904 KL(0.25)=0.0950
k= 64 offset3=p3eqp2 half=1 coverage=0.736 KL(1)=0.3400 KL(0.5)=0.1825 KL(0.25)=0.0904
k= 64 offset3=drop half=0 coverage=0.753 KL(1)=0.2485 KL(0.5)=0.1277 KL(0.25)=0.0614
k= 64 offset3=drop half=1 coverage=0.732 KL(1)=0.2391 KL(0.5)=0.1223 KL(0.25)=0.0584
k= 128 offset3=p3eqp2 half=0 coverage=0.797 KL(1)=0.3480 KL(0.5)=0.1877 KL(0.25)=0.0935
k= 128 offset3=p3eqp2 half=1 coverage=0.788 KL(1)=0.3391 KL(0.5)=0.1820 KL(0.25)=0.0901
k= 128 offset3=drop half=0 coverage=0.791 KL(1)=0.2419 KL(0.5)=0.1239 KL(0.25)=0.0593
k= 128 offset3=drop half=1 coverage=0.780 KL(1)=0.2362 KL(0.5)=0.1207 KL(0.25)=0.0576
k= 256 offset3=p3eqp2 half=0 coverage=0.823 KL(1)=0.3440 KL(0.5)=0.1853 KL(0.25)=0.0922
k= 256 offset3=p3eqp2 half=1 coverage=0.827 KL(1)=0.3408 KL(0.5)=0.1834 KL(0.25)=0.0911
k= 256 offset3=drop half=0 coverage=0.832 KL(1)=0.2432 KL(0.5)=0.1248 KL(0.25)=0.0599
k= 256 offset3=drop half=1 coverage=0.810 KL(1)=0.2309 KL(0.5)=0.1179 KL(0.25)=0.0562
T1: predicted KL(s=1) in [0.231, 0.353] (observed shared-MTP gap = 0.392) -> CONSISTENT
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
T2 entropy: CE-only 4.366 vs shared-MTP 5.635 (Δ+1.269; mixture predicts +0.3..0.4)
T3 future CE: shared-MTP better at t+2 by +2.950, t+3 by +2.581 (mixture predicts shared-MTP BETTER at t+2/t+3)
Saved /root/perrow-gradient-interference/analysis/kl_estimate_results.json
=== Ceilings + control (re-measures CE-vs-L1 control + CE-vs-CE ceiling on the step-1000 snapshots) ===
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
CONTROL CE-vs-L1 (active): median -0.029, %|cos|>0.3 1.0% (full-vocab 36.2%)
GPT2 initialized (124M params):
Total params: 124,354,560 (124.4M)
Muon params: 84,934,656 (84.9M) [68% Q/K/V/O + MLP per layer]
Adam params: 39,419,904 (39.4M) [32% wte/wpe/lm_head]
CE-vs-CE ceiling (active median): CE-only -0.236, MTP-trained -0.231 (CE-vs-MTP active median ~0.53 -> near ceiling)
1 batch(es): active 8.9%, median 0.541
2 batch(es): active 14.5%, median 0.539
Traceback (most recent call last):
File "/root/perrow-gradient-interference/analysis/measure_ceilings.py", line 132, in <module>
ns = measure_norm_support(mtp_model, acc)
File "/root/perrow-gradient-interference/measurement/measure_norm_support.py", line 52, in measure_norm_support
mtp3 = F.cross_entropy(logits[:, :-3].reshape(-1, V), batch[:, 3:].reshape(-1))
File "/usr/local/lib/python3.10/dist-packages/torch/nn/functional.py", line 3494, in cross_entropy
return torch._C._nn.cross_entropy_loss(
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 9.19 GiB. GPU 0 has a total capacity of 79.18 GiB of which 8.30 GiB is free. Including non-PyTorch memory, this process has 70.87 GiB memory in use. Of the allocated memory 65.25 GiB is allocated by PyTorch, and 4.89 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)