Self-play vs. text pretraining:
what kind of in-context learning
does each produce?

A follow-up to Self-Play Pretraining with Zero Data

Abstract

Does pretraining on self-generated programs teach the same in-context learning as pretraining on web text? We compare the same byte-level transformer architecture across six model sizes, matching learner tokens. They learn different kinds of ICL. Self-play leads on procedural tasks; text leads on lookup and word meaning. Starting text training from self-play also brings word-level ICL earlier.

Models are at most 24M parameters. Text runs have 1–2 seeds. Learner tokens are matched, rather than total compute: self-play also trains a generator.

1. Different data, different strengths

Explore the measurements ↓

(a) Procedural tasks

Printable-text format

(b) Lookup & word meaning

Word format

(c) The paper’s byte format

Raw-byte format

Figure 1 values & download

2. Explore the measurements

Every point is a released result

Figure 2 values & download

3. What the comparison measures

Same architecture. Three prompt formats.

The learner is a byte-level Llama-style transformer with a 4,096-byte context. Each task gives worked examples, then a query. Evaluation uses greedy decoding with no fine-tuning.

Printable-text and word answers require every byte to be correct. The raw reverse task uses per-byte accuracy, following the original paper. Suite scores average fixed task cells at m = 32–128; the word headline excludes word-order reversal. The explorer still includes the zero-scoring tasks.

Illustrative word prompt · arbitrary labels
cat=t
mouth=b
rabbit=

Query word unseen in the examples; predict t.

Read the uncertainty with the result.

Self-play has four released seeds. Bands and whiskers are 95% Student’s t intervals over seeds, clipped to [0, 1] for display. Text has two seeds at 1M, 3M and 6M; one at the other sizes. Text runs appear individually, without a seed interval. Exact, unclipped bounds are in the tables.

Warm starts use one seed and begin at self-play round 8191 (51.54B prior learner tokens). Additional text tokens exclude that cost. The universal prior uses different program counts per round and has no released 24M checkpoint.

These are small-model observations across a narrow task set. The size trends are not established scaling laws, and the token-matched comparison does not establish a general advantage per unit of total compute.

All limitations in the writeup →

4. Sources & downloadable data

This companion uses the five published CSV tables and original figures from the released repository. Values retain the source precision. Browse full evaluation JSONs and training logs for trial counts and experiment metadata.

Original publication figures · SVG & PDF

Built on Self-Play Pretraining with Zero Data by Cowsik et al., with released code and learners. Text data: DCLM-Baseline 1.0.