Extended clm on harvey LAB benchmark
This CLM projection head was trained on successful Parthenon and CAFL agent traces. It requires the frozen Qwen3-8B encoder and the CLM code.
Data and training
Parthenon supplied 408 traces from Claude Code (Sonnet 4.6 and Haiku 4.5) and Codex (GPT-5.5 and GPT-5.4 mini). CAFL supplied 290 traces from DeepSeek, Gemini, and GLM configurations. The combined pool contains 310 tasks and 32,801 causal state/next-action pairs. The trainer uses 279 tasks and 29,455 pairs for updates. Its internal validation split contains 31 tasks and 3,346 pairs.
Training starts from the released CLM-v0.1-8B head.
Both heads preserve the released architecture: 4,096 input dimensions,
width 1,536, depth 3, 512 output dimensions, GELU, layer normalization,
and no residual connections. Qwen3-8B stays frozen at revision
b968826d9c46dd6066d109eabc6255188de91218.
States use the chat template and final 8,191 tokens.
Actions use the first 8,191 tokens without special tokens.
Embeddings use normalized last-token pooling.
The run uses masked bidirectional InfoNCE, AdamW, batches of 2,048, zero weight decay, and a 20-epoch OneCycle cosine schedule. The warm-start learning rate is 0.0011547005383792516. The trainer caps logit scale after each optimizer step. Seed is 1234; early-stopping patience is five epochs. Internal within-task retrieval selects epoch 20 at 17.30% accuracy.
Evaluation
The unchanged upstream evaluator selects among up to eight traces per task. It uses mean cosine similarity over the final 12 steps.
| Head | Validation: 52 tasks | Test: 54 tasks |
|---|---|---|
| Released starting head | 14/52 (26.92%) | 15/54 (27.78%) |
| Extended clm on harvey LAB benchmark | 18/52 (34.62%) | 19/54 (35.19%) |
The lower learning rate tied the temperature-only change on validation. The test cohort was inspected earlier, so its result is exploratory. Both outer cohorts contain Parthenon traces and require a successful candidate. Their task families are excluded from the training pool. The internal split separates task IDs but shares eight related task families.
Reproduce
Data and cached embeddings include the exact conversations, labels, row order, and task splits. Reproduction scripts provide download, verification, evaluation, training, and embedding commands. See PR #12.
head.pt contains both heads, logit scale, configuration, epoch, and metrics.
All tensors match the recorded training checkpoint exactly.
Two local filesystem paths were removed from checkpoint metadata.
checksums.json records SHA-256 hashes.
provenance.json records the starting checkpoint hash and training settings.
training_history.json contains all epoch metrics.
val_results.json and test_results.json contain evaluator results and selected trace IDs.