Julian-600M Base
600M parameter language model trained from scratch on 39B tokens (70% EN, 30% FR) on Google Cloud TPUs with JAX.
This is the base model (text completion, not instruction-following). Start typing and the model will continue your text.
50 500
0.1 1.5
10 100
0.1 1
1 2
Try these examples
| Start typing (the model will continue) | Max Tokens | Temperature | Top-K | Top-P | Repetition Penalty |
|---|
Benchmark Results (zero-shot)
| Benchmark | Score |
|---|---|
| HellaSwag | 53.5% (acc_norm) |
| PIQA | 66.8% (acc) |
| LAMBADA | 37.3% (acc) |
| ARC-Easy | 54.0% (acc) |
| ARC-Challenge | 27.6% (acc_norm) |
| WinoGrande | 52.7% (acc) |
| BoolQ | 57.9% (acc) |
Outperforms OPT-1.3B (41.5% HellaSwag) despite having less than half the parameters.
Model Details
| Parameter | Value |
|---|---|
| Parameters | 600M |
| Architecture | LLaMA-style (RoPE, SwiGLU, RMSNorm) |
| Training | 39B tokens (70% EN, 30% FR) |
| Context Length | 2048 tokens |
| Vocab | 50,000 (SentencePiece BPE) |
| Framework | JAX/Flax on TPU v4-32 |
Julian Model Family
| Model | Type | Training | HellaSwag |
|---|---|---|---|
| JULIAN-100M | Base | 4.5B tokens | ~33% |
| julian-600m-10b | Base | 10B tokens | 45.8% |
| Julian-600M-40B | Base | 39B tokens | 53.5% (Current) |
| julian-600m-40b-instruct-sft100k | SFT | +2.47M (100K steps) | 41.6% |
Trained with Google TPU Research Cloud (TRC) program | Research Paper