Julian-600M Base

600M parameter language model trained from scratch on 39B tokens (70% EN, 30% FR) on Google Cloud TPUs with JAX.

This is the base model (text completion, not instruction-following). Start typing and the model will continue your text.

Model Card | Instruct Versions | Paper

50 500
0.1 1.5
10 100
0.1 1
1 2
Try these examples
Start typing (the model will continue) Max Tokens Temperature Top-K Top-P Repetition Penalty

Benchmark Results (zero-shot)

Benchmark Score
HellaSwag 53.5% (acc_norm)
PIQA 66.8% (acc)
LAMBADA 37.3% (acc)
ARC-Easy 54.0% (acc)
ARC-Challenge 27.6% (acc_norm)
WinoGrande 52.7% (acc)
BoolQ 57.9% (acc)

Outperforms OPT-1.3B (41.5% HellaSwag) despite having less than half the parameters.

Model Details

Parameter Value
Parameters 600M
Architecture LLaMA-style (RoPE, SwiGLU, RMSNorm)
Training 39B tokens (70% EN, 30% FR)
Context Length 2048 tokens
Vocab 50,000 (SentencePiece BPE)
Framework JAX/Flax on TPU v4-32

Julian Model Family

Model Type Training HellaSwag
JULIAN-100M Base 4.5B tokens ~33%
julian-600m-10b Base 10B tokens 45.8%
Julian-600M-40B Base 39B tokens 53.5% (Current)
julian-600m-40b-instruct-sft100k SFT +2.47M (100K steps) 41.6%

Trained with Google TPU Research Cloud (TRC) program | Research Paper