Gemma 2 2B, domain-instruction-tuned on legal/financial tasks: summarize, extract, rewrite, classify, draft, obey format constraints.
2.6B
parameters
6,461
instructions
10
task types
1.90
val ppl
Validation metrics along the Gemma 2 2B lineage
Each perplexity is measured on that stage's own validation set, so read the trend as 'how well the model fits its own stage's data', not as one curve on one dataset. DPO and RLAIF optimize preferences rather than likelihood, so they log preference margin and reward instead of perplexity. Click a stage to open that model.
An instruction-tuned stage of the Gemma 2 2B: the closed-book QA model was fine-tuned on ~6.5k domain-grounded synthetic instructions (summarize, extract, rewrite in plain English, classify, explain, draft, enumerate, format-constrained answers), every example compliance- and groundedness-judged. Lineage: base → QA SFT → instruction SFT.
Served scale-to-zero on Modal, so the first request may take ~20–60s while
the model wakes.