Olmo 3 & Olmo Hybrid
For details on reproducing models, see:
Tokenizer Settings
When releasing multiple models (instruct and think) and using two different codebases for post-training (allenai/olmo-core for SFT and allenai/open-instruct for DPO and RL), there are many steps needed to get exact chat templates right. The final step is getting the chat templates right for public release, which can entail different system prompts to maintain model identity.
This document is a reference for the settings used for Olmo 3, based on the best available information.
Olmo 3 Instruct Models (7b, 32b): Tokenized AND intermediately evaluated with allenai/olmo-3-tokenizer-instruct-dev, released with:
- 7b instruct: the above tokenizer on HuggingFace (no Olmo identity modified system prompt, appropriate tool use chat template/special tokens), the above tokenizer and Olmo identity system prompt modification for demos such as Ai2's playground.
- 32b instruct: the above tokenizer on HuggingFace for SFT and DPO, and for RL/Playground the final model has the chat template with modified Olmo system prompt on playground.
- Reason for the above discrepancy: We needed to match the chat templates for the final models relative to what was used in demos, which in the case for these models involved system prompt edits to improve model reliability. Matching models in demos and on HuggingFace is ideal.
Olmo 3 Thinking Models:
- 7b: Training data tokenized with the
olmo_thinker_no_think_7bchat template (that has the olmo identity in the prompt), but there was a minor miscommunication in transition to the next training stages, so the DPO and RL models have a slightly different chat template, all reflected in the final released models. - 32b: Training data tokenized with the
olmo_thinker_no_think_sft_tokenizationchat template (otherwise identical, doesn't have olmo identity in the prompt), released with that chat template + the think token inadd_generation_prompt. - Reason for the difference between 7b and 32b: we learned as we went to not have the identity baked into the prompt (so it was easier to fix at the time of the demo in the form of a system prompt) but couldn't afford to retrain 7b thinking model at that point.
Olmo 3.2+ models (also used for Olmo Hybrid):
- Think SFT and DPO data are tokenized with the Instruct chat template
allenai/olmo-3-tokenizer-instruct-dev. This template does not include<think>, which prevents<think>from being masked out during tokenization so the model learns to generate it. This remains necessary for unannotated Hub templates; see #1882. - Think evaluation should use
allenai/olmo-3.2-tokenizer-think-dev, which is the instruct chat template plus<think>inadd_generation_prompt(new models should combine tool use abilities from the instruct template with<think>for reasoning). Named3.2to distinguish from the original Olmo 3 think tokenizers, which did not include function calling. - Think release models should use
allenai/olmo-3.2-tokenizer-think-release, which is the same as the think-dev template but with the Olmo identity system prompt. - Instruct release models should use
allenai/olmo-3-tokenizer-instruct-release, which is the same asinstruct-devbut with the Olmo identity system prompt. This is analogous to howthink-releasediffers fromthink-dev.
Assistant labels and Olmo 3.5: SFT and DPO can use {% generation %} blocks to mark rendered assistant output. Put the block after the assistant header and before any generated <think> tag; include separate reasoning_content, answer content, serialized tool calls, and the closing token. Keep the inference-only add_generation_prompt outside the block. For Olmo 3.5, a tool call trains <|im_end|> to hand off control, while an answer trains eos_token. Keep inter-turn separator whitespace outside the block.
For external templates, last-turn labeling locates blocks between stable renders before and through the final assistant message. Turn endings must not change when later messages are appended. Several blocks per assistant message are supported; matching block and message counts alone does not establish ownership. Ambiguous ownership and EOS/ChatML closing tokens left just outside a block cause the row to be logged and dropped. Truncation is handled separately by over_length_strategy.
Other special tokens immediately following a block produce a warning: they may be a missing assistant terminator or a correctly masked next-turn header. External annotations still need review for header exclusion and complete content coverage; generation blocks are not yet checked against all of the prefix path's content invariants. With left truncation, drop removes truncated rows, while terminate preserves the intact tail instead of replacing an existing tool-handoff token with EOS.
Generation annotations preserve rendered text but are saved with the tokenizer and by scripts/tokenizers/export_chat_template.py. Consumers must support the generation extension (as Transformers does); a plain Jinja environment cannot render these tags without that extension. The Olmo 3.5 tokenizer includes these annotations in both its source template and synced config. Use it with generation-range-aware labeling code; older prefix-based trainers can still mask out the opening <think>. Other unannotated Hub templates still require the workaround above. The offline Olmo 3.5 fixture checks the same generation-range contract.
Use Olmo 3.5 tokenizer revision 8b9717061fae09d5be814373d189919a62a9a00d or a later revision containing these annotations. Older revisions fall back to prefix labeling and can mask the opening <think> even with the updated trainer. #1882 intentionally remains open for unannotated Hub templates. The range renderer, transformers.utils.chat_template_utils.render_jinja_template, is a Transformers-internal API; rerun the rendering and labeling tests when upgrading Transformers. The separate tokenizer-loader discrepancy is tracked in #1896.
Note on chat_template.jinja vs tokenizer_config.json: When a HuggingFace repo contains both a chat_template.jinja file and a chat_template field in tokenizer_config.json, transformers prioritizes chat_template.jinja. Keep both in sync, or only use one. The diff_tokenizers.py script compares both files.
There are two main issues that lead to all the floating chat templates: one, the
TLDR until these two issues are resolved:
allenai/olmo-3-tokenizer-instruct-devis the primary chat template for tokenizing both instruct and think models that have tool use abilities.- For Instruct evaluation/training, use
allenai/olmo-3-tokenizer-instruct-dev. For release, useallenai/olmo-3-tokenizer-instruct-release(adds Olmo identity). - For Think SFT and DPO tokenization, use
allenai/olmo-3-tokenizer-instruct-dev(avoids the<think>masking bug). For Think evaluation and RL prompts, useallenai/olmo-3.2-tokenizer-think-dev(adds<think>toadd_generation_prompt). For release, useallenai/olmo-3.2-tokenizer-think-release(adds Olmo identity).
To verify that two tokenizer repos differ only where expected, use the diff tool:
python scripts/tokenizers/diff_tokenizers.py allenai/olmo-3-tokenizer-instruct-dev allenai/olmo-3-tokenizer-instruct-release