Skip to content

Olmo 3 & Olmo Hybrid

For details on reproducing models, see:

Tokenizer Settings

When releasing multiple models (instruct and think) and using two different codebases for post-training (allenai/olmo-core for SFT and allenai/open-instruct for DPO and RL), there are many steps needed to get exact chat templates right. The final step is getting the chat templates right for public release, which can entail different system prompts to maintain model identity.

This document is a reference for the settings used for Olmo 3, based on the best available information.

Olmo 3 Instruct Models (7b, 32b): Tokenized AND intermediately evaluated with allenai/olmo-3-tokenizer-instruct-dev, released with:

  • 7b instruct: the above tokenizer on HuggingFace (no Olmo identity modified system prompt, appropriate tool use chat template/special tokens), the above tokenizer and Olmo identity system prompt modification for demos such as Ai2's playground.
  • 32b instruct: the above tokenizer on HuggingFace for SFT and DPO, and for RL/Playground the final model has the chat template with modified Olmo system prompt on playground.
  • Reason for the above discrepancy: We needed to match the chat templates for the final models relative to what was used in demos, which in the case for these models involved system prompt edits to improve model reliability. Matching models in demos and on HuggingFace is ideal.

Olmo 3 Thinking Models:

  • 7b: Training data tokenized with the olmo_thinker_no_think_7b chat template (that has the olmo identity in the prompt), but there was a minor miscommunication in transition to the next training stages, so the DPO and RL models have a slightly different chat template, all reflected in the final released models.
  • 32b: Training data tokenized with the olmo_thinker_no_think_sft_tokenization chat template (otherwise identical, doesn't have olmo identity in the prompt), released with that chat template + the think token in add_generation_prompt.
  • Reason for the difference between 7b and 32b: we learned as we went to not have the identity baked into the prompt (so it was easier to fix at the time of the demo in the form of a system prompt) but couldn't afford to retrain 7b thinking model at that point.

Olmo 3.2+ models (also used for Olmo Hybrid):

  • Think SFT and DPO data are tokenized with the Instruct chat template allenai/olmo-3-tokenizer-instruct-dev. This template does not include <think>, which prevents <think> from being masked out during tokenization so the model learns to generate it. This remains necessary for unannotated Hub templates; see #1882.
  • Think evaluation should use allenai/olmo-3.2-tokenizer-think-dev, which is the instruct chat template plus <think> in add_generation_prompt (new models should combine tool use abilities from the instruct template with <think> for reasoning). Named 3.2 to distinguish from the original Olmo 3 think tokenizers, which did not include function calling.
  • Think release models should use allenai/olmo-3.2-tokenizer-think-release, which is the same as the think-dev template but with the Olmo identity system prompt.
  • Instruct release models should use allenai/olmo-3-tokenizer-instruct-release, which is the same as instruct-dev but with the Olmo identity system prompt. This is analogous to how think-release differs from think-dev.

Assistant labels and Olmo 3.5: SFT and DPO can use {% generation %} blocks to mark rendered assistant output. Put the block after the assistant header and before any generated <think> tag; include separate reasoning_content, answer content, serialized tool calls, and the closing token. Keep the inference-only add_generation_prompt outside the block. For Olmo 3.5, a tool call trains <|im_end|> to hand off control, while an answer trains eos_token. Keep inter-turn separator whitespace outside the block.

For external templates, last-turn labeling locates blocks between stable renders before and through the final assistant message. Turn endings must not change when later messages are appended. Several blocks per assistant message are supported; matching block and message counts alone does not establish ownership. Ambiguous ownership and EOS/ChatML closing tokens left just outside a block cause the row to be logged and dropped. Truncation is handled separately by over_length_strategy.

Other special tokens immediately following a block produce a warning: they may be a missing assistant terminator or a correctly masked next-turn header. External annotations still need review for header exclusion and complete content coverage; generation blocks are not yet checked against all of the prefix path's content invariants. With left truncation, drop removes truncated rows, while terminate preserves the intact tail instead of replacing an existing tool-handoff token with EOS.

Generation annotations preserve rendered text but are saved with the tokenizer and by scripts/tokenizers/export_chat_template.py. Consumers must support the generation extension (as Transformers does); a plain Jinja environment cannot render these tags without that extension. The Olmo 3.5 tokenizer includes these annotations in both its source template and synced config. Use it with generation-range-aware labeling code; older prefix-based trainers can still mask out the opening <think>. Other unannotated Hub templates still require the workaround above. The offline Olmo 3.5 fixture checks the same generation-range contract.

Use Olmo 3.5 tokenizer revision 8b9717061fae09d5be814373d189919a62a9a00d or a later revision containing these annotations. Older revisions fall back to prefix labeling and can mask the opening <think> even with the updated trainer. #1882 intentionally remains open for unannotated Hub templates. The range renderer, transformers.utils.chat_template_utils.render_jinja_template, is a Transformers-internal API; rerun the rendering and labeling tests when upgrading Transformers. The separate tokenizer-loader discrepancy is tracked in #1896.

Note on chat_template.jinja vs tokenizer_config.json: When a HuggingFace repo contains both a chat_template.jinja file and a chat_template field in tokenizer_config.json, transformers prioritizes chat_template.jinja. Keep both in sync, or only use one. The diff_tokenizers.py script compares both files.


There are two main issues that lead to all the floating chat templates: one, the token chopping in the tokenization script where our code incorrectly masks the first token as part of the prompt, and two, the identity issue which means we should train and release with different system prompts.

TLDR until these two issues are resolved:

To verify that two tokenizer repos differ only where expected, use the diff tool:

python scripts/tokenizers/diff_tokenizers.py allenai/olmo-3-tokenizer-instruct-dev allenai/olmo-3-tokenizer-instruct-release