A model that thinks needs vLLM's --reasoning-parser to split that
thinking out of the answer. Getting it wrong fails in two ways, and only one
of them is visible.
Missing parser: the user sees the model's thinking in the
reply — obvious, and quickly fixed. Wrong parser: it claims
the entire output, the request finishes with stop, and
content comes back empty. Nothing errors.
This reads the repo's chat template only — a few kilobytes, fetched from your browser — and asks the one question that decides it: does the assistant's turn carry a marker some parser closes on?
Two models from one family, one release apart, answer differently:
Qwen3.6-35B-A3B emits <think> and needs the
qwen3 parser; Qwen3-VL-30B emits nothing of the sort
and is broken by it. A quantized pack inherits its source's template,
so the same split runs through every repack of them.
The markers below were read out of vLLM's vllm/reasoning/ modules,
not from documentation — a parser matches what its code matches. Only markers
that discriminate a reasoning section are listed: an earlier
version of this table included <|end|>, gpt-oss's turn
terminator, and recommended a parser for a model that does no thinking at all.
| Parser | Closes on |
|---|