</think> every response

#36
by pathosethoslogos - opened

Screenshot_20260804_194925_Fennec

Unsure why it's doing this, but this is also happening in INT4, NVFP4, DFlash versions, etc. I searched around for answers, but no luck. I tried changing chat template as well. This is on DGX Spark via vLLM.

Interesting that the last discussion talks about instead. There is definitely something wrong here.

I am seeing this too when using open webUI.

Probably using the wrong reasoning parser.

Probably using the wrong reasoning parser.

Using poolside_v1 results in this.

I had the same issue, solution was to enable thinking explicitly. Here is my full startup script for reference (fitting on single DGX Spark):
export CUTE_DSL_ARCH=sm_121a # arch string for FP4 kernel JIT
export PATH=/usr/local/cuda/bin:$PATH # nvcc for JIT
export MAX_JOBS=4 # cap JIT fan-out; see warning below
source ~/vllm025/bin/activate

vllm serve poolside/Laguna-S-2.1-NVFP4
--speculative-config '{"model":"poolside/Laguna-S-2.1-DFlash-NVFP4","num_speculative_tokens":7,"method":"dflash"}'
--default-chat-template-kwargs '{"enable_thinking": true}'
--enable-auto-tool-choice
--tool-call-parser poolside_v1
--reasoning-parser poolside_v1
--max-model-len 262144
--gpu-memory-utilization 0.90
--kv-cache-dtype fp8
--max-num-batched-tokens 2048
--max-num-seqs 16
--host 0.0.0.0 --port 8000

I had the same issue, solution was to enable thinking explicitly.

Interesting. I wonder what makes it different than having it enabled implied, by default. I'll have to try this later.

Sign up or log in to comment