Skip to content

[Bug] Model Randomly Outputting "The User's Message Is Empty" Thinking Blocks #12327

Description

@largeapple52

Issue

There have been a bunch of times now where the model just randomly outputs a thinking block that contains something like "The user's message is empty — there's no actual query content, just a conversation start marker. There's nothing to reason about or solve. I should respond appropriately, but the system requires me to invoke a function call in response. Since there's no prior thinking of substance, I'll use the audit tool with what little reasoning I have.".

I have just got this thinking block from the model that you see in the image, from the first prompt of the chat that I sent to the model. The chat is a part of an Unsloth sandbox project. The model just randomly writes these thinking blocks where it says the user's message is empty and it does not have context but then just carries on normally after it outputs that thinking block like nothing happened.

Image Transcription

User message: Read ./README.md

Model response:

<think>The user's message is empty — there's no actual query content, just a conversation start marker. There's nothing to reason about or solve. I should respond appropriately, but the system requires me to invoke a function call in response. Since there's no prior thinking of substance, I'll use the audit tool with what little reasoning I have.</think>
Image

LLama Logs

This message appeared right after prompt processing has finished and has also shown up a few messages later:

W srv    operator(): got exception: {"error":{"code":500,"message":"\n------------\nWhile executing CallExpression at line 159, column 24 in source:\n... not messages %}↵    {{- raise_exception('No messages provided.') }}↵{%- endif %...\n

Is this a JINJA template issue, an Unsloth issue or a llama.cpp issue?

Server Logs

{"timestamp": "2026-09-30T02:44:05.155323Z", "level": "info", "event": "engine_stats", "gen_tok_s": 0.0, "prompt_tok_s": 0.0, "running": 1, "waiting": 0, "decode_calls_s": 0.1}
{"timestamp": "2026-09-30T02:44:06.132574Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/rag/linked-folders", "status_code": 200, "process_time_ms": 11.2}
{"timestamp": "2026-09-30T02:44:09.125181Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT", "status_code": 200, "process_time_ms": 7.02}
{"timestamp": "2026-09-30T02:44:09.152473Z", "level": "info", "event": "request_completed", "method": "PUT", "path": "/api/chat/threads/__LOCALID_t1YZDoT/messages", "status_code": 200, "process_time_ms": 12.83}
{"timestamp": "2026-09-30T02:44:10.179352Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT/forks", "status_code": 200, "process_time_ms": 15.64}
{"timestamp": "2026-09-30T02:44:15.158381Z", "level": "info", "event": "engine_stats", "gen_tok_s": 0.0, "prompt_tok_s": 0.0, "running": 1, "waiting": 0, "decode_calls_s": 0.2}
{"timestamp": "2026-09-30T02:44:17.228724Z", "level": "info", "event": "request_completed", "method": "PUT", "path": "/api/chat/threads/__LOCALID_t1YZDoT/messages", "status_code": 200, "process_time_ms": 13.31}
{"timestamp": "2026-09-30T02:44:25.161357Z", "level": "info", "event": "engine_stats", "gen_tok_s": 0.0, "prompt_tok_s": 0.0, "running": 1, "waiting": 0, "decode_calls_s": 0.1}
{"timestamp": "2026-09-30T02:44:25.411100Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT", "status_code": 200, "process_time_ms": 7.19}
{"timestamp": "2026-09-30T02:44:25.438227Z", "level": "info", "event": "request_completed", "method": "PUT", "path": "/api/chat/threads/__LOCALID_t1YZDoT/messages", "status_code": 200, "process_time_ms": 12.49}
{"timestamp": "2026-09-30T02:44:26.472899Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT/forks", "status_code": 200, "process_time_ms": 22.05}
{"timestamp": "2026-09-30T02:44:33.601363Z", "level": "info", "event": "request_completed", "method": "PUT", "path": "/api/chat/threads/__LOCALID_t1YZDoT/messages", "status_code": 200, "process_time_ms": 12.86}
{"timestamp": "2026-09-30T02:44:35.163762Z", "level": "info", "event": "engine_stats", "gen_tok_s": 0.0, "prompt_tok_s": 0.0, "running": 1, "waiting": 0, "decode_calls_s": 0.1}
{"timestamp": "2026-09-30T02:44:36.718475Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/rag/linked-folders", "status_code": 200, "process_time_ms": 13.45}
{"timestamp": "2026-09-30T02:44:41.722013Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT", "status_code": 200, "process_time_ms": 9.09}
{"timestamp": "2026-09-30T02:44:41.739240Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT/messages", "status_code": 200, "process_time_ms": 6.68}
{"timestamp": "2026-09-30T02:44:41.767346Z", "level": "info", "event": "request_completed", "method": "PUT", "path": "/api/chat/threads/__LOCALID_t1YZDoT/messages", "status_code": 200, "process_time_ms": 9.84}
{"timestamp": "2026-09-30T02:44:42.787616Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT/forks", "status_code": 200, "process_time_ms": 10.11}
{"timestamp": "2026-09-30T02:44:44.613635Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/models/list", "status_code": 200, "process_time_ms": 15.36}
{"timestamp": "2026-09-30T02:44:44.620464Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/chat/threads/__LOCALID_t1YZDoT/messages", "status_code": 200, "process_time_ms": 19.88}
{"timestamp": "2026-09-30T02:44:44.633858Z", "level": "info", "event": "request_completed", "method": "GET", "path": "/api/inference/status", "status_code": 200, "process_time_ms": 34.9}
{"timestamp": "2026-09-30T02:44:45.166209Z", "level": "info", "event": "engine_stats", "gen_tok_s": 0.0, "prompt_tok_s": 138.9, "running": 1, "waiting": 0, "decode_calls_s": 9.5}

Environment

Where are you using Unsloth?

  • Unsloth desktop application
  • Unsloth web UI (unsloth studio)
  • Unsloth CLI
  • Python package or notebook
  • Colab or Kaggle

Operating system and version:
Ubuntu 26

GPU model(s) and accelerator backend:
NVIDIA CUDA

Versions:
Unsloth Version
GitHub Main
Package Version
2026.9.12

Model and operation involved:
Model: https://huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF/tree/main/Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64
JINJA template: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates/blob/main/chat_template.jinja (v22.5)

Llama settings:
Arguments: -lzm on --cache-prompt --no-context-shift
KV cache: BF16
Speculative Decoding: Ngram

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions